Skip to main content

How we validate synthetic panels

A synthetic panel is only worth running if its predictions track what real visitors actually do. Anyone can build a panel that returns a ranking; the question a buyer should ask is whether that ranking has ever been checked against reality, and what happens when it is wrong. This page is Erdo’s answer: the public benchmark we score the panel against, the correction that turns a model’s raw guess into a calibrated effect, and the falsification test that refuses a result it cannot stand behind.

Methodology summary

Erdo validates its synthetic panels against the Upworthy Research Archive — 32,487 real headline A/B tests, each with the true click-through rate that reality measured — the same public benchmark the leading academic studies use. The panel predicts which variant will win, several times over, and its predictions are corrected against real outcomes with a calibration curve fitted so that no test ever helps score itself. A falsification test then checks whether the corrected predictions are statistically consistent with the truth: when they are not, the benchmark reports that the configuration failed rather than publishing a number it cannot defend. The result is not just an accuracy figure but a system that states, on evidence, when it knows and when it doesn’t — which is the property that matters when a prediction is about to steer real ad spend.

The benchmark: a public archive with real outcomes

You cannot check a prediction against reality without reality to check it against. The Upworthy Research Archive is the rare public dataset that provides it: for tens of thousands of real headline A/B tests run on a high-traffic news site, it records every variant that was shown and the click-through rate each one actually earned. That makes it a ground truth a synthetic panel can be graded on directly — predict the winner, then compare against the outcome the archive already measured. Two points of hygiene travel with the data. The archive is published under the Creative Commons Attribution 4.0 licence (Upworthy Research Archive, J. Nathan Matias, Kevin Munger, et al., osf.io/jd64p), and Erdo carries that attribution wherever the benchmark is reported. And a June 2024 erratum identified randomisation problems in a window of the earliest tests; every test created between 25 June 2013 and 10 January 2014 is excluded from the benchmark, so the number is measured only on tests whose outcomes can be trusted.

Prediction: ask several times, then average

A single model prediction is noisy — ask the same question twice and the answers wobble. So the panel does not read a variant once. It predicts the outcome several times over (five independent draws is the default), and averages them. Averaging several draws is a well-established way to recover a stabler estimate than any single read, and it is cheap: the whole exercise costs a few dollars and a couple of minutes per variant, against the hundreds of dollars of paid traffic it takes to read one variant for real.

Calibration: correct the scale against reality

Models are far better at direction than at magnitude — good at saying which variant is stronger, unreliable at saying by how much, and they tend to exaggerate the size of an effect. Using their raw numbers to steer spend would import that exaggeration wholesale. So Erdo never uses the raw prediction directly. It fits a calibration curve that maps predicted effects onto the real effect scale, learned from the pairs of prediction and measured outcome the benchmark provides. The curve is monotone — a higher predicted effect always maps to a higher calibrated one, because inverting that order would be fitting noise, not signal — and it is fitted with cross-validation, so the correction applied to any one test is learned only from other tests and never from the test’s own outcome. That discipline is what keeps the reported accuracy honest: a model that got to see the answer before grading itself would look better than it is.

The falsification gate: the system says when it doesn’t know

Calibration can make predictions look accurate; it cannot, by itself, prove they are trustworthy. The falsification test is the check that can. After calibration, it asks a single statistical question: are the corrected predictions consistent with being unbiased estimates of the real effects? If the corrected predictions systematically miss — too confident, biased in one direction, drifting away from the measured truth — the test detects it and the configuration fails the gate. A failed gate is not a number to be spun; it is the benchmark refusing to publish a result it cannot defend. This is the honest differentiator, and it is why Erdo publishes the method rather than only a figure. A flat accuracy claim tells you how a system did on someone’s chosen slice of data; it tells you nothing about when to distrust it. A gate that can fail tells you the system will decline to guess when its predictions stop tracking reality — and that is a harder thing to copy than any single accuracy number, because it requires the reality-paired ledger to run against in the first place.

The gate demonstrated in both directions

The gate is only meaningful if it can actually fail, so the calibration run tested it on two configurations of the same pipeline. A deliberately cheap configuration — a small model with minimal effort — scored around 0.60 directional and failed the falsification test (z = 3.69): its corrected predictions were provably inconsistent with the truth, and the gate refused them. A production-tier configuration on the same tests passed (p = 0.64), and recovered the effect scale from 0.27 to 0.68 of the true magnitude. One pipeline, two configurations, correctly rejected one and accepted the other. That discrimination — accepting the good configuration and refusing the cheap one on the same data — is the property the whole system leans on.

Current benchmark result

The full Upworthy holdout run that produces Erdo’s headline directional hit-rate — the single accuracy figure an agency report would quote — is a deliberate, staged run that has not yet been executed; this page will state that number, with its confidence interval and the gate verdict it passed, once it has. What is already established is the gate itself, demonstrated above on real archive data in both directions: the calibration correctly failed the cheap configuration and passed the production-tier one. Until the full run lands, that is the honest state of the benchmark — the method is proven and the accuracy figure is pending, and Erdo would rather say so than publish a number it has not yet measured. When the run completes, and only if the gate passes, the hit-rate drops in here as the maintained, provenance-stamped number; if the gate fails, that failure is the finding and it will be reported as such. The number is a property of the engine, produced and published platform-side, so every product that runs on the panel inherits the same validated method. White-label and agency reports may quote the methodology summary above and, once published, the benchmark figure — a client does not need to understand isotonic regression to rely on a panel whose predictions are checked against reality and whose gate declines to guess when they stop tracking it.