Combining Synthetic and Human Respondents: The Math

Methodology · 11 min read

TL;DR: Combining synthetic and human respondents usually means pooling both into one dataset and averaging. That hides the synthetic panel's offset instead of removing it, and the more synthetic rows you add the more the offset dominates. The defensible move is rectification: run a small human sample on the identical instrument, measure the gap between synthetic and human answers in matched census cells, subtract that gap from the synthetic estimate, and widen the interval to price the correction. The width of that interval is set by how much the two sources disagree, not by how many personas you ran. This guide gives the estimator, the allocation rule, the cell-matching constraint, and the cases where no correction is available.

What does it mean to combine synthetic and human respondents?

Three distinct designs get called the same thing. Pooling stacks both sources into one dataset and averages. Routing runs synthetic upstream and human downstream as two separate readings. Rectification uses a small human sample to measure the synthetic panel's offset, then corrects the synthetic estimate and reports one number with a wider interval.

The word combine is doing too much work in most vendor material.

Pooling is the default because it is the easiest thing a spreadsheet can do. Stack 600 synthetic rows under 150 human rows, average the column, report the result. The number you get is a blend of two different measurement processes, and the blend weight is whichever sample happened to be larger.

Routing is the design this category already recommends, including our own head-to-head comparison of synthetic users and real respondents. Screen broadly on synthetic, confirm the shortlist on human. It is sound advice, and it has a cost: you finish with two numbers and no principled way to state one.

Rectification is the third option and the only one with a statistical literature behind it. It comes from prediction-powered inference, published in Science in 2023 as a method for producing valid confidence intervals when a large set of model predictions is paired with a small gold-standard sample.¹ Survey research has since adopted the same machinery under the name rectification.³

Only one of the three produces a number you can defend line by line.

Why does averaging synthetic and human responses give you the wrong number?

Averaging treats the two sources as interchangeable draws from one population. They are not. A synthetic panel carries a systematic offset on most questions, and pooling averages that offset into the estimate rather than removing it. Doubling the synthetic sample makes the pooled number more precise and further from the human reading.

Work it through on an illustrative purchase-intent read.

Say the synthetic panel returns 41 percent top-two-box on a concept and 150 human completes return 34 percent. Pool them and you land near 39 percent, closer to the synthetic figure because that sample is four times larger.

Now field 2,000 personas instead of 600. The pooled figure moves to roughly 40.5 percent. Your interval narrowed and your estimate walked further from the human reading.

The offset itself is documented. Bisbee and colleagues examined language model responses as substitutes for human survey data and reported divergences that are unstable across model versions.⁵ Argyle and colleagues, who established that conditioned language models can mirror human samples on some tasks, were explicit that the correspondence is uneven.⁶ Neither result says synthetic responses are useless. Both say the offset is real, structured, and available to measure.

Measuring it is the entire move. The offset is not noise to be averaged away.

How do you correct a synthetic estimate with a small human sample?

Run the identical instrument on a small human sample matched to the same census cells as the panel. Compute the gap between the synthetic and human answers in those cells. Subtract the gap from the full synthetic estimate. Report the corrected figure with an interval that reflects how much the gap varied.

Four steps, in order.

1. Field the same instrument twice. Same wording, scale, order and randomization, on the synthetic panel and on the human sample. Any wording difference contaminates the gap you are about to measure, so the rules for writing synthetic questionnaires apply to both sides.

2. Match on cells, not on people. You cannot pair a persona to a person, so pairing happens at the level of census cells: age band by gender by region, or whatever quota frame the study uses. This is the real constraint of the method, and most write-ups skip it.

3. Compute the gap per cell, then aggregate. In the illustration above, synthetic reads 40 percent in the matched cells against 34 percent human, so the gap is plus 6 points. The full synthetic estimate of 41 percent becomes a corrected 35 percent.

4. Price the correction in the interval. The corrected figure carries two sources of uncertainty: sampling error on the synthetic mean, which is small because the panel is large, and estimation error on the gap, which is set by your human sample size and by how much the gap moves across cells. The second term dominates, and PPI++ sets out the weighting between them.²

A counterintuitive result falls straight out of that arithmetic. Ten thousand personas with 150 human completes buys about the same interval as 600 personas with 150 human completes.

How many human respondents do you actually need?

Enough to estimate the gap in every cell you intend to correct. That is fewer than a standalone human study and more than a token check. A total-market correction can rest on roughly a hundred completes. Segment-level corrections need their own completes in each segment you report.

The allocation question has live research attached to it, including 2026 work on rectification difficulty and optimal sample allocation in language-model-augmented surveys.⁴ The working rule is simpler than the papers.

You need human completes wherever you want a corrected number. One national figure means one gap. Four segments means four gaps, each with its own completes. A 150-complete sample that supports a clean national correction supports nothing at segment level once it is cut four ways.

So the design decision comes before the fieldwork order. Decide which cells will appear in the final deck, then buy human completes for those cells and no others.

The other half of the answer is the quality of the human side. A contaminated benchmark produces a contaminated correction, and the estimate inherits every bot and every AI-written open-end in the sample you measured against. Grade the human benchmark before you trust the gap. The ESOMAR and GRBN online sample quality guideline is the standard reference for what to ask a supplier.⁷

Smallest useful next step: take one question from a study you already fielded to humans, run it unchanged on the matching census-grounded panel, and compute the gap. That single number tells you more about whether synthetic research works in your category than any vendor benchmark. A free PersonaHive account includes 250 credits, which covers it.

When is no correction available?

When there is no human reading to measure the gap against, the method has nothing to work with. That covers rare-event incidence, genuinely new categories, populations too small to quota, and any question where the human answer is the thing being discovered rather than the thing being predicted.

Rectification is a correction, not a source of evidence. It needs a human reading to correct toward.

Three cases where it does not apply. Rare-event incidence, because a hundred completes will not contain the event. Genuinely new categories, where no prior human data exists and no census cell carries signal for the behaviour. And exploratory work whose output is a hypothesis, where correcting a number you were never going to quote adds cost and nothing else.

A fourth case is subtler. If the gap is unstable across cells, an aggregate correction misleads even when it computes cleanly. A panel that runs 6 points hot on average, 2 points cold in one segment and 14 points hot in another has no aggregate correction worth the name. Read the spread of the gap, not only its mean. That instability sits among the documented failure modes where synthetic evidence breaks.

Hullman and colleagues make the broader version of the argument: a simulation counts as behavioural evidence only under conditions you have to state, not by default.⁸ A correction does not change what a method can measure. It changes how honestly you can report the part it does.

What do you put in the report?

Four things: the uncorrected synthetic estimate, the human sample that produced the gap, the gap itself with its spread across cells, and the corrected estimate with its interval. A reader who disagrees with your correction should be able to recompute the uncorrected number from what you published.

Publish both numbers. Synthetic and corrected, side by side, with the gap stated. Burying the uncorrected figure invites the suspicion that the correction was chosen to land somewhere.

Publish the gap's spread. One number for the mean gap, one for how much it varied across cells. The second is what tells a reader whether the correction was a measurement or an average of unrelated things.

Publish the run record. A correction is valid for the model version, provider settings, panel definition and instrument it was measured on, and none of those hold across quarters. The six fields that make a synthetic study re-runnable are the same six that make a correction auditable.

Where PersonaHive fits is specific. Its personas are grounded in national census data on a country-specific basis, built from aggregated public statistics and validated against real surveys, across nine national panels covering the United States, Germany, France, Austria, Czech Republic, Hungary, Romania, Denmark and Finland. Cell matching is only possible when the panel is composed against the same census frame you quota a human sample on. Every response also ships with a written rationale, which is how you investigate why one cell's gap is large instead of accepting it. The published validation work against the U.S. Consumer Financial Protection Bureau 2024 National Age-Friendly Banking Survey is a worked example of a gap measured against a named public source.

For the general version of that exercise rather than the per-study correction described here, the validation study protocol covers benchmarking at aggregate, segment and question level. That grades the instrument once. This post is about the number you have to ship this quarter.

Start with one question and one gap. Open a free PersonaHive account, run a question whose human answer you already hold, and read the difference.

Frequently asked questions about combining synthetic and human respondents

**Can I weight the synthetic responses instead of correcting them?**

Weighting fixes composition, not accuracy. If the panel over-represents a cell, weighting repairs that. If the panel answers differently from people in the same cell, weighting leaves the offset untouched.

**Does this mean synthetic research needs human fieldwork to be useful?**

No. Ranking, screening and instrument testing need no correction because they do not quote absolute levels. Correction is for numbers you plan to state as levels.

**How often should the gap be re-measured?**

Whenever the model version, provider settings, panel definition or instrument changes, and at least once per tracking wave. A correction from two quarters ago is a claim about a system that no longer exists.

**Is this the same as a validation study?**

No. A validation study grades an instrument once and reports where it agrees. Rectification is a per-study estimator applied to the specific number you are about to report.

**What if the corrected estimate crosses my decision threshold and the uncorrected one does not?**

Then the study answered its question, and the answer is that the concept sits at the threshold. Widen the human sample or route the decision to fieldwork. Do not pick the number that reads better.

Sources

  • Prediction-powered inference — Anastasios N. Angelopoulos, Stephen Bates, Clara Fannjiang, Michael I. Jordan, Tijana Zrnic, Science, vol. 382, pp. 669-674 (2023)
  • PPI++: Efficient Prediction-Powered Inference — Anastasios N. Angelopoulos, John C. Duchi, Tijana Zrnic, arXiv preprint 2311.01453
  • Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification — Proceedings of ACL 2026 (Main Conference), ACL Anthology
  • Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys — arXiv preprint 2604.17267
  • Synthetic Replacements for Human Survey Data? The Perils of Large Language Models — James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, Jennifer M. Larson, Political Analysis, vol. 32, no. 4, pp. 401-416 (2024), Cambridge University Press
  • Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, David Wingate, Political Analysis (2023), Cambridge University Press
  • ESOMAR/GRBN Guideline on Online Sample Quality — ESOMAR and GRBN, ESOMAR and the Global Research Business Network
  • Validating LLM simulations as behavioral evidence — Jessica Hullman and colleagues, MU Collective, Northwestern University

Related Articles

  • How to Run a Validation Study for AI Synthetic Consumer Research — A practical methodology for validating a synthetic consumer research panel against a live national survey: what to measure, how to design a fair benchmark, and how to present the evidence to skeptical stakeholders.
  • Synthetic Users vs. Real Respondents: A Head-to-Head Comparison — Synthetic users vs. real respondents across speed, cost, bias, and reliability, and how census-calibrated personas complement traditional panels.
  • Survey Panel Data Quality: Grading Your Human Benchmark — Survey panel data quality is now a benchmark problem: bots, fraud, and AI-written answers. How to grade a human sample before you validate against it.