Significance Tests on Synthetic Panels: What Free n Hides
Methodology · 11 min read
TL;DR: A statistical significance test on a synthetic panel answers a question you did not ask. A p-value measures sampling error, and on a synthetic panel you set the sample size, so the interval narrows to whatever you are willing to spend while the model error that dominates the estimate stays fixed. Run enough personas and every gap reads significant. This guide shows the arithmetic behind that, names the error a synthetic sample actually carries, and gives the three things to report in its place: the gap in points, the range across independent replicate runs, and a decision threshold set before the run. It also names the cases where the number still has to come from people.
Can you run a statistical significance test on a synthetic panel?
You can compute one, but it will not mean what a significance test means on a human sample. A p-value measures sampling error only. On a synthetic panel you choose the sample size, so sampling error shrinks to whatever you want it to be, while the model error that dominates the result stays exactly where it was.
Every research lead who moves a study onto a synthetic panel eventually hits the same moment in a methods review. Concept A scored 62. Concept B scored 57. Is that difference real?
The honest answer starts by separating two things one number hides. A confidence interval on a survey estimate describes a single source of error: the chance that this particular draw of respondents differs from the population it came from. It says nothing about whether the instrument measures the right thing, and nothing about whether the respondents resemble the population at all.
On a human panel those other errors exist too, and survey statisticians spend most of their time on them. On a synthetic panel the balance changes. Sampling error becomes the smallest term in the equation, because you set it. Sample size is a budget decision rather than a fieldwork constraint, which is why choosing n by saturation rather than by power is the more useful frame.
The American Association for Public Opinion Research made the underlying point about nonprobability samples years before synthetic respondents existed. Its 2013 task force concluded that a margin of sampling error should not be reported for a nonprobability sample, because the calculation assumes a random draw from a defined population that never took place.¹ A synthetic panel is not a nonprobability sample in the usual sense. It is not a sample of people at all. The objection applies with more force, not less.
Why does a bigger synthetic sample make the p-value less informative?
Because the interval narrows while the error does not. At n of 400, a 95 percent interval around a 50 percent result runs about 4.9 points either side. At n of 40,000 it runs about 0.5 points. Personas cost credits rather than fieldwork, so any difference you care about eventually clears the threshold, and significance stops separating findings.
The arithmetic is fixed and worth seeing in one place. The table below shows the 95 percent interval around a single proportion sitting near 50 percent, and the smallest gap between two equally sized cells that would be called significant at the same level.
Read the last row carefully. At 40,000 personas per cell, a gap of seven tenths of one point passes a significance test. No product decision turns on seven tenths of a point, and no researcher would defend one on a human sample. The test has stopped doing the job people believe it is doing.
This creates a specific failure that is easy to walk into without noticing. A first run at n of 400 returns a gap that does not clear the bar. Rerunning at n of 8,000 costs almost nothing, and the same gap now clears it. Nothing about the market changed. Nothing about the concepts changed. The only thing that changed was the denominator, and the denominator was under your control the whole time.
The arithmetic in this table is not evidence about anything. It is a property of the formula. That is the point: a number that moves whenever you decide to spend more credits is not measuring the world.
What error does a synthetic sample actually carry?
Model error, and it does not shrink with n. Bisbee and colleagues found in Political Analysis in 2024 that language model responses understate the variance of human survey data and move with the wording of the persona prompt. Adding personas multiplies the same conditioned distribution. Ten thousand draws from a biased generator give a precise estimate of the bias.
Bisbee, Clinton, Dorff, Kenkel and Larson compared model-generated survey responses against human benchmarks and reported two findings that matter here. Synthetic responses were less variable than the humans they stood in for, and they shifted when the prompt describing the respondent was reworded.² Neither of those errors is reduced by generating more responses. Both are reproduced by it.
Dominguez-Olmedo, Hardt and Mendler-Dunner reported a related instability at NeurIPS 2024: model answers to standard survey instruments move with the ordering and labelling of the answer options, in ways that a human sample would not show.³ The measurement is sensitive to the instrument in a way the arithmetic of the confidence interval never sees.
The error budget on a synthetic study therefore has two lines. Sampling error, which you control and can drive to near zero. Model error, which you do not control and cannot reduce by spending. A significance test reports the first and stays silent on the second, which is the larger one.
AAPOR's explainer on credibility intervals makes the general version of the point: such an interval is not a margin of sampling error, and it rests on modelling assumptions that have to be stated.⁶ The same holds for any interval you put around a synthetic estimate.
What should you report instead of a p-value?
Report the size of the gap, the decision threshold you set before the run, and the spread across independent replicate runs. Run the same study three to five times on freshly drawn personas, record the topline each time, and report the range. That range is the honest uncertainty band on a synthetic estimate, and it is usually wider than any sampling interval.
The replicate range is the single most useful number a synthetic study can produce, and almost nobody reports it. It is cheap: five runs at n of 400 cost less than one run at n of 2,000, and they tell you something the larger run cannot.
If concept A leads concept B by five points in every one of five runs, you have a finding worth acting on. If the lead is five points in two runs, one point in two more, and reversed in the fifth, you have a coin flip dressed as a result. No p-value computed inside any single run would have told you that.
Record the run conditions alongside the range, or the replicate test cannot be repeated by anyone else. The six fields that make a synthetic run reproducible cover what to log. Reproducibility and validity are different claims, and the replicate range speaks to the first.
How do you set a decision threshold before the run?
Name the smallest difference that would change what you do, and write it down before you look at the output. If a five point gap moves a concept forward and a three point gap does not, then three points is noise for this decision regardless of any p-value. Wasserstein and Lazar made the general case in 2016.
The ASA statement on p-values is blunt about the limit: a p-value does not measure the size of an effect or the importance of a result, and scientific conclusions should not be based on whether a value passes a threshold.⁴ That guidance was written for human data. It applies with more force where the threshold is purchasable.
The threshold you set should come from the decision, not the data. Ask what gap would change the launch call, the pricing move, or the creative you take forward. Write that number in the study plan next to the research question. Then the analysis has one job: report the gap and the replicate range, and say whether the gap clears the line in most runs.
Gelman and Carlin give the reason this ordering matters. When estimates are noisy and true effects are small, results that clear a significance filter systematically overstate the size of the effect, and can get its direction wrong.⁵ Synthetic subgroup estimates are exactly that kind of noisy, which is one of the six failure modes where a synthetic panel matches a national topline and reverses inside a segment.
A threshold set afterwards is not a threshold. It is a description of the result you already saw.
When does the number still have to come from people?
When the decision needs a level rather than a ranking, when the gap sits inside your replicate range, and when the subgroup driving the call is small or thinly represented. Synthetic panels are strongest at ordering options and weakest at absolute magnitude, and absolute magnitude is what a forecast or a business case usually needs.
Three cases send you back to human respondents. A volume forecast or a revenue model needs a level, and a level is what a synthetic panel estimates least well. A gap inside the replicate range is not a finding at any sample size. And a call resting on a small or thinly represented subgroup sits where model error is largest.
When you need the level and not just the order, the correction runs through a human anchor rather than a bigger synthetic sample. The estimator for combining synthetic and human respondents sets out how many human completes each corrected cell needs.
PersonaHive builds personas grounded in national census data, country by country across nine countries, and validates them against real surveys. Every response ships with a written rationale, so a gap you cannot explain can be read back to the reason each persona gave. Neither makes a p-value meaningful. Both make the replicate test cheap to run and the result auditable afterwards.
The smallest useful next step is that replicate test. Take a study you have already run, run it again on a fresh persona draw, and compare the two toplines. If the gap you reported to stakeholders is smaller than the gap between your own two runs, you have learned that before anyone else does. PersonaHive's free tier includes 250 credits and needs no card.
Frequently asked questions about significance testing on synthetic panels
Short answers to the five questions that come up most in methods reviews: whether a t-test is ever valid here, how many replicate runs to budget, what to say when a stakeholder asks for a margin of error, whether temperature settings substitute for replicates, and how to report a synthetic result in a deck.
**Is a t-test on synthetic responses ever valid?** As a description of the spread inside one run, yes. As evidence that a difference would appear among people, no. Report it as a within-run statistic if you report it at all, and never as the basis for the decision.
**How many replicate runs should I budget?** Three to five for a routine screening study, and five or more when the decision is expensive or the gap is close to your threshold. The runs must draw fresh personas, not re-ask the same ones.
**A stakeholder asked for the margin of error. What do I say?** Give them the replicate range instead and explain what it covers. AAPOR's guidance that a margin of sampling error should not be reported for a nonprobability sample gives you the reference to cite.¹
**Does varying temperature or seed replace replicate runs?** No. Changing sampling settings perturbs one persona set. A replicate run rebuilds the persona draw, which is the part carrying the model error you are trying to see.
**How should a synthetic result appear in a deck?** As a rank order with the gap in points, the replicate range in brackets, and the pre-set threshold shown as a line. Absolute scores belong in the appendix with the run conditions.
Sources
- Report of the AAPOR Task Force on Non-Probability Sampling — AAPOR Task Force on Non-Probability Sampling, American Association for Public Opinion Research
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models — James Bisbee, Joshua Clinton, Cassy Dorff, Brenton Kenkel and Jennifer Larson, Political Analysis, Cambridge University Press
- Questioning the Survey Responses of Large Language Models — Ricardo Dominguez-Olmedo, Moritz Hardt and Celestine Mendler-Dunner, Advances in Neural Information Processing Systems (NeurIPS 2024)
- The ASA Statement on p-Values: Context, Process, and Purpose — Ronald L. Wasserstein and Nicole A. Lazar, The American Statistician, American Statistical Association
- Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors — Andrew Gelman and John Carlin, Perspectives on Psychological Science
- Understanding a credibility interval and how it differs from the margin of sampling error in a public opinion poll — American Association for Public Opinion Research, American Association for Public Opinion Research