Low-Incidence Audiences: What a Synthetic Panel Knows
Methodology · 11 min read
TL;DR: Specifying a rare audience on a synthetic panel costs nothing, which is exactly why it needs a rule. In a human study, screening cost rises with the inverse of incidence, and that cost quietly checked whether the group was documented well enough to sample. A census-grounded panel removes the cost and the check together. This guide gives a three-tier test: audiences published directly by a national statistical office, audiences adjacent to published variables, and audiences defined by behaviour with no public statistical base. It shows what each tier supports, why the third tier fails silently rather than loudly, and how to use a synthetic panel as a rehearsal harness for the expensive human study you still have to run.
Can a synthetic panel represent a low-incidence audience?
Sometimes. A census-grounded synthetic panel can represent a low-incidence audience when the attribute defining it is published by a national statistical office at that level of detail. When the defining attribute is a purchase, a brand relationship, or an attitude with no public statistical base, the panel has no anchor and produces plausible text rather than evidence.
Every research plan for a rare audience starts with the same question in a different form. Can we even reach these people, and what will it cost?
In a human study that question has a price attached, and the price does useful work. It forces you to check whether the group is documented well enough to sample before you commit a budget. Remove the price and the check goes with it. A synthetic panel will accept any audience definition you type, including definitions no data source can support.
So the useful question is not whether the panel can produce answers for a rare audience. It always can. The question is what those answers are anchored to. A census-grounded panel is calibrated against tables a national statistical office publishes: age, gender, region, income, education, household composition. Where the tables stop, calibration stops.
That boundary is the subject of this post. It sits next to the failure modes where synthetic evidence is not valid, but it is a different check, and it comes earlier. Failure modes ask what kind of question you are asking. This one asks who you claimed to be asking.
Why does a rare audience cost so much to reach in a human study?
Because you pay for everyone you screen out. Kalton and Anderson showed in 1986 that the screening burden for a rare population scales with the inverse of its prevalence. At 5 percent incidence you screen about 20 people per complete. At 1 percent you screen about 100. The questionnaire cost is unchanged; the recruitment cost is what moves.
The cost structure of a low-incidence study is not really a cost of interviewing. It is a cost of screening out.
Graham Kalton and Dallas Anderson set the arithmetic out in the Journal of the Royal Statistical Society in 1986. To find members of a rare population by screening a general sample, the number of screening contacts you need scales with the inverse of that group's prevalence.¹ Questionnaire length barely matters. The incidence rate decides the budget.
The table below shows what that means for a target of 200 completed interviews.
Pew Research Center has documented the same problem from the design side. Its 2023 review of approaches to surveying small populations notes that screening a general population sample becomes impractical as the group gets smaller, which pushes researchers toward list frames, oversampling and non-probability sources, each with its own trade-off.² None of those options is free of bias and all of them take time.
That is the honest case for looking at a synthetic panel first. Not because it is more accurate, but because the realistic alternative for a one percent audience is often no research at all.
Which rare audiences are actually inside a census-grounded panel?
Three tiers. Tier one is a cell the statistical office publishes directly, so the panel is calibrated on it. Tier two is a cell the office does not publish but that correlates with variables it does, so the panel infers it. Tier three is defined by behaviour with no public statistical base, and the panel cannot reach it.
Rarity by itself is not the problem. A group can be one percent of a country and still be well described in official statistics. Another group can be ten percent of a country and invisible in every published table.
The test that separates them has three parts. Is the defining attribute a variable the national statistical office collects? Is it published at the level of cross-tabulation your audience needs? And is the behaviour you plan to ask about documented anywhere in public statistics?
Lisa Argyle and colleagues showed in Political Analysis in 2023 that conditioning a language model on detailed demographic backstories reproduces recognisable subgroup patterns from human survey data, a property they named algorithmic fidelity.⁴ That result is why tier one works. It also marks where it stops. The fidelity came from conditioning on attributes described in aggregate somewhere in the training data. Nothing in it promises fidelity for attributes nobody has published.
Two practical notes. Published does not mean available at every cross. Statistical offices suppress or widen small cells to protect confidentiality and to avoid releasing an estimate they cannot stand behind. The United States Census Bureau's 2020 guidance for American Community Survey users is explicit that estimates for small populations carry margins of error large enough to make them unreliable for many uses, and asks users to check the margin before drawing a conclusion.³ Your panel inherits that limit rather than repairing it.
The second note is that the tier is a property of the audience definition, not of the topic. The same study can be tier one for one cell and tier three for the cell next to it.
What happens when you specify a cell the census does not publish?
The model fills the gap from its training data, and that is where the documented errors live. Wang and colleagues reported in Nature Machine Intelligence in 2025 that language models standing in for human participants can misportray and flatten identity groups. Santurkar and colleagues found at ICML in 2023 that opinion alignment varies sharply by group. Both effects concentrate in small cells.
Two findings matter more here than any vendor claim, and both are about subgroups rather than about averages.
Angelina Wang, Jamie Morgenstern and John Dickerson published a study in Nature Machine Intelligence in 2025 on using language models as substitutes for human participants. They found that models misportray identity groups and flatten the variation inside them, presenting a group as more homogeneous than it is.⁵ Flattening is invisible in a topline. It shows up exactly where a low-incidence study looks, inside the small cell.
Shibani Santurkar and colleagues reported a related result at ICML in 2023. Comparing model output against United States public opinion data, they found substantial and uneven misalignment across demographic groups, with some groups reflected considerably worse than others.⁶ Uneven is the operative word. An error that varies by group in ways you cannot predict is an error you cannot correct for.
Put the two together and the risk for a rare audience is specific. The panel will return a tight, coherent, low-variance picture of the group it knows least about, and the tightness will read as signal. Wide disagreement would at least be a warning. Consensus is not.
This is why replicate runs matter more here than anywhere else. Run the same thin cell three times with fresh persona draws. If the spread across runs is wider than the effect you are chasing, you have your answer without needing a methodological argument. The rules for choosing sample size by saturation still apply, with one caveat: saturation on a thin cell arrives early and means less.
How do you use a synthetic panel for an audience it cannot represent?
Rehearse the study instead of running it. Use the panel to test whether your screener separates the audience from the near miss, whether the questionnaire holds up, and how long it runs. Then buy the human sample once, with a screener that already works. The panel de-risks the expensive fieldwork rather than replacing it.
There is a use for tier three worth more than the study you cannot run, and it is the one most teams skip.
Low-incidence fieldwork usually fails in the field, not in the analysis. The screener over-qualifies or under-qualifies. The incidence assumption turns out to be wrong. The questionnaire runs long once the qualified respondent finally arrives, and you have no cheap way to find any of that out before the invoice.
A synthetic panel lets you rehearse it first.
1. Write the screener exactly as you would send it to a recruiter, same wording, same order.
2. Build two synthetic cells: one matching your intended audience on every census attribute you can specify, and one deliberate near miss that differs on a single attribute.
3. Run the screener on both. If it qualifies the near miss at a similar rate, the screener is not separating anything, and it would have failed in field at full cost.
4. Run the full questionnaire on the qualifying cell and read the rationales. Ambiguous items and unanswerable questions surface here.
5. Fix the instrument, then buy the human sample once.
Nothing in that sequence claims the synthetic answers describe your rare audience. The panel is a test harness for the instrument, which is a claim it can support.
PersonaHive is built for the first half of that loop. Personas are grounded in national census data, country by country across nine countries, and validated against real surveys, so a tier one cell is specified against the same variables the statistical office publishes rather than against a free-text description. Every response ships with a written rationale, which is what makes step four readable: you can see why a persona qualified, not only that it did. The rules for writing survey questions for synthetic personas cover the instrument side in more depth.
The smallest useful next step is a tier test on one audience. Take the audience definition you were about to send to a recruiter and mark each defining attribute as published, adjacent or behavioural. If every attribute is published, run it. If one is behavioural, you have found the part of the study that has to stay with people. PersonaHive's free tier includes 250 credits and needs no card.
When does a low-incidence answer still have to come from people?
When the audience sits in tier three, when the decision needs a level rather than a ranking, and when the rare group is rare for a reason the census does not record. A panel calibrated on published margins cannot tell you what a group that public statistics do not describe actually believes.
Four cases, and the first is the tier test failing.
The second is a decision that needs a level. Market sizing, incidence estimation and volume forecasting all ask how many, and how many is what a synthetic panel estimates least well. On a rare audience it is worse, because the published base you would calibrate against is thin to begin with.
The third is a group that is rare for a reason official statistics do not record. Recent switchers, lapsed users of one brand, people in the middle of a specific life event. The census records who they are demographically and nothing about the state that defines them.
The fourth is any study whose finding will be attributed to that group in public. A flattened portrait of a small group presented as research about that group is the failure Wang and colleagues described, and it carries a reputational cost separate from being wrong.⁵
In each case the correction is human data, not a larger synthetic sample. Country coverage matters here too, because a cell that is tier one in one market can be tier two in another when the two statistical offices publish different tables. Why census grounding has to be country by country sets out what that means for multi-market work.
As of 2026 no published validation work establishes how far a synthetic panel can be trusted on a cell its calibration source does not cover. Until that exists, the tier test is the conservative position, and it is the one you can defend in a methods review.
Frequently asked questions about low-incidence audiences on synthetic panels
Short answers to five questions that come up when a rare audience lands on a synthetic panel: whether B2B niches behave differently, what to do when a cell is tier one in one market and tier two in another, whether more personas fix a thin cell, how to report a tier two result, and what to tell a stakeholder who asks why the panel cannot simulate anyone.
**Do B2B niches follow the same three tiers?** Partly. Firmographics such as company size, sector and region appear in business registers and official statistics, so they behave like tier one. Job role, buying authority and vendor relationships are rarely published at any useful cross and behave like tier two or tier three.
**A cell is tier one in one market and tier two in another. What then?** Report them separately and do not average. A cross-market topline mixing a calibrated cell with an inferred one hides which half the finding came from.
**Will more personas fix a thin cell?** No. Adding personas narrows sampling error, and sampling error is not what is wrong with a thin cell. The problem is the missing anchor, and it does not shrink with n.
**How should a tier two result appear in a report?** As a hypothesis, with the tier stated next to it, the replicate range shown, and the human study it is meant to shape named underneath. Tier two output belongs in the design section, not the findings section.
**A stakeholder asks why the panel cannot simulate anyone. What do I say?** That it can produce answers for anyone, and can only be checked for people the country's statistics describe. The limit is the published data, not the model.
Sources
- Sampling Rare Populations — Graham Kalton and Dallas W. Anderson, Journal of the Royal Statistical Society, Series A
- When surveying small populations, some approaches are more inclusive than others — Pew Research Center, Pew Research Center
- Understanding and Using American Community Survey Data: Understanding Error and Determining Statistical Significance — United States Census Bureau, United States Census Bureau
- Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting and David Wingate, Political Analysis, Cambridge University Press
- Large language models that replace human participants can harmfully misportray and flatten identity groups — Angelina Wang, Jamie Morgenstern and John P. Dickerson, Nature Machine Intelligence
- Whose Opinions Do Language Models Reflect? — Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang and Tatsunori Hashimoto, Proceedings of the 40th International Conference on Machine Learning (ICML 2023)