Market Sizing on a Synthetic Panel: What Breaks
Methodology · 10 min read
TL;DR: A market size is a chain of multiplications, and a synthetic panel can carry only one link in it honestly. You set the panel's composition at the draw, so any share you read back on a census attribute is the number you specified, returned to you. Shares the census does not publish, category usage, purchase frequency, brand repertoire, come from model priors rather than measurement. That makes incidence, penetration and volume unsafe to size on. What a census-grounded panel does carry is the work around the sizing study: screener pretesting, segment definition, and scenario ranges on parameters sourced elsewhere.
Can a synthetic panel size a market?
No, not the numbers that go into the model. A market size rests on prevalence, and prevalence is the one quantity you fix yourself when you draw a synthetic panel. The panel returns your own specification on census attributes and a language prior on everything else. Use it around the sizing study, not for it.
The request usually arrives in one sentence. How many people in Germany would buy this, and what is that worth.
Both halves look like survey questions. Neither is. A market size is an arithmetic chain, and fieldwork supplies only part of it. The rest comes from official statistics and from commercial measurement, and a synthetic panel touches neither.
The trap is that the panel answers anyway. Ask a census-grounded panel what share of German households bought a premium dishwasher tablet last quarter and you get a clean percentage, with a written rationale attached to every response. Nothing in the output marks that percentage as something other than a measurement.
This is the same structural problem behind significance testing on a synthetic panel. When the constraint that used to discipline a number disappears, the number keeps its old appearance and loses its old meaning. Free sample size does that to a p-value. A specified panel does it to a base rate.
So the useful question is not whether the panel will produce a size. It will. The question is which links in the chain it can carry, and that turns out to be one of them.
What is a market sizing number actually made of?
Three kinds of input drawn from three different sources: a population base from official statistics, an incidence or penetration rate from fieldwork or commercial measurement, and a value per buyer from pricing and transaction data. Most sizing arguments fail on the middle term, because incidence is the only link teams habitually estimate from a survey.
Write the chain out before you argue about the tool. A market size is a population base multiplied by a qualifying rate, multiplied by a purchase rate, multiplied by a value, and then multiplied by whatever share you think you can capture.
Each multiplication has a home. The base sits in national statistics: the American Community Survey publishes US household and demographic counts¹, and Eurostat and the national registers do the same job across the EU². Those figures are free, current and auditable.
The middle terms are the expensive ones. A qualifying rate comes from a screener fielded to real people, and the ESOMAR and GRBN guideline on online sample quality sets out the reporting buyers are entitled to expect from that sample⁸. Penetration and frequency come from retail measurement or a consumer panel. Value per buyer comes from transaction records.
Notice what that means for the tool choice. Four of the five links are already outside the survey. Only the qualifying rate is a survey question, and it is the one link a synthetic panel cannot supply.
Why does a synthetic panel hand back your own denominator?
Because composition is an input, not a finding. You draw the panel to match published census distributions on age, income, region and household type, so a share read back on any of those attributes recovers what you specified. The panel is not measuring the population. It is reproducing the file you calibrated it against.
Run it through once. You draw 500 German personas calibrated to the national distribution of household size. You ask what share live alone. The answer lands near the published figure, because you built it to.
That is a working calibration check and worth running. It is not evidence about Germany, and it cannot be. Two-layer calibration is what makes the loop tight, which is a feature for panel composition and a trap for sizing.
Now step one attribute off the census. Ask what share of those households bought a premium dishwasher tablet in the last three months. No statistical office publishes that rate. The panel has no specification to return, so the model supplies a figure from what its training data makes plausible.
The evidence on that second case is not encouraging. Writing in Political Analysis, James Bisbee and colleagues found that estimates generated by language models diverge from human survey data unevenly across subgroups rather than shifting by a constant amount⁵. A 2026 arXiv paper on LLM consumer panels frames the same problem as a diagnostics question: a panel may be usable, but only once you have measured the gap against something real⁷. Separate work on uncertainty quantification for LLM-based survey simulations makes the matching point about inference, that a simulated share needs an interval built for simulation rather than the interval a survey would carry⁶.
A sizing model has nowhere to put that uncertainty. It multiplies a point estimate by four other numbers and prints a currency figure.
Which sizing inputs can a synthetic panel produce?
One of the five, and it is the link you already had. A census-grounded panel reproduces the population base because it was built from it. Qualifying incidence, penetration, frequency and value per buyer all need a source outside the panel. Treating any of them as a synthetic read puts a language prior inside a board number.
Two rows in the table below are worth separating from the rest, because they are where synthetic work earns its place in a sizing project.
Relative ordering between segments you have defined yourself can survive when levels do not. If three concepts are put to the same panel and one is consistently ranked ahead of the other two across independent replicate runs, that ordering is a finding about the panel that may hold in the market. It is still not a share, and the same discipline applies as in low-incidence work: the cell has to be one the census documents before the ranking means anything.
Scenario ranges are the second. Once the incidence rate comes from a real screener, the arithmetic around it is yours. A synthetic panel can tell you which qualitative conditions plausibly move a parameter and in which direction, which sharpens the range you model. It cannot tell you the parameter.
Everything else in the table belongs to fieldwork or to commercial data. NIQ, whose retail measurement supplies the penetration and frequency rows above, publishes its own briefing on the rise of synthetic respondents in market research⁹.
What should you run on a synthetic panel before a sizing study?
Three jobs, all of them upstream of fieldwork. Pretest the screener that will set your incidence rate, settle the segment definitions the model is built on, and map the need states that decide what counts as in-market. Each one shortens the human study you still have to field and lowers the chance it misfires.
The screener pretest is the highest-value of the three. Your incidence rate is the most expensive line in the whole chain, and a screener that reads ambiguously inflates or deflates it. On live fieldwork you discover that after the invoice.
Run the same question on a synthetic panel first and read the rationales rather than the percentage. If personas who clearly belong in the category are talking themselves out of it, the wording is routing people wrongly and you have found it for free. PersonaHive covers nine countries with panels grounded in each country's own national census, so a multi-market screener can be pretested in the market and language it will be fielded in. That is the same mechanic as cross-cultural pretesting, applied to the one question that sets your denominator.
Segment definition is the second job. Sizing models break when two segments overlap or when a definition turns out to be unusable in the field. A panel session surfaces that quickly, because you can see whether personas assign themselves consistently.
Need-state mapping is the third. Before you can size a market you have to decide what counts as being in it, and that boundary is a judgement, not a measurement. Ask a panel when people reach for the category and when they reach past it, and the occasions that fall outside your definition show up plainly. Move the boundary on paper, before it is baked into a screener that costs money to field.
The smallest useful next step is one screener question. Draw a panel in a single market you are sizing, run your qualifying question unchanged, and read the written rationales. PersonaHive's free tier includes 250 credits, which covers that test.
When does the sizing number still have to come from people?
Whenever the number leaves the room. Anything landing in a board pack, a business case, an investor deck or a regulatory file needs a chain of evidence behind it, and a synthetic read has no fieldwork to point back to. Four cases make that rule concrete: a commitment, a new category, an evidential claim, and anything tracked over time.
Take them in order.
A number that becomes a commitment needs real respondents. Volume forecasts, revenue plans and investment cases all get audited later against what happened, and a synthetic incidence rate gives an auditor nothing to check. Both industry bodies have published on where that boundary sits: the Market Research Society's Delphi Group in its report on using synthetic respondents for market research³, and AAPOR in its task force report on responsible AI integration in survey research⁴.
A new category needs real respondents, because there is no published base rate for the model to have learned and no census cell to anchor the draw.
A regulated or evidential claim needs real respondents by definition. That is a harder line than a methodological preference, and it is covered separately in claim substantiation.
A number that will be tracked over time needs a consistent human instrument, because you cannot separate market movement from model movement without one.
Outside those four, synthetic work is doing something useful in a sizing project. It is just doing it before the study, not instead of it. Start with the screener.
What else do teams ask about market sizing on a synthetic panel?
The recurring questions are whether a larger panel fixes the base-rate problem, whether census grounding is enough on its own, whether ordering is safer than levels, whether TAM and SAM behave differently, and what belongs in the methods note. Short answers follow; the mechanism behind each is covered in the sections above.
**Does drawing a bigger synthetic panel make the incidence estimate more reliable?**
No. Sample size reduces sampling error, and sampling error is not the problem here. Ten thousand personas return the same base rate as five hundred, with a tighter interval around a number that was never a measurement.
**If the panel is census-grounded, does that not fix prevalence?**
Only for the attributes the census publishes, and for those it is circular. Census grounding fixes who is in the panel. It says nothing about whether the panel's answers about category behaviour match the real people in each cell.
**Is relative ordering between segments safe to use?**
More often than levels, but it needs evidence. Run the comparison across independent replicate runs and treat a ranking that flips between runs as absent. A stable ranking is a hypothesis for fieldwork, not a substitute for it.
**Does TAM behave differently from SAM and SOM?**
TAM is usually the safest of the three, because it leans hardest on the population base and least on behaviour. SAM and SOM both depend on qualifying and capture rates, which is where synthetic estimates are weakest.
**What belongs in the methods note if synthetic work touched the sizing project?**
State which links came from the panel and which from fieldwork or commercial data. In practice that means recording the panel only against screener wording, segment definitions and scenario structure, and naming the external source for every multiplied parameter.
Sources
- American Community Survey (ACS) — U.S. Census Bureau
- Database: Population and demography — Eurostat
- Using synthetic respondents for market research (MRS Delphi Report) — MRS Delphi Group, Market Research Society
- Responsible AI Integration in Survey Research (Task Force Report) — AAPOR
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models — James Bisbee and colleagues, Political Analysis
- Uncertainty Quantification for LLM-Based Survey Simulations — arXiv
- When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels — arXiv
- ESOMAR/GRBN Guideline on Online Sample Quality — ESOMAR and GRBN
- The rise of synthetic respondents in market research — NIQ