Which LLM Runs Your Synthetic Panel: What It Changes
Methodology · 11 min read
TL;DR: The question arrives in every evaluation: which model do you run. No model is validated as the best choice for simulating a population, and the published work gives no reason to expect the largest one to win. Scale has not closed the gap between model answer distributions and real demographic groups, and tuning with human feedback narrows coverage rather than widening it. What model identity does control is comparability. Named models change behaviour between releases, so an upgrade nobody commissioned can move a tracker reading while the market stands still. Treat the model as a logged study parameter, ask a vendor six specific questions about it, and keep a fixed battery you can re-run.
Which LLM is best for synthetic survey respondents?
No model holds that title. Published evaluations of population simulation test specific models under specific conditions, and none of them establishes a general winner. Model identity still matters, because it sets the answer distribution your panel starts from and the comparability of readings over time. Treat it as a parameter you record, not a quality tier you buy.
The question is usually asked as though the answer were a brand name. It is closer to a question about an instrument's calibration.
Two different properties get bundled into it. Capability is whether the model can read your questionnaire, follow a scale and write a coherent rationale. Every current frontier model can.
Fit is whether the answers it produces, drawn across a population you specified, land where that population lands. Your study depends on the second one, and it is not a function of how well the model writes.
The founding paper in this area is explicit about the scope of its own result. The 2023 Political Analysis study that introduced algorithmic fidelity demonstrated a conditioned model reproducing patterns in US survey data, and framed fidelity as a property of a particular model under particular conditioning rather than a general capability of language models.⁷
That framing has not changed. Fidelity is measured, per model, per population, per instrument. Nobody has shown it transfers when you swap the engine.
The buyer guides that rank synthetic research platforms almost never report which model sits underneath, and the vendors rarely volunteer it. So the practical question is not which model is best. It is what the model you are running does to your numbers, and what happens the day it changes.
Does a bigger model make synthetic research more accurate?
No published evidence says it does. The benchmark work on population simulation reports that moving to a larger general model does not bring its answer distributions closer to real demographic groups. Fit improves when a model is adapted to survey response data, which is a different purchase from a frontier upgrade.
This is where most evaluations go wrong, because capability benchmarks and population benchmarks point in different directions.
The Stanford work behind OpinionQA measured model answer distributions against the Pew American Trends Panel and reported substantial divergence from the views of US demographic groups. Models tuned with human feedback skewed toward a narrower set of groups rather than toward the population average.¹
A 2024 NeurIPS paper tested models against the US American Community Survey and found a systematic pull toward the first answer option offered, a bias that moving to a larger model did not resolve.² The defect lives in how the model reads a question, and a better writer reads it the same way.
What does move fit is specialisation. A 2025 NAACL study adapted models on survey response data and improved their match to response distributions across global populations.³ Qualtrics makes a vendor claim along the same line, reporting that its purpose-built survey model outperforms general-use models on survey research tasks.⁸ The gain comes from fitting a model to the task, not from buying a larger one.
The table below separates what a frontier upgrade reliably gives you from what it does not.
Whose opinions does the model bring to your panel?
Its own. A model carries an answer distribution learned from text and then reshaped by tuning, and that distribution is not the population you are studying. Census grounding fixes who is in the sample. It does not replace what the model already believes about the questions you are putting to it.
A synthetic run is two layers stacked. Underneath sits a general model with a distribution of its own. On top sits whatever the platform does to aim that distribution at a specified population.
The OpinionQA result reads the bottom layer. Measured against a national probability panel, the models' views sat closer to some demographic groups than others, and human-feedback tuning sharpened rather than softened that tilt.¹ A newer model from the same provider changes the tilt without announcing it.
The top layer is what a research platform is for. PersonaHive draws each panel against its own national census, country by country across nine markets, and adds four design layers on top of that grounding to counter the flat, middle-of-scale answering that untreated models default to. Every response carries a required written rationale, so a tilt arrives as text you can read rather than an average you cannot interrogate.
None of that erases the bottom layer. Grounding sets the composition of the sample; the model still supplies the opinion inside each cell. That is also why it already holds a view of your brand before the study starts.
The honest statement about model choice is therefore narrow. A different model gives you a different starting tilt, and no available evidence ranks those tilts for your category.
What happens when the vendor upgrades the model mid-study?
Your series loses its baseline. Named models have been shown to change behaviour between releases on identical tasks, and synthetic survey estimates shift with small changes to prompt or version. A wave-on-wave difference then contains a model change nobody commissioned, mixed into whatever the market did.
This is the part of model choice that costs real money, and it is rarely in the contract.
The Harvard Data Science Review study on behaviour change over time ran identical tasks against the same named ChatGPT versions months apart and documented substantial differences in how they responded.⁶ The label on the model stayed the same. The behaviour did not.
The survey-specific evidence shows the same fragility. The Political Analysis paper on synthetic replacements for human survey data generated responses with ChatGPT and found means landing near the human benchmark while variance was compressed, with subgroup conclusions shifting under changes to the prompt and the model version.⁵ A TACL study tested whether models reproduce the response biases that survey methodology has documented in people, reported that they largely do not, and found human-feedback-tuned models more sensitive to question changes that leave people unmoved.⁴
Read those three together and the operational conclusion is blunt. The model is the most volatile input in your study, and it is the one input you do not control.
This is why a synthetic tracker needs more care than a human one. A wave-on-wave difference on a synthetic panel already mixes the market, the model and the instrument, and an unannounced version change removes your ability to separate them after the fact.
What should you ask a vendor about the model?
Six questions, all answerable in writing. Which model family and version runs a study, who decides when that changes, whether a prior version stays available for a re-run, what grounding sits on top of it, what validation exists against a named human survey, and what each run records about all of this.
Most evaluation checklists stop at the demo. These six decide whether your data is still readable in six months.
One. Which model family and version answers a question today. A vendor that will not say has handed you an unrecorded study parameter.
Two. Who decides when it changes, and do you get notice. The decision usually sits with the model provider, not the vendor.
Three. Can a prior version be re-run. This is the only clean way to separate a model shift from a market shift later.
Four. What grounding sits on top. PersonaHive builds personas from aggregated public statistics calibrated to each country's own national census, which makes the composition of a panel checkable rather than asserted.
Five. What validation exists, against what. A named, held-out human survey is the only answer worth scoring. PersonaHive publishes a held-out benchmark against the 2024 US CFPB National Age-Friendly Banking Survey, a specific comparison rather than a general accuracy claim.
Six. What does a run record. Model identity, grounding version, instrument and seed belong in the output file, not in a support ticket. The six fields worth recording on every run set out the minimum.
This is the methodology column of your diligence, not a replacement for it. The fifty-question RFP framework covers security, pricing and support.
The smallest useful step is cheap. Write a ten-question battery you would repeat unchanged for a year, run it once on a single market, and file the output. Signing up gives free credits and asks for no card, so the baseline costs an afternoon.
How do you test whether a model change moved your numbers?
Re-run a fixed instrument. Keep a short battery whose wording never changes, run it at every wave, and compare it against its own history. If the fixed battery moves while the market story stands still, the instrument moved. Without that held baseline, the two are not separable after the fact.
Six steps, and the first has to happen before you need it.
One. Build the fixed battery now. Ten to fifteen items, neutral wording, no reference to a current campaign or price. Its only job is to stay the same.
Two. Run it at every wave, in the same session as the live questionnaire. A battery run separately picks up a different context.
Three. Record what produced each run: model family and version, grounding version, date and run settings, stored with the output.
Four. Replicate within the wave, so you know how much movement the method produces when nothing has changed.
Five. Compare against the stored history, not against expectation. A shift larger than your within-wave spread is a signal about the instrument. A shift inside it is noise.
Six. On a confirmed shift, re-run the previous version if the vendor still offers it. If it does not, the honest report says the series restarts here and states why.
The limit matters. This procedure detects that something moved. It does not tell you which direction is correct, because neither run is a measurement of people. Deciding which reading to trust still needs a human benchmark, and a model change is a good moment to field one.
What else do teams ask about model choice?
Five questions recur once the version problem is clear: whether running several models helps, whether an open-weight model fixes the drift, whether reasoning modes improve fit, whether model choice shows up in the methods note, and what to do with a study that has already shipped.
Does running the same study on two models and averaging help? It gives you a sensitivity check, which is useful, and a combined number, which is not. Report the spread between models as a range and leave the average out, for the same reason a p-value does not apply when you set the sample size.
Does an open-weight model you host yourself solve the drift problem? For comparability, yes, because you decide when the weights change. It does not improve fit to a population, and the benchmark work gives no reason to expect a smaller self-hosted model to do better there.²
Do reasoning or extended-thinking modes produce better respondents? They produce longer rationales. No published population benchmark shows deliberation length improving distributional fit, and an elaborate justification for the same answer is easy to mistake for evidence.
Should model identity appear in the methods note? Yes, named and versioned, next to the grounding source and the field dates. A reviewer who finds it missing is right to discount the result.
A study already shipped without any of this recorded. What now? Do not retract it. Re-run the instrument today, report both readings with their dates, and treat the gap as the cost of the missing parameter. Then write the fixed battery, run it once on one market, and store what produced it.
Sources
- Whose Opinions Do Language Models Reflect? — Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang and Tatsunori Hashimoto, Proceedings of Machine Learning Research (ICML 2023)
- Questioning the Survey Responses of Large Language Models — Ricardo Dominguez-Olmedo, Moritz Hardt and Celestine Mendler-Dunner, Advances in Neural Information Processing Systems (NeurIPS 2024)
- Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations — NAACL 2025, ACL Anthology
- Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design — Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkar and Graham Neubig, Transactions of the Association for Computational Linguistics
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models — James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel and Jennifer M. Larson, Political Analysis
- How Is ChatGPT's Behavior Changing over Time? — Lingjiao Chen, Matei Zaharia and James Zou, Harvard Data Science Review
- Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle and colleagues, Political Analysis
- Qualtrics' AI Outperforms General-use LLMs on Survey Research — Qualtrics