Synthetic Open-Ends: What Counts as a Finding

Methodology · 11 min read

TL;DR: A synthetic open-end is generated text, not a report of experience. That one fact decides what you may take from it. Counting how often a theme appears measures the model's generation distribution, and published work on AI-written survey text finds it homogenises toward a common centre, so frequency understates the tails. A stated reason is not a cause: chain-of-thought explanations can be unfaithful to what actually drove the answer. What survives is coverage. Synthetic open-ends map the objection space, surface the vocabulary a category uses, and build the codebook you take into human fieldwork. Read them as a hypothesis inventory, never as consumer voice, and never quote one in a deck.

What can you conclude from a synthetic open-end?

A synthetic open-end tells you which arguments a census-grounded persona finds available and plausible for its profile. It does not tell you how many people hold that view, how strongly they hold it, or what they would say unprompted. Treat it as a map of the argument space, not as a measurement of a population.

Human open-ends are reports. A person read your question, retrieved something from their own life, and wrote a compressed version of it. The text is evidence about a mind that existed before the question arrived.

A synthetic open-end has no such prior. The text is produced at the moment of asking, conditioned on a persona description and the wording in front of it. Argyle and colleagues named the useful version of this in Political Analysis: conditioning a model on demographic detail produces response patterns that correlate with the matching human subpopulation⁵. Correlation at the level of patterns is a real finding. It is not testimony.

The practical consequence is a rule about verbs. You can say a synthetic panel surfaced an objection. You cannot say consumers raised it. The first claim is about coverage of an argument space and your data supports it. The second is about incidence in a population and your data cannot carry it.

Everything below follows from that one distinction.

Why does theme frequency mislead on a synthetic panel?

Because the count measures the model, not the market. Text generated by a language model clusters toward a common centre, so the same two or three framings recur across personas while genuine minority framings thin out. The rank order of themes is directionally useful. The percentages attached to them are an artefact of decoding, not an estimate of prevalence.

Zhang, Xu and Alvero studied what happens when survey open-ends are written with generative AI rather than by the respondent alone, and published the result in Sociological Methods and Research in 2025. The AI-assisted answers were more similar to one another than the unassisted ones, and the distinctive markers that separate one writer from another faded¹.

Padmakumar and He found the same shape in a different setting. Writers who composed alongside a language model produced text that was less diverse across writers than text written alone, presented at ICLR in 2024². Both studies describe a human writer using a model. A synthetic panel removes the human from the loop entirely, so there is no reason to expect a smaller effect.

Bisbee and colleagues reported the numeric version in Political Analysis in 2024: simulated survey responses carried markedly less variance than the human data they were meant to stand in for, and moved when the prompt was reworded³. Compressed variance in ratings and homogenised text are one phenomenon read on two instruments.

The loss lands in the tail. A theme mentioned by three personas out of two hundred is exactly the theme a real panel of two hundred might have raised twenty times, or never once. Your data cannot separate those two cases.

Can you trust a persona's stated reason for its answer?

Treat the reason as a hypothesis about the answer, not as its cause. Turpin and colleagues showed at NeurIPS in 2023 that models write explanations which omit the features actually driving their output. A rationale is still worth requiring, because it is inspectable and it changes the numeric answer. It is not a causal account.

The finding is specific. When researchers inserted a biasing feature into the prompt, models shifted their answers toward that bias and then wrote explanations that never mentioned it⁴. The explanation was fluent, internally consistent, and wrong about its own origin.

This bites hardest where teams reach for open text first: diagnosis. A concept scores 2.8 out of 5 and the rationales say the price feels high. The temptation is to treat that as the reason for the score and go and reprice. What you have is a plausible story produced alongside the score, not the mechanism behind it.

The test that separates the two is cheap. Change the price in the stimulus, hold everything else constant, and rerun. If the score moves, price is a driver. If the score holds while the rationales still talk about price, the rationale was narration. PersonaHive ships a written rationale attached to every response, which makes that check a matter of reading rather than instrumentation.

Can you quote a synthetic verbatim in a deck?

No. A synthetic verbatim on a slide reads as a consumer voice, and it is not one. Use it only inside a clearly labelled methods appendix, or paraphrase it as an objection the panel surfaced. Disclosure of how the data was produced is a standing question for buyers, not a footnote you add if challenged.

The failure mode here is drift, not deception. Someone lifts a well-written line into a summary slide, the label falls off in the third revision of the deck, and by the time it reaches a steering group it is being read as what customers said. Nobody decided to mislead. The format did the work.

ESOMAR publishes buyer guidance on augmented synthetic data that puts how the data was generated, and how that is disclosed, among the questions a buyer should ask before commissioning⁶. Answering it in the appendix of the deck is late. Answering it on the slide is not.

Two conventions prevent the drift. First, synthetic text never appears in quotation marks with a persona name attached. Write that the panel surfaced an objection, then describe the objection in your own words. Second, every artefact that leaves the research team carries a method line: which panel, which country, which run, generated on which date.

Verian published an industry study on synthetic sample in social research reporting significant limitations in AI generated responses⁷. Whatever your own read on the strength of the method, that debate is live, and a reader who discovers the provenance late will discount everything else in the deck. Labelling early costs less than defending late.

The same discipline applies to group formats. Synthetic focus groups converge, so a transcript that reads like a lively discussion may be recording agreement rather than opinion.

How do you code synthetic open-ends so the output holds?

Code presence, not frequency, across at least three replicate runs, and keep only the themes that appear in every run. Build the codebook from the first run, freeze it, then apply it blind to the rest. Report the theme list with run conditions attached, and report the range across runs rather than a single percentage.

Step one. Run three to five replicates of the same question on the same panel, changing only the seed. Replication is cheap on a synthetic panel and it is the only way to see which themes are stable.

Step two. Build the codebook from run one alone. Build it from all runs at once and you will fold rare themes into common ones without noticing you did it.

Step three. Apply the frozen codebook to the remaining runs without knowing which run you are coding. This is the step teams skip, and it is the step that keeps the codebook from growing to fit the data.

Step four. Keep themes that appear in every run. Mark themes appearing in some runs as unstable and carry them forward as questions rather than findings.

Step five. Record the run conditions beside the output: panel, country, model version, date, question wording, replicate count.

Step six. Report presence and stability. A line reading five of five runs is honest. A line reading 34 percent of respondents is not, for the same reason a synthetic panel reports a replicate range instead of a p-value.

The protocol costs an afternoon and it is what makes the output survive a methods review by someone who came in sceptical.

What do synthetic open-ends give you that a rating scale cannot?

An objection inventory before you spend on fieldwork. Coverage is the thing a synthetic panel is genuinely good at: it reaches framings your team has not thought of, across every market you plan to enter, at a cost that makes running it in nine countries reasonable. You then take that inventory into human research as the thing to measure.

A concept screen usually dies of an objection nobody wrote a question about. Scales cannot find it, because a scale measures only what the questionnaire already anticipated. Open text is where the unanticipated arrives.

That is a coverage job, and coverage is where compressed variance stops hurting you. You are not estimating how many people hold an objection. You are asking whether the objection is reachable at all from a profile grounded in national census data. A panel that overproduces the centre still reaches the centre reliably.

The handoff is the point. What you produce is a codebook and an objection list, not a result. Both go into human fieldwork, where the frequency question gets answered by people. Teams that skip this step write the human questionnaire from their own assumptions instead, which is the expensive version of the same mistake.

PersonaHive runs this on census-grounded panels in nine countries, each calibrated to its own national statistics, and the free tier includes 250 credits with no card. The smallest useful test: take one concept you have already fielded with humans, run the open-end on a matched panel, and check whether the objections your human respondents raised appear in the synthetic list. That single comparison tells you more about whether the method fits your category than any vendor claim will.

What else do teams ask about synthetic open-ends?

Five questions come up in nearly every methods review: whether a bigger panel fixes homogenisation, whether temperature restores diversity, whether a human coder is still needed, whether open-ends can stand in for depth interviews, and what belongs in the methods note. Short answers follow, with the reasoning in the sections above.

Does a larger panel fix the homogenisation problem? No. More personas give you more draws from the same generator, so the centre gets more precise and the tails stay empty. The fix is replication and honest reporting of presence, not sample size.

Does raising temperature restore diversity? It adds variation in wording more than variation in position, and it degrades instruction following at the same time. Prefer diversity built into panel composition over diversity bought from the sampler.

Do you still need a human coder? Yes, for the codebook and for the decisions about what counts as a distinct theme. A model can apply a frozen codebook at scale, and should, because that part is mechanical and benefits from consistency.

Can synthetic open-ends stand in for depth interviews? No. An interview earns its value from the follow-up question nobody planned, asked of someone with a life the interviewer cannot predict. A synthetic panel has no such life behind it, which is one of the six failure modes worth knowing before you design a study.

What belongs in the methods note? Panel and country, model version, run date, question wording, replicate count, and one sentence stating that responses were generated by census-grounded synthetic personas and are not human responses. That sentence is the one that survives the deck being forwarded.

If you want to test the method on your own category, run one open-end question on a matched panel and compare it against a study you have already fielded. It takes an afternoon and it settles the question for your team.

Sources

  • Generative AI Meets Open-Ended Survey Responses: Research Participant Use of AI and Homogenization — Simone Zhang, Janet Xu, AJ Alvero, Sociological Methods and Research
  • Does Writing with Language Models Reduce Content Diversity? — Vishakh Padmakumar, He He, ICLR 2024
  • Synthetic Replacements for Human Survey Data? The Perils of Large Language Models — James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, Jennifer M. Larson, Political Analysis, Cambridge University Press
  • Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting — Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman, NeurIPS 2023
  • Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, David Wingate, Political Analysis, Cambridge University Press
  • 5 Topics of Discussion to Help Buyers of Augmented Synthetic Data — ESOMAR, ESOMAR
  • Synthetic Sample in Social Research: significant limitations of AI generated responses — Verian, Verian Group

Related Articles

  • Forced Rationale: Why Every Synthetic Response Should Ship With a Written Justification — A rating without a rationale is a black-box output. Forcing every synthetic persona to write a justification grounded in its own backstory changes both the response and the auditability. Here is how the mechanism works and why it is not optional.
  • Synthetic Focus Groups Converge: Six Design Rules — Synthetic focus groups converge because AI personas conform to each other. The mechanism behind it, six design rules, and when to use a real group.
  • Synthetic Ad Testing: What a Persona Cannot See — An ad test measures two layers, and a synthetic panel reaches only one. The reception and response test for what persona creative feedback supports.