Synthetic Interviews: What Probing Actually Adds
Methodology · 11 min read
TL;DR: A synthetic interview reads like a good one. The persona answers, you probe, it elaborates, and the transcript fills with quotable material. That elaboration comes from the same representation that produced the first answer, so the second answer is longer rather than better informed. A human depth interview earns its cost because the respondent holds a memory you cannot reach any other way. A persona holds no memory, so probing recovers nothing. Two different things are sold under one label: an AI interviewer putting questions to real people, which the published work supports, and a simulated respondent answering them, which is a different instrument. Keep the depth work with people and send the structured battery to the panel.
Do synthetic interviews work?
Not as interviews. A depth interview is an instrument for retrieving something only the respondent holds: an episode, a reason, a memory of deciding. A census-grounded persona holds a demographic profile and whatever the underlying model has read, so a follow-up question returns more text from the same source rather than new evidence.
Two things are sold under the same word, and the difference decides whether the output is evidence.
An AI interview puts a language model in the interviewer's chair. A real person answers. The model asks the follow-ups and the respondent supplies the content.
A synthetic interview puts the model in both chairs. The questions come from your guide and the answers are generated. Nobody is on the other side of the table.
The most ambitious published work in this area makes the dependency hard to miss. A Stanford and Google DeepMind team built generative agents of 1,052 people and reported those agents reproducing their originals' survey answers at 85 percent of the participants' own test-retest consistency.¹ The input for each agent was a two-hour interview with the person it modelled.
That result is the argument in one line. The fidelity came out of the interviews. Interviews are how information gets into the system, not something the system can produce on its own.
What does probing actually retrieve from a synthetic respondent?
Nothing that was unavailable when the first answer was written. Probing works on people because recall is partial and a second question reaches material the first one missed. A model composes each answer in full from a single representation, so a follow-up returns more sentences drawn from the same place.
Consider what a probe does in a human interview. You ask why someone switched banks. They give the tidy reason. You ask what happened in the week they decided, and a different account comes out, because the tidy reason was a summary and the episode was still sitting there to be recovered.
The persona has no week. It has a profile and a model behind it. Ask the follow-up and the system runs again, writing a longer answer conditioned on the one before it. Nothing is recovered because nothing was withheld.
Measured work on generation diversity points the same direction. NoveltyBench evaluates whether repeated generations from a model are genuinely distinct rather than reworded, and reports models returning markedly fewer distinct outputs than the number of samples drawn.² A 2026 analysis of the same effect traces the bottleneck to calibration and finds that drawing more samples yields diminishing new content rather than proportionally more.³
Apply that to a transcript. Six follow-ups do not give you six layers. They give you one answer in six lengths. Adding personas does not help either, for the same reason it does not settle how many personas a study needs: more draws from one generator tighten the centre and leave the edges empty.
Why does a synthetic transcript still read so well?
Because fluency and informativeness are separate properties, and a model is optimised for the first. Each transcript is coherent, specific and quotable on its own terms. The weakness shows up only across the set, where themes, phrasings and emotional register repeat far more than they would across real respondents.
This is why the method survives internal review. Nobody reads twenty transcripts side by side. They read two, find them vivid, and sign off.
A controlled experiment published in Science Advances in 2024 found the same shape in a different task. Writers working with generative assistance produced individually more creative stories while the collective diversity of what the group produced fell.⁴ Richer per unit, narrower in aggregate.
A set of synthetic interviews behaves the same way. The transcript gains and the spread shrinks, and spread is the thing a qualitative study is sampling for.
Two consequences follow. Themes appear to saturate early, because the repetition arrives for generative reasons rather than because the population agrees. And a striking quote is no evidence that anyone holds the view, so a generated verbatim does not belong in a deck as a customer voice.
What is the difference between an AI interview and a synthetic interview?
One replaces the moderator, the other replaces the respondent. An AI interviewer questioning real people is a fieldwork method, and published evaluations of it are encouraging. A simulated respondent answering your guide is a modelling exercise. Both get marketed as AI interviews, and only one produces data about people.
The evidence splits cleanly along that line.
On the interviewer side, an LLM-in-the-loop study of AI-generated follow-up questions examines how models produce probes inside real interviews and what those probes ask of the person answering.⁵ A CESifo working paper on conducting qualitative interviews with AI compares AI-led interviews against human-led ones and examines the quality of what respondents disclose.⁶ A case study on collecting qualitative data at scale documents the same configuration working as a collection instrument.⁷
All three keep a person on the answering side. That is not incidental. It is the reason there is data at the end.
On the respondent side there is no equivalent literature, because there is no equivalent object to validate. You can test whether a simulated answer matches a known population statistic. You cannot test whether a simulated memory happened.
The table below is the version to put in front of whoever asks why the two quotes differ by an order of magnitude.
Which qualitative jobs still hold up on a synthetic panel?
The ones that need no memory. Mapping the span of reasons a category produces, pressure-testing a discussion guide before it reaches a paying respondent, building a codeframe, and comparing how different markets frame the same problem. Those are language and attitude tasks, and a census-grounded panel represents them.
Four jobs hold, and together they are worth more than a convincing fake transcript.
Hypothesis spread. Ask a panel what could make someone abandon a checkout and you get the space of candidate reasons, drawn wide across demographics. Treat it as a list to test, never as a count. Prevalence is a property of people.
Guide design. Most discussion guides are inherited and too long. A synthetic pass finds the double-barrelled question, the term nobody outside the company uses, and the warm-up that earns nothing. Fix those before a recruited respondent sits down.
Codeframe construction. Build the coding scheme on synthetic language, then apply it to human transcripts. The scheme is a hypothesis and the counts come from people. The rules for reading that text are set out in the guidance on what counts as a finding in synthetic open-ends.
Framing across markets. Each PersonaHive panel is drawn against its own national census, country by country across nine markets, so you can see how a problem is described in Hungary and in Finland before either market is fielded. The comparison is about language and attitude, which is the half that travels.
What does not hold is the group version of the same idea. Putting several personas in conversation restores none of the missing information and adds a second problem, because synthetic focus groups converge toward agreement.
How do you turn a discussion guide into something a panel can answer?
Split it by what each question requires. Anything asking the respondent to recall an event stays with people. Anything asking what the category expects, what words it uses, or how a type of person frames a decision becomes a structured item with a written rationale attached. Run that half, then field the shortened guide.
Six steps keep the boundary clean.
One. Mark every question containing the words your, last, remember or recent. Those are episode questions and they stay with recruited respondents.
Two. Rewrite the remainder as structured items rather than open conversation. A rating with a required justification gives you something inspectable. An open monologue gives you prose.
Three. Require a written rationale on every response. PersonaHive enforces this at platform level rather than through prompt wording, which is what makes the reasoning auditable rather than hidden behind an average. If a quarter of the rationales describe a retail channel that does not exist in that market, the run is reporting on the model.
Four. Run each item more than once and record the spread. One pass reads like a person. Several show how much of the answer came from your brief and how much from the draw.
Five. Code the output into a scheme, then cut the human guide using what you learned. The synthetic stage should pay for itself in fieldwork minutes removed.
Six. Label the outputs in the methods note. If a quote or a count in your deck is attributed to a customer, it came from a customer.
The smallest useful version of this takes an afternoon. Take the guide you were about to field, split it by that first rule, and run the attitude half on a single market. Signing up gives free credits and asks for no card, so the test costs time rather than budget.
What else do research teams ask about synthetic interviews?
Five questions recur once the split is clear: whether a richer persona brief fixes the problem, whether a synthetic interview can stand in for a pilot, whether B2B changes the answer, what to do with a transcript that already shipped, and whether a human still has to code the output.
Does a longer, richer persona brief fix it? No. A brief can state that a persona cancelled a subscription last spring. It cannot supply what happened that spring, so the model writes it at answer time. More specification makes the invention more confident, not more accurate.
Can a synthetic interview stand in for a pilot? For part of one. The instrument-testing half transfers. The half that checks whether real respondents tolerate the guide does not, because tolerance is a behaviour rather than an opinion.
Does B2B change the answer? The mechanism is identical and the grounding is thinner, since less is written about business buyers. Expect wider spread between runs and weaker footing on named vendors and narrow job titles.
What if a synthetic transcript already went into a report? Relabel rather than retract. Report it as generated material with the method attached, and move anything presented as a customer voice into the next human wave.
Does a human still have to code the output? Yes. Work examining language models as qualitative analysts finds model coding useful as a first pass and unreliable as the last one.⁸ The scheme stays yours to own.
Start with the guide you already have. Split it by episode and attitude, run the attitude half, and see how much shorter the human session gets.
Sources
- Generative Agent Simulations of 1,000 People — Joon Sung Park and colleagues, arXiv
- NoveltyBench: Evaluating Language Models for Humanlike Diversity — arXiv
- Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs — arXiv
- Generative AI enhances individual creativity but reduces the collective diversity of novel content — Anil R. Doshi and Oliver P. Hauser, Science Advances
- Ethics and Social Responsibility in AI-Assisted Interviewing: An LLM-in-the-Loop Study of AI-Generated Follow-Up Questions — arXiv
- Conducting Qualitative Interviews with AI — Felix Chopra and Ingar Haaland, CESifo Working Papers
- Collecting Qualitative Data at Scale with Large Language Models: A Case Study — arXiv
- Examining Large Language Models as Qualitative Data Analysts — arXiv