Synthetic NPS: Measuring an Experience Nobody Had
Methodology · 11 min read
TL;DR: A synthetic panel will return an NPS, a CSAT and a CES for any brand you name. None of the three is a satisfaction measure, because each one asks a respondent to grade something that happened to them, and nothing happened to a persona. What comes back is a reputation estimate wearing a customer experience label. NPS is hit hardest, because it nets two tail proportions and throws the middle away, so the response style every language model carries moves the headline by points rather than decimals. The expectation side of a satisfaction model does survive, since expectations are a population attitude rather than a private event. Run the synthetic pass before you field, and keep the experience items with people.
Can a synthetic panel measure NPS?
Not as a satisfaction measure. Net Promoter Score asks how likely you are to recommend a company you have dealt with, and the question depends on having dealt with it. A census-grounded persona holds no account, no purchase and no service history, so the eleven-point answer it returns reports reputation rather than experience.
The score itself is simple. Frederick Reichheld introduced it in Harvard Business Review in December 2003 as the share of respondents answering 9 or 10 minus the share answering 0 to 6, on a single likelihood-to-recommend question.¹ Everything else about it is argued over. A 2007 study in the Journal of Marketing re-ran the growth analysis across industries and did not find the predictive advantage over other satisfaction measures that the original claim rested on.²
That argument is about people. It assumes a respondent who bought something and formed a view. Put the same question to a synthetic panel and you are one step further back, because the respondent never bought anything.
The persona still answers. Models rarely decline. You get a clean number, a distribution across the eleven points, and a set of verbatims explaining the rating. All of it is produced from what the underlying model has read about the brand, conditioned on the demographic brief you supplied.
That is a usable quantity for some jobs. It is not customer satisfaction, and labelling it NPS in a deck invites a decision nobody would make if the label were accurate.
What exactly is missing when there is no episode?
The event the question refers to. Customer experience instruments are recall instruments: they ask someone to grade a purchase, a call, a delivery or a year of use. A persona has a biography, not a transaction log, so when a question names an episode the model supplies a plausible one and then grades its own invention.
Satisfaction measurement has a settled structure. The American Customer Satisfaction Index builds its score from expectations, perceived quality and perceived value, and it collects those from people with recent experience of the company being scored.³ The model only works because one side of it is a lived comparison.
A synthetic panel can hold the expectation side. It cannot hold the comparison, because there is nothing to compare against.
What happens instead is quiet and easy to miss. Asked why they rated a bank a 4, a persona writes a short account of a branch visit or a failed transfer. The account reads well. It was generated at the moment of answering, to justify a number the model had already landed on.
This is where a platform-level control earns its place. PersonaHive requires a written rationale on every response, which makes the invention inspectable rather than hidden behind an average. If a third of your detractor rationales describe service events in a market where you have no retail presence, the number is telling you about the model and not about the market.
The related question of whether a panel can represent your buyers at all is separate, and answered in what a census-grounded panel can and cannot say about your customers. This piece assumes you have solved that and asks what the instrument does next.
Why does NPS break worse than an average rating?
Because NPS is a difference between two tails. It counts 9s and 10s against 0s to 6s and discards the 7s and 8s entirely, so a response distribution that is slightly too kind or slightly too flat moves the headline by several points while the mean barely registers the change.
Work the arithmetic once and the sensitivity is obvious. Take a panel where 10 percent of responses sit at 8. Move that 10 percent to 9. The mean rises by 0.1. The NPS rises by 10 points, because a passive has become a promoter. Nothing about the underlying view changed by a tenth of a point.
Now consider what a language model does to that distribution. Two documented pressures push in opposite directions.
The first is agreement. Anthropic researchers reported in 2023 that models trained with human feedback exhibit sycophancy, tending to match what the asker appears to want.⁴ A brand-owner framing pulls ratings up.
The second is hedging. Models trained to avoid strong positions cluster toward the safe middle of a scale, the effect covered in neutrality bias in synthetic research. On an eleven-point scale that means 7s and 8s, which NPS throws away.
So a synthetic NPS is the net of two biases whose balance shifts with the model, the wording and the category. A mean can sometimes be corrected with a constant offset measured against a human benchmark. A net of two tails cannot, because the offset is not constant across the distribution.
The table below is the short version for the four instruments teams usually ask about.
Which customer experience questions still work on a synthetic panel?
The ones that name no episode. Category expectations, the language people use for service failure, what a segment believes about a brand it has never bought, and whether your questionnaire reads cleanly. Those are attitudes and wording, which a census-grounded panel represents. Recollections are not, and it does not.
Four jobs survive, and they are worth more than the broken score.
Expectations. What should a mid-market insurer answer a claim in, how long is too long for a grocery delivery window, what does a reasonable return policy look like. These are population attitudes held by people who have never bought from you, and they are the input half of a disconfirmation model.³ Run them country by country and the comparison holds, because each panel is drawn against its own national census.
Driver vocabulary. Ask a panel what makes a service interaction go badly in a category and you get the space of reasons people give. Use it to build your codeframe before fielding, never to estimate how common each reason is. Prevalence is a count of real people.
Instrument pretesting. A customer experience tracker is usually long and usually inherited. A synthetic pass finds double-barrelled items, untranslatable idioms and scale points nobody distinguishes, before the questionnaire reaches a paying customer.
Perception at a distance. Comparing how a category talks about two brands is legitimate, with one control. Run the stimulus unbranded first and branded second, for the reasons set out in blind versus branded stimulus testing. The gap between the two runs is the brand effect. The branded run on its own is a measure of how much has been written about the name.
What does the published evidence say about this class of question?
That broad attitudes replicate better than specifics. Since Argyle and colleagues showed language models reproducing population-level attitude distributions,⁵ later work has found the picture uneven: a 2026 psychometric audit reports synthetic responses that look plausible and fail validity checks,⁶ and a cross-domain benchmark documents where simulated respondents diverge from people.⁷
Three readings of the literature matter for a customer experience brief.
Aggregates travel further than individuals. The early simulation work established that a model conditioned on demographics can approximate how a population distributes across an attitude.⁵ Nothing in that result extends to a private event in one respondent's past.
Plausibility is not validity. The 2026 audit of language models as synthetic survey respondents makes the distinction its title: responses can read as plausible while failing the psychometric checks a researcher would run on any new instrument.⁶ Customer experience batteries are instruments, and they carry that exposure.
Performance varies by domain, so it has to be measured per study. The cross-domain benchmark published in 2026 found simulated survey responses failing unevenly across subject areas,⁷ a replication study by Verasight tested how well models reproduce omnibus survey data across topics,⁸ and a 2026 diagnostics paper proposes checks and corrections for consumer panels built this way.⁹ The common thread is that you establish fit for your questions, in your category, rather than inheriting a vendor's headline.
Read together, the evidence points the same direction as the mechanism. Stable, widely held, publicly discussed attitudes are the strong case. A rating of something that happened to one person last Tuesday is the weak case, and customer experience work lives there.
How do you run a customer experience study with a synthetic panel in it?
Put the synthetic pass before fielding, not in place of it. Pretest the instrument, build the codeframe, read the expectation items, then send the experience items to customers on a shorter questionnaire. The synthetic run earns its place by cutting what you ask people, not by replacing who you ask.
A six-step sequence keeps the boundary clean.
One. Split the questionnaire into episode items and attitude items. Anything containing the words your, last or recent is usually an episode item. Those stay with customers.
Two. Run the attitude items on a synthetic panel, country by country if you field in more than one market. Record the panel specification, the model, the run date and the number of replicate runs.
Three. Pretest the full instrument on the panel for wording defects. Log the defects, fix them, and do not report any of the pretest numbers.
Four. Build your open-end codeframe from the synthetic driver vocabulary, then apply it to human verbatims. The codeframe is a hypothesis list. The counts come from people.
Five. Field the shortened instrument to customers. The synthetic stage should have removed enough items to pay for itself in completion rate.
Six. Label the outputs honestly in the methods note. A synthetic expectation score is a synthetic expectation score. If a number in your deck is called NPS, it came from customers.
That last rule is the one that gets broken under deadline, and it is the one that costs you the room when a stakeholder finds out later.
What else do insight and CX teams ask?
Five questions recur once the boundary is clear: whether a larger panel helps, whether a tightly written persona can carry an experience, whether trending a synthetic score is safe, whether B2B changes the answer, and what to do with a tracker that already shipped a synthetic number. Short answers follow.
Does a bigger synthetic panel fix the score? No. More personas draw more responses from the same generator. The centre tightens and the missing episode stays missing. Sample size was never the constraint here.
Can a detailed persona brief carry an experience? A brief can state that a persona has been a customer for three years. It cannot supply what those three years contained, so the model fills the gap at answer time. Specifying the fact of a relationship makes the invention more confident, not more accurate.
Is it safe to trend a synthetic customer experience score over time? Only against a frozen control. A wave-on-wave difference contains the market, the model version and the instrument at once, which is the same problem a synthetic brand tracker has. Without a control wave you cannot say which one moved.
Does B2B change the answer? The mechanism is identical and the reputational prior is thinner, because less is written about most business brands. Expect wider spread between runs and weaker grounding for named vendors.
What if a synthetic NPS already shipped? Relabel it rather than retract it. Report it as a reputation index with the method attached, say what it does and does not cover, and put the experience items into the next human wave.
The practical next step is small. Take the customer experience questionnaire you are about to field, split it into episode items and attitude items, and run the attitude half on a synthetic panel before the human wave goes out. Either you cut questions from the instrument, which pays for the exercise, or you confirm the instrument is already tight, which is worth knowing before you spend on completes. PersonaHive gives new accounts 250 credits without a card, which covers that pass.
Sources
- The One Number You Need to Grow — Frederick F. Reichheld, Harvard Business Review
- A Longitudinal Examination of Net Promoter and Firm Revenue Growth — Timothy L. Keiningham, Bruce Cooil, Tor Wallin Andreassen and Lerzan Aksoy, Journal of Marketing
- The Science of Customer Satisfaction — American Customer Satisfaction Index
- Towards Understanding Sycophancy in Language Models — Mrinank Sharma and colleagues, arXiv
- Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle and colleagues, Political Analysis, Cambridge University Press
- Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents — arXiv
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses — arXiv
- Can Large Language Models Replicate Survey Data Across Topics? — Verasight
- When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels — arXiv