Synthetic Purchase Intent Is Not a Sales Forecast

Methodology · 11 min read

TL;DR: A synthetic purchase intent score is two forecasts stacked on top of each other, and most teams read it as one. The first is old and well measured: people overstate what they will buy when nothing is at stake, and the gap varies by category, horizon and wording. The second is newer: the respondent is a language model conditioned on a persona, so part of the answer reflects the elicitation, not a plan to buy. Stacking the two is why a synthetic top-two-box figure should never be read as a trial forecast. This guide separates the gaps, gives the run design that makes the score usable, and names the four cases that still need people.

Does a synthetic purchase intent score predict sales?

Not as a level. A synthetic purchase intent score is a comparative instrument. It ranks concepts against each other well enough to be useful, and it carries two stacked biases that push the absolute number in a predictable direction. Use it to choose between options, not to size a launch.

Two different errors sit inside one number, and only one of them is new.

The first error is old. People overstate what they will buy, and marketing science has spent four decades measuring by how much and under what conditions. It lives in human surveys and has nothing to do with AI.

The second error is new. A synthetic respondent is a language model conditioned on a persona, so the output is the text such a person would plausibly produce. That is a separate error with a separate direction and no fixed size.

Teams that treat a synthetic panel as a cheaper survey inherit both errors and correct for neither. The fix is not a better model. The fix is reading the number as what it is, which is an ordering.

What is the intention-behaviour gap, and why does it apply before AI enters?

The intention-behaviour gap is the distance between what someone says they will do and what they do. Sheeran and Webb reviewed the evidence in 2016 and found that intentions account for only part of later behaviour, with many people who form an intention never acting on it. Every survey-based intent figure carries it.

Forecasters mapped this gap long before language models existed, and Sheeran and Webb's 2016 review is the standing summary of how far intention falls short of action.²

Morwitz, Steckel and Gupta examined when stated intentions actually track later sales.¹ The relationship holds better for products people already know than for genuinely new ones, and it weakens as the distance between the question and the purchase grows. That is a conditional result, not a constant, which is the important part.

Economists measured the same thing from the money side. Murphy and colleagues, reviewing stated preference valuation studies, found that hypothetical statements systematically exceed what people pay when real money is at stake.³ List and Gallet had already shown that the size of the disparity depends on how the question was run.⁴

So the same person gives you a different number depending on category, horizon and elicitation. None of that is about AI. It is why a raw top-box score has never been a forecast in any panel, human or otherwise.

What second gap does a synthetic panel add on top?

A model gap. The synthetic respondent is not reporting an intention, it is generating the response a persona like that would plausibly give. Recent work shows the elicitation design moves that response substantially, which means part of any score reflects how you asked rather than what the persona would do.

The evidence here is young and moving, so treat any single figure as provisional.

Work published in late 2025 showed that changing how a Likert rating is elicited, by scoring semantic similarity against anchor statements instead of asking for a number, changes how closely model output tracks human purchase intent distributions.⁵ The elicitation method is a free parameter that moves the answer.

A 2026 paper on validating language model simulations as behavioural evidence makes the general version of the point.⁶ A simulation that reproduces a published finding is evidence about the simulation. Getting the same answer is not the same as getting it for the same reason.

Willingness to pay behaves no better. Researchers inferring willingness to pay from model choices in 2026 recovered the quantity from subjective choices, which tells you the quantity exists inside the model.⁷ It does not tell you it matches a market.

The practical consequence: the model gap is not an offset you can look up. It varies by category, by wording, and by model version, which is why the six fields of a reproducibility record belong in every run.

How do you convert a synthetic purchase intent score into a decision?

Do not convert it into a sales number. Convert it into a rank order, then a shortlist. Run every concept in one session on one panel with identical wording, compare scores only against each other, and carry forward the gap between concepts rather than the level of any single one.

Five steps. The first and the third are the ones teams skip.

1. Fix the instrument before you run anything. Same wording, same scale, same option order for every concept. The rules for writing survey questions for synthetic personas matter more here than in a human study, because a model is more sensitive to phrasing than a person is.

2. Run all concepts in one session on one panel. Scores from two sessions three weeks apart are not comparable, because the model underneath may have moved between them.

3. Include a control concept with a known outcome. Something you already launched, ideally alongside something that failed. The control is the only external anchor the run has.

4. Report differences, not levels. A concept scoring 12 points above your control is a finding. A concept scoring 61 percent top-two-box is a number without a unit.

5. Take the top two or three to human respondents before anything expensive happens.

Step 3 is what separates a research programme from a tool. After four or five studies with a control in each, you hold a rough offset for your own category, which is the only calibration constant worth having. Nobody can sell you that number, because it is a property of your category and your instrument.

The mechanics of running the test itself, the metric stack and the scoring frame, sit in automated concept testing with AI personas. This guide is about what the resulting number means.

Where PersonaHive fits is steps 1 to 3. Personas are grounded in national census data on a country-specific basis, built from aggregated public statistics and validated against real surveys, across nine national panels covering the United States, Germany, France, Austria, Czech Republic, Hungary, Romania, Denmark and Finland. Running a full concept set against the same census cells in one session is what makes step 2 possible. Every response also ships with a written rationale, so a high intent score arrives with its reason attached, and you can see whether a persona is responding to the proposition or to one word in the copy. That is the honest differentiator against a prompt-built persona set. It is not a claim that the level is right.

Smallest useful next step: take one concept you launched and one that failed, run both unchanged on the matching census-grounded panel, and check whether the gap between them points the right way. A free PersonaHive account includes 250 credits, which covers it.

What does a defensible synthetic intent report look like?

It reports a rank order, a difference against a control, and the conditions the run happened under. It never presents a single top-box percentage as a standalone finding. A reader should be able to see the panel definition, the model version, the elicitation wording, and what the control scored.

The reporting format does most of the governance work here, because a number stated without its conditions will be quoted without them later.

One rule covers the rest. If a figure would survive being pasted into a slide with no methods note attached, it is a figure you should not have written down.

When should the intent number come from people instead?

Four cases. When the number feeds a financial model. When the category is genuinely new to the market. When the decision is expensive or hard to reverse. And when an outside party, a regulator, a retailer or a board, will be asked to rely on it. In each, the level matters and the ordering does not.

A financial model needs a level. A volume forecast built on a synthetic top-box figure inherits both gaps and then hides them inside a spreadsheet where nobody sees them again.

A new category removes the control. Step 3 depends on having a launched concept to anchor against, and a first-in-market proposition has none. Morwitz and colleagues found the intent to sales relationship is weakest for exactly this case, before any model was involved.¹

An expensive decision does not change the accuracy of the instrument. It changes the cost of being wrong.

External reliance is a governance question rather than a measurement one. The wider set of situations where synthetic evidence does not carry is in the six failure modes where synthetic research is not valid.

Start with the control concept. It is the cheapest change in this guide and the one that makes every run after it readable.

Frequently asked questions about synthetic purchase intent

**Is there a published multiplier I can apply to synthetic intent scores?**

No. The stated preference literature shows the disparity between hypothetical and real varies with elicitation and commodity type,⁴ and the model gap adds a second source of variation. Any usable factor is one you estimate from your own controls.

**Does forcing the persona to justify its answer make the score more accurate?**

It makes the score auditable, not accurate. A rationale tells you what the persona responded to, which is how you catch a score driven by a word in the copy. It does not move the level closer to a market.

**How many personas does a concept test need?**

Fewer than a segmentation and more than a single read. Because you are comparing concepts rather than estimating a population figure, stability of the ordering is the stopping rule. Track whether the ordering stops changing as you add personas, and stop there.

**Is a synthetic intent score worse than a human survey intent score?**

It carries one more gap. Both need calibration before anyone treats them as a forecast, and a human panel with poor sample quality can be the weaker of the two.

**Can I use synthetic willingness to pay instead, since it names a price?**

It is not safer. Willingness to pay is a stated preference measure with a longer record of hypothetical bias,³ and 2026 work recovering it from model choices demonstrates the quantity exists in the model rather than that it matches a market.⁷ Treat it the same way: comparative, not absolute.

Sources

  • When do purchase intentions predict sales? — Vicki G. Morwitz, Joel H. Steckel, Alok Gupta, International Journal of Forecasting, vol. 23, no. 3, pp. 347-364 (2007), Elsevier
  • The Intention-Behavior Gap — Paschal Sheeran, Thomas L. Webb, Social and Personality Psychology Compass, vol. 10, no. 9, pp. 503-518 (2016), Wiley
  • A Meta-analysis of Hypothetical Bias in Stated Preference Valuation — James J. Murphy, P. Geoffrey Allen, Thomas H. Stevens, Darryl Weatherhead, Environmental and Resource Economics, vol. 30, no. 3, pp. 313-325 (2005), Springer
  • What Experimental Protocol Influence Disparities Between Actual and Hypothetical Stated Values? — John A. List, Craig A. Gallet, Environmental and Resource Economics, vol. 20, no. 3, pp. 241-254 (2001), Springer
  • LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings — arXiv preprint 2510.08338
  • This human study did not involve human subjects: Validating LLM simulations as behavioral evidence — arXiv preprint 2602.15785
  • Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices — arXiv preprint 2602.09802

Related Articles

  • Automated Concept Testing: How to Validate Product Concepts in Hours, Not Weeks — A guide to automated concept testing with AI personas: the 5-step workflow, scoring metrics, comparison to traditional tests, and when to validate live.
  • Price Elasticity Surveys in FMCG: How AI and Synthetic Research Are Changing the Game — How FMCG brands use surveys to derive price elasticity of demand, and how AI respondents and synthetic research accelerate and improve pricing decisions.
  • How to Run a Validation Study for AI Synthetic Consumer Research — A practical methodology for validating a synthetic consumer research panel against a live national survey: what to measure, how to design a fair benchmark, and how to present the evidence to skeptical stakeholders.