Van Westendorp on Synthetic Panels: What the Curves Mean
Pricing · 10 min read
TL;DR: Pricing research on a synthetic panel returns clean curves and a specific number, and the number is mostly a reading of your prompt. Van Westendorp asks for four absolute prices and supplies no reference point, so a model brings whatever your framing made salient. Gabor-Granger asks for a yes at ascending prices, and a monotone acceptance curve is close to guaranteed before the study runs. This guide separates what a synthetic price study can carry, which is rank order, direction and language, from what it cannot, which is the price you charge. It gives a two-anchor design that makes the anchor visible, and the cases that stay with people.
Can you run a Van Westendorp study on a synthetic panel?
You can run it, and it will return four clean curves, but the price points come mostly from your prompt. The price sensitivity meter asks for four absolute numbers and supplies no reference point, so the respondent has to bring one. A person brings a memory of the category. A model brings whatever your framing made salient.
Peter van Westendorp presented the price sensitivity meter to ESOMAR in 1976.¹ It asks for the price at which a product is too expensive, so cheap you doubt the quality, expensive but still worth considering, and a bargain. Plot the four cumulative curves, read the crossings, and you have an acceptable range and an optimal point.
The instrument has lasted fifty years because it is quick and needs no experimental design, which is also why it travels badly. Nothing inside the four questions constrains the answer. The constraint sits outside, in the respondent, who has paid for this category before and remembers what the shelf looked like.
A model has no receipt. It has a distribution over plausible prices for the product you described, conditioned on the words you used to describe it. Experimental work documents anchoring bias in large language models directly,² and a second study reports both behavioural and attributional evidence of the same effect.³ Numbers present in the context pull the numbers that come out.
So the question is not whether a synthetic price sensitivity meter produces an answer. It will produce four, and they will plot. The question is what moves them, and the honest answer is your brief.
Why does the too-cheap question fail first?
Because it rests on a price and quality inference a person holds and a model only describes. Gabor and Granger reported in 1966 that some buyers read a low price as a signal of poor quality, and that the effect varies sharply by buyer and by category. A model reproduces the textbook version of that rule, not one shopper's version of it.
The finding is old and it is specific.⁴ Price acts as a quality cue for some people, in some categories, at some points in the range. It is a dispersion result, not an average one, and the dispersion is the part a synthetic panel loses.
The too-cheap curve sets the lower bound of the acceptable range and helps fix the optimal point, so an error there moves the headline number the study exists to produce.
A 2026 benchmark measures the decision fidelity of language model user simulators against real purchase outcomes and states the pattern in its title: simulated customers never walk away.⁵ A respondent that rarely rejects will rarely name a price low enough to be suspicious either.
Run the four questions and you will still get four curves. They will look like every Van Westendorp chart you have seen. That resemblance is the trap.
Does Gabor-Granger hold up better than Van Westendorp?
It fails differently. Gabor-Granger reads a yes or no at ascending prices, and a model will nearly always accept at the bottom of the ladder and refuse at the top, so the shape of the curve is settled before the study runs. Only the inflection carries information, and the inflection moves with the ladder you chose.
Monotonicity is the whole output of a Gabor-Granger study, and monotonicity is the one thing a language model will reliably give you. Ask about a higher price and the probability of a yes goes down. That is a property of the instruction, not a measurement of a market.
The ladder is also the anchor. Present prices from 3 to 9 and the inflection lands inside 3 to 9. Present 8 to 20 and it lands inside 8 to 20. Neither run can say which is closer to a market, and the acceptance rate at a price is not a share of market. Whether willingness to pay can be inferred from a model's subjective choices at all is still an open research question.⁶
One layer sits underneath both instruments. Stated values exceed actual payments among human respondents, a gap established across the stated preference literature.⁷ A synthetic study stacks a second gap on the first, in the same direction. The related failure for intent scores is set out in why synthetic purchase intent is not a sales forecast.
What can a synthetic price study legitimately tell you?
Rank order, direction, and language. A census-grounded synthetic panel can tell you which of five price points a concept is ranked best at, whether adding a feature moves stated appeal up or down, and how people justify a price in their own words. It cannot tell you the price to charge.
The split is about which outputs survive when the absolute level is unreliable but the comparison runs under identical conditions. Hold the ladder, the persona specification and the wording constant across concepts, and the ordering does work the level cannot.
Language is the underrated output. The reasons personas give for calling a price high are hypotheses about what a real buyer will object to, collected while you are still writing the questionnaire. Treat them as candidate objections, never as frequencies.
Everything below the line in the table should leave the study as nothing at all. Not as a number with a caveat. A number with a caveat gets copied into a deck without the caveat.
How do you design a synthetic price study so the anchor is visible?
Run the same study twice at two different anchor ranges and report the gap between the two results. If the two runs disagree by more than the threshold you set beforehand, the study measured your framing rather than a market. That control costs one extra run and it is the cheapest way to see the anchor.
Five steps make a price study readable by someone who did not run it.
Fix the ladder or the reference set and write it into the run record. Run every concept against that same ladder. Then run the whole study a second time at a deliberately different range, and report both results side by side rather than averaging them.
Use five replicate runs per condition instead of one large run. The spread across replicates is the error bar a synthetic study actually has, and it costs less than scaling the sample, as set out in what free n hides on a synthetic panel. Hold the persona specification constant, and set your decision threshold before you look at any output.
PersonaHive runs pricing studies on personas grounded in national census data, country-specific and validated against real surveys. That grounding fixes the income and household distribution the panel is drawn from, which is not the same as fixing the price a persona names, a distinction worth checking in national census-grounded personas across countries. Where a price must be paid against something rather than named in the open, use a trade-off format: the platform supports choice-based conjoint, whose own limits are set out in what holds for conjoint and MaxDiff.
The smallest useful next step: take one price question you have already run with people, run it twice on a synthetic panel at two different anchor ranges, and compare the two synthetic results to each other before comparing either to the human result. The free tier is 250 credits and needs no card.
When does the price still have to come from people?
Whenever the number leaves the room. Any price you will publish, any figure entering a revenue forecast, any value-for-money claim, and any launch gate need human evidence. Recent work on synthetic consumer panels frames the question as one of diagnostics and corrections rather than trust by default.⁸
Four cases are not close calls. Setting a launch price. Feeding a price into a forecast someone is accountable for. Supporting a value-for-money claim. Choosing between two prices whose difference sits inside the replicate range.
In each of those the synthetic study is preparation. It narrows the ladder you will field and tells you which concepts are worth live respondents at all. That is real work, and it is not the same work as producing the number.
A 2026 paper on validating language model simulations states the standard in its title: this human study did not involve human subjects.⁹ Keep that line in the methods slide, not in a footnote.
Frequently asked questions about pricing research on synthetic panels
Four questions come up in every methods review of a synthetic pricing study: whether the output can set a launch price, whether census grounding fixes the anchor, whether trade-off formats are safer, and how many replicate runs a price study needs.
Answers below assume a census-grounded synthetic panel running a stated-price instrument, as of September 2026.
Can a synthetic Van Westendorp set a launch price?
No. The optimal price point produced by a synthetic price sensitivity meter is a function of the range implied by your brief, and it will move when the brief moves. Use the study to choose the ladder and the concepts you take into live fieldwork, and let the live study set the price.
The check is simple. If the optimal point changes materially between two anchor conditions, it was never a market reading.
Does grounding personas in census data fix the anchor problem?
No. Census grounding fixes who is in the panel, which is the sampling frame: the distribution of age, income, region, education and household type. Anchoring happens after that, when a persona is asked to produce a number and takes its reference from the prompt rather than from a purchase history it does not have.
The two problems are independent. A panel can be correctly composed and still return a price that reflects your wording.
Is choice-based conjoint safer than direct price questions?
Relatively, because the respondent pays a price against a feature instead of naming one in the open, which removes the free-floating reference point. It is not a clean pass. A synthetic panel randomises text rather than a decision, so conjoint carries its own set of limits on a synthetic panel.
The full breakdown of which parts of a trade-off study survive is in the conjoint and MaxDiff guide linked above.
How many replicate runs does a synthetic price study need?
Five per anchor condition is a workable floor, which means ten runs for a two-anchor design. Five small runs cost less than one large run and tell you something the large run cannot: how much the answer moves when nothing about the study changes except the draw.
Report the spread across those runs next to every figure you show. A price point without a replicate range is a single draw presented as a finding.
Sources
- A New Approach to Study Consumer Perception of Price — Peter van Westendorp, Research World
- Anchoring Bias in Large Language Models: An Experimental Study — arXiv
- Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs — arXiv
- Price as an Indicator of Quality: Report on an Enquiry — Andre Gabor and Clive W. J. Granger, Economica
- Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes — arXiv
- Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices — arXiv
- A Meta-analysis of Hypothetical Bias in Stated Preference Valuation — James J. Murphy, P. Geoffrey Allen, Thomas H. Stevens and Darryl Weatherhead, Environmental and Resource Economics
- When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels — arXiv
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence — arXiv