Conjoint and MaxDiff on Synthetic Panels: What Holds

Methodology · 10 min read

TL;DR: A conjoint asks a person to trade one attribute against another and reads part-worth utilities off the choices they make. That reading depends on levels being randomised to a person who then chooses under a real constraint. A synthetic panel randomises the text of a choice task, not the experience of a decision, so the utilities it returns describe how a model weights described attributes rather than how a market makes trade-offs. This guide splits a trade-off study into four parts, names which ones a census-grounded synthetic panel can carry, explains why MaxDiff behaves better than choice-based conjoint, and gives the three checks that grade synthetic utilities before anyone builds a simulator on them.

Can you run a conjoint analysis on a synthetic panel?

Partly, and the part matters. A trade-off study has four components: the attribute shortlist, the design and instrument, the part-worth utilities estimated from the choices, and the market simulator built on those utilities. A census-grounded synthetic panel is strong on the first, useful on the second, and unreliable on the last two.

Most teams asking this have one word for four different claims. The shortlist is a judgement about what is worth testing. The design is an instrument. The utilities are a measurement. The simulator is a forecast.

Error compounds downstream. A wrong shortlist costs you one wasted attribute. A utility that is materially off produces a share estimate nobody can bound.

So the useful question is not whether a synthetic panel can run a conjoint. It can produce choices, and choices produce utilities. The question is which of those four claims you are willing to put your name to.

Why does a synthetic conjoint stop being an experiment?

Because nothing is randomised to anyone. A conjoint earns its causal reading from levels assigned at random to a person who then chooses under a real constraint. A synthetic panel varies the text of a task and reads back a generated response, which recovers the association between described attributes and stated preference in the text the model was trained on.

The distinction is easy to miss because the output looks identical. Both produce a choice matrix. Both feed the same estimation code. Only one was an intervention.

A human respondent facing a choice task gives something up. The trade-off is the measurement. A model reads a description of two products and returns the option its training distribution associates with the profile it was given. Nothing was surrendered, so nothing was measured about willingness to surrender it.

The empirical record supports treating the output as associational rather than causal. Bisbee and colleagues, writing in Political Analysis, documented divergences between synthetic and human survey data that were themselves unstable across model versions.² Instability across versions is the signature of a property of the model rather than a property of the market.

Chris Chapman, presenting through Sawtooth Software, makes the stronger version: synthetic responses are not data in the sense a research method requires, because no sampling frame connects them to a population.³ Conjointly, which sells conjoint software, calls synthetic respondents the homoeopathy of market research.⁶ That is unfair to the upstream uses in this guide and fair to the downstream ones.

What breaks in synthetic part-worth utilities?

Four things, roughly in the order you will notice them. Price and brand get overweighted. Attribute non-attendance disappears, because a model reads every line of the task. Preference heterogeneity compresses, which starves any individual-level model. And the utilities move when the model version moves.

Salience overweighting is the first and the most predictable. A model's sense of what matters in a category comes from text, and text about products is dominated by price and brand. Attributes that people care about but rarely write about, such as pack format or service terms, arrive underweighted.

The direction of effect often survives even when the magnitude does not. Brand, Israeli and Ngwe found that language models asked willingness-to-pay questions produce demand curves that slope downward and respond sensibly to product features.¹ That is a real positive result, and it is a result about direction and ordering. A simulator needs magnitude.

Attribute non-attendance is the second and the least discussed. Real respondents routinely ignore attributes in a large design, and that behaviour is informative: it tells you the design is too big. A model attends to every line of the task, so an overloaded design comes back looking viable.

Compressed heterogeneity is the third. Hierarchical Bayes exists to estimate the spread of preference across individuals. Run it on synthetic choices and it returns a spread that is largely a property of the generation procedure, not of the market. Adding personas does not repair it, the same trap described in what free sample size hides on a synthetic panel.

Does MaxDiff behave better than choice-based conjoint?

Yes, on one axis. MaxDiff asks which item in a set is best and which is worst, and returns relative importance rather than absolute utilities, so it never feeds a share simulator that can be wrong in units. The ordering is what survives. The spread between respondents still does not.

MaxDiff, also called best-worst scaling, presents small sets of items and asks for the extremes. Sawtooth Software's documentation sets out the standard form and what it returns: scores that rank items against each other on a common scale.⁵

That output shape is a better match for what a synthetic panel can honestly produce. A ranking is a claim about order, and order is the part of the signal that survives salience overweighting most often. A part-worth utility is a claim about magnitude, and magnitude is the part that does not. MaxDiff also has no simulator attached, so no downstream step silently amplifies the error.

The limit is the same one that limits every synthetic method here. MaxDiff across a census-grounded panel will give you a plausible ordering of 30 message claims. It will not tell you that the ordering differs sharply between two segments, because the between-persona spread is compressed for the reason described in what survives a segmentation on a synthetic panel.

What is a synthetic trade-off study actually good for?

Four jobs, all before fieldwork. Cutting an oversized attribute list to a fieldable design. Checking that level wording reads the way you intended. Killing impossible level combinations before they reach a human design. And rehearsing the analysis plan on data shaped like the real thing.

Attribute cutting is the highest-value job and the one nobody markets. Conjoint designs fail from size more than from anything else. A product team arrives with 14 attributes, the design can carry six, and the argument about which eight to drop is usually settled by seniority rather than evidence.

Run a MaxDiff on the 14 across a census-grounded panel and you get an ordering to argue with. Drop the bottom five, keep the contested middle, take a six-attribute design to humans. That is a saving on scope, not a substitute for the wave.

Instrument pretesting is the second. Level wording carries assumptions that only surface when someone reads it back, and the rules for writing survey questions for synthetic personas apply to a choice task before any human sees it.

There is one route to a number, and it is not synthetic data alone. Wang, Zhang and Zhang, writing in Marketing Science, set out a data-augmentation framework in which model-generated responses supplement a small human sample under an explicit correction rather than replacing it.⁴ The correction is the point. Pooling both sources and averaging hides the offset instead of removing it, which is the method described in combining synthetic and human respondents.

Where PersonaHive fits is the first two jobs. Personas are grounded in national census data on a country-specific basis, built from aggregated public statistics and validated against real surveys, across nine national panels covering the United States, Germany, France, Austria, Czech Republic, Hungary, Romania, Denmark and Finland. Every response ships with a written rationale, which matters more in a trade-off task than anywhere else: the rationale names which attribute the persona actually read, and a choice count cannot tell you that. It does not make the utilities market-true, and no census frame does.

Smallest useful next step: take the attribute list from a conjoint you are scoping, run a MaxDiff on the matching census-grounded panel, and compare the importance ordering against your team's prior. A free PersonaHive account includes 250 credits, which covers it.

How do you grade synthetic utilities before trusting them?

Three checks, in increasing cost. Compare the attribute importance ordering against any human trade-off study you already own. Re-run the same design on a second seed and a second model version and see whether the ordering holds. Then score synthetic predictions against human holdout tasks.

The first check is nearly free and almost nobody runs it. Most categories have an old conjoint sitting in a drive somewhere. Rerun its attribute list synthetically and correlate the two importance orderings. A rank correlation that collapses is a finding about your category, and it arrives before you have spent anything.

The second grades the instrument rather than the market. Same design, different seed, different model version. If the top three attributes reorder, the ordering was never stable enough to cut a design with, and no amount of sample repairs it.

The third is the only one that grades prediction. Hold out a set of choice tasks from a human wave, predict them from the synthetic utilities, and report the hit rate against a first-choice baseline. This is the number a methods reviewer asks for.

One reporting obligation sits on top of all three. The ICC/ESOMAR International Code, revised in 2025, addresses synthetic data and transparency directly.⁷ If synthetic responses shaped the attribute list or the design, the method section says so, along with the model version and the panel definition.

Start with the rank correlation against an old study. It is the cheapest check in this guide and the one most likely to change your plan.

Frequently asked questions about conjoint and synthetic panels

Five questions that come up in every scoping call: whether a synthetic wave can replace a human one, whether more personas fix the estimates, whether hierarchical Bayes works, whether design pretesting is safe, and how direct price questions such as van Westendorp behave on the same panel.

**Can I replace a conjoint wave with a synthetic one to save budget?**

No. You can shrink the wave by cutting attributes first, which reduces design size and respondent burden. That is a saving on scope, not a substitution.

**Will more personas fix the utility estimates?**

No. More personas reduce sampling error, and the problem here is bias. A larger synthetic sample gives a tighter interval around the wrong number.

**Does hierarchical Bayes work on synthetic choice data?**

It runs and it converges. The individual-level variance it estimates is mostly a property of the generation procedure, so treat the segments it implies as hypotheses.

**Can I use synthetic responses to test my conjoint design before fielding?**

Yes, and this is the strongest use in the guide. Check task readability, level plausibility, and whether the design produces enough variation before a human sees it.

**What about van Westendorp or Gabor-Granger on a synthetic panel?**

Direct price questions are stated preference with no trade-off structure, so they inherit the salience problem with less to check it against. The workflow context sits in how AI changes price elasticity work in FMCG.

Sources

  • Using LLMs for Market Research (Working Paper 23-062) — James Brand, Ayelet Israeli, Donald Ngwe, Harvard Business School
  • Synthetic Replacements for Human Survey Data: The Perils of Large Language Models — James Bisbee and colleagues, Political Analysis, Cambridge University Press
  • Synthetic Survey Data? It's Not Data — Chris Chapman, Sawtooth Software
  • Large Language Models for Market Research: A Data-Augmentation Approach — Mengxin Wang, Dennis Zhang, Heng Zhang, Marketing Science, INFORMS
  • Introduction to MaxDiff — Sawtooth Software, Sawtooth Software
  • Synthetic respondents are the homoeopathy of market research — Conjointly, Conjointly
  • ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics — ICC and ESOMAR, International Chamber of Commerce and ESOMAR

Related Articles

  • Segmentation on a Synthetic Panel: What Survives — A segmentation depends on the joint distribution, not the marginals. Which parts of a segmentation study a synthetic panel can run, and which it cannot.
  • Price Elasticity Surveys in FMCG: How AI and Synthetic Research Are Changing the Game — How FMCG brands use surveys to derive price elasticity of demand, and how AI respondents and synthetic research accelerate and improve pricing decisions.
  • Combining Synthetic and Human Respondents: The Math — Combining synthetic and human respondents by averaging hides the bias. Measure the gap on a matched human sample, subtract it, widen the interval.