Three Worked Examples: How Synthetic Consumer Research Runs in Practice

Methodology · 13 min read

TL;DR: This article walks through three worked examples of how a research study runs on a census-calibrated synthetic panel: white-space identification in a category, new product concept screening and iteration, and pricing via choice-based conjoint. Each example is illustrative and clearly labeled as such, not a client case study. For each, the article covers the business question, how the study is set up, the type of output produced, and the explicit limits where live validation is still required. The purpose is to give research and insights leaders a concrete map of what synthetic research actually looks like in practice, so the method can be evaluated on its mechanics rather than on claims. Any numbers shown are illustrative and hypothetical.

How should these three examples be read?

These three examples are illustrative worked walkthroughs of how a study is designed and executed on a census-calibrated synthetic panel, not case studies of real customers. No named clients, brands, testimonials, or claimed real-world outcomes appear. Any figures shown are hypothetical and labeled as such. The purpose is to make the method inspectable, not to argue results.

Insights leaders evaluating a new research method reasonably want two things: worked examples specific enough to see the mechanics, and honesty about what the method does and does not do. Vendor case studies routinely deliver the first while obscuring the second. This article separates the two on purpose.

Three research designs are covered because they represent the three most common demand-side questions a consumer insights function fields: 'where should we play' (white-space identification), 'is this specific idea any good and how do we improve it' (concept screening and iteration), and 'what should we charge' (pricing). Each design has an established live methodology with published norms, which gives the worked example a clear reference frame.

For each design, the walkthrough covers the same four fields: the business question, how the study is set up on a synthetic panel, what output it produces, and the explicit limits where live validation is required. That last field is the one most vendor materials skip. It is the one that matters most for a research lead deciding when to route work to which instrument.

What is a census-calibrated synthetic panel, in one paragraph?

A census-calibrated synthetic panel is a set of AI personas composed so that the panel's joint distribution of demographic and behavioral attributes reproduces the target population as reported in national statistics (ACS in the U.S., Eurostat in the EU, national equivalents elsewhere). Each persona is a structured profile grounded in that reference distribution, and the panel is queried like a survey sample would be, one respondent at a time, with written rationales attached to each response.

The census-calibration step is what separates a synthetic panel from a general-purpose language-model chatbot with a persona prompt. The reference distributions are published and auditable, the composition method is documented, and the panel's actual composition can be shown alongside the target for inspection.⁷ Two structural properties follow: the panel can be queried at the segment level with the confidence that the segment weights match the reference population, and the panel's outputs can be routed to comparison against a live benchmark using the metrics documented in the public-opinion research literature.⁸

Everything below assumes this baseline. When a worked example says 'the panel was composed to U.S. adults 25 to 54 who purchased in the category in the last six months', the assumption is that the census-derived reference for that population is the anchor and that the actual composition of the queried panel is documented against it. The rest of the study design layers on top.

Worked example 1: how does white-space identification run on a synthetic panel?

White-space identification uses a synthetic panel to surface unmet-need territories in a category by combining a Jobs to be Done-style elicitation with structured need-statement rating across segments. The output is a ranked map of need clusters showing where importance is high, satisfaction with current alternatives is low, and the pattern holds across the target segments. The purpose is to compress a wide territory into a shortlist worth exploring further with live fieldwork.

The business question is directional: in a mature category, where are there unmet-need territories the incumbents have not fully addressed, and which of those territories look attractive enough for the team to explore further.

Study setup, in an illustrative run. Compose the panel to the target category population using census-derived demographic anchors plus a category-behavior filter (for example, U.S. adults 25 to 54 who bought in the category in the last six months). Field a two-stage instrument. Stage one asks each persona to articulate the jobs it is hiring the category to do, in its own language, with a written rationale.⁵ Stage two takes the union of extracted need statements from stage one, dedupes them, and rates each need on importance and current-alternative satisfaction across the same panel, segmented by relevant subgroups.

Output. A two-axis map of need clusters (importance x satisfaction gap) with segment overlays. Clusters in the top-left quadrant (high importance, low satisfaction) are candidate white spaces. Clusters that show consistently in the top-left across multiple segments are stronger candidates than clusters that appear only in one narrow segment. Written rationales accompany each need statement, which lets a strategy team read the underlying language rather than just the numbers.

Explicit limits. Synthetic panels are directional on need articulation; the ranking of importance is the useful signal, not the specific gap width. The top three or four candidate territories from the synthetic pass are the ones worth investing live qualitative (in-depth interviews or ethnography) into before committing product or brand investment. The synthetic step compresses the territory. Live fieldwork confirms the territory is real.⁶

In an illustrative run, the panel might return 40 unique need statements clustering into 12 need territories, of which 3 to 5 sit in the top-left quadrant across two or more segments. Those 3 to 5 are the shortlist for live follow-up. Any specific gap-width or market-size number attached to those clusters is hypothetical and should not be treated as a demand estimate.

Worked example 2: how does concept screening and iteration run on a synthetic panel?

Concept screening and iteration on a synthetic panel takes a set of early-stage concepts, runs them through a standardized evaluation frame (relevance, differentiation, believability, purchase intent, likes and dislikes) across the target segments, and returns a comparative ranking with structured verbatim feedback. The output is the two or three concepts worth iterating further, plus the specific rewrites the panel suggests. Absolute purchase-intent scores are not the deliverable; comparative signal is.

The business question is comparative: given a pipeline of eight to twenty early-stage concepts, which two or three are worth iterating and putting in front of live respondents, and what are the specific weaknesses to address before that live round.

Study setup, in an illustrative run. Compose the panel to the target segments the concepts are aimed at, using the same census anchors and category-behavior filters as the white-space example. Field a within-subject evaluation instrument: each persona sees the concept stimulus (headline, benefit statement, brief description, image reference) and responds on a standard concept-evaluation frame consistent with published concept-testing practice.⁴ Include the standard measures (relevance, differentiation, believability, purchase intent on a labeled scale) plus two open-ended fields (what the persona likes, what the persona would change) with written rationales.

Iteration is where the synthetic panel adds distinctive value. After the first pass, the analyst rewrites the two or three top-ranked concepts using the specific 'what would you change' rationales, then re-fields the rewritten versions on the same panel composition. Two or three iteration cycles at synthetic speed and cost typically compress the same option-refinement work that would take a live team weeks.

Output. A comparative ranking on the standard concept-evaluation frame, segment-level cuts, and a structured verbatim library that shows exactly which phrases and features personas responded to positively or negatively. The comparative ranking is the primary signal. The absolute purchase-intent number is not.

Explicit limits. Two failure modes to plan around. First, absolute purchase-intent scores from a synthetic panel are not comparable to live BASES-style norms and should not be presented as if they were.⁴ The relative ranking of concepts within the same synthetic run is the read to use; the absolute score is not. Second, category-first-of-kind concepts (a new category the panel has no behavioral referent for) get flatter reads than incremental innovations do, so the synthetic method is more reliable for line extension and repositioning work than for genuinely category-defining novelty. Those novelty cases need earlier live qualitative in the process.

In an illustrative run, a team screening 12 concepts might land on 3 concepts worth live testing after two iteration cycles. The synthetic step has compressed the option set and improved the specific concepts before live budget is committed. The live step still runs, on a smaller and better-briefed cohort.

Worked example 3: how does choice-based conjoint pricing run on a synthetic panel?

Choice-based conjoint on a synthetic panel presents each persona with a sequence of forced-choice tasks over product profiles that vary on price and a small set of relevant features. The panel's choices are aggregated into utilities and share-of-preference simulations across price points, which is the standard conjoint output. On a synthetic panel this is a directional pricing map and starting-point calibration; committed pricing decisions still require live conjoint fieldwork.

The business question is quantitative: given a product with three or four meaningful features and a plausible price range, what does the demand curve look like across price points, and which feature-price combinations deliver the strongest share of preference against a defined competitive set.

Study setup, in an illustrative run. Compose the panel to the target category buyer using census and category-behavior anchors as before. Define the attributes and levels using the standard discipline of conjoint design (a small number of attributes, three to four levels each including a labeled price ladder, orthogonal design so the effects can be estimated cleanly).¹ Field the choice tasks across the panel with the standard task discipline: each task shows the persona a small choice set (usually three to four alternatives plus a 'none' option), and each persona completes a battery of tasks. Aggregate the choices into part-worth utilities using the estimation method appropriate to the design, then simulate share-of-preference across the price ladder for the competitive set of interest.

Output. Part-worth utilities per attribute level per segment, a demand curve across the price ladder with segment cuts, and a share-of-preference simulator that lets the team ask counterfactual questions ('what happens if we lower the mid-tier price by 10 percent while raising the premium tier by 5 percent'). These outputs match what a live conjoint run produces in structure, which is one reason conjoint is a natural first quantitative use case for synthetic panels.

Explicit limits. Two limits are important to name up front. First, synthetic panels are known to smooth the tails of price-sensitivity distributions, so the extreme ends of the demand curve (very steep price rejection at the top, very flat sensitivity at the bottom) are the least trustworthy regions and should not be relied on for committed pricing decisions.⁶ Second, the aggregate share-of-preference from a synthetic conjoint is a directional read, not a market forecast; committing final pricing needs the same design run on a live conjoint panel, ideally with the synthetic run used to narrow the tested price ladder before the live study is fielded.

In an illustrative run, a synthetic conjoint might identify a price-feature combination that dominates within a segment and a competitive set. The team would then narrow the live conjoint to test that combination and its two closest neighbors, saving the live study from testing the full grid. The synthetic step has compressed the design space. The live step confirms the specific pricing that will be committed.¹

What do the three examples have in common, and where do they differ?

All three examples follow the same operating rule: use the synthetic panel to compress a wide option set into a specific shortlist, then route the shortlist to live fieldwork for the commitment. They differ in what the compression produces (need clusters, iterated concepts, or a narrowed price grid) and in how directly the output can be acted on before live validation. The rule to hold in every case is that synthetic evidence is for exploration and comparative signal, and live evidence is for the specific commitment.

The common structure across the three examples is compression. In white-space, the compression is from many need statements to a few candidate territories. In concept work, the compression is from many concepts to a few iterated finalists. In pricing, the compression is from a wide price-feature grid to a narrowed set of combinations worth live testing. Compression is the operating benefit that justifies synthetic research economically. It is the reason a team can afford to explore more options before committing to any.

The differences matter for study design. White-space output is qualitative in character even when it uses rating scales, so the analyst reads the rationales as carefully as the scores. Concept work produces a comparative ranking that can carry directly into a live cohort brief. Pricing work produces the most structured output (utilities, demand curves) but also has the strictest live-validation requirement, because the eventual commitment is the most consequential.

A research function running all three of these designs on the same platform ends up with a stable operating pattern: synthetic compresses, live commits, and the routing rule is explicit. That routing rule is the durable capability. Any single study is only as good as the rule that decided how to route it.²³

One discipline holds across all three. Every study is presented with its instrument map: what the panel was composed against, what was queried, which reads to trust as-is, which to interpret with caution, and which are pending live validation. Presenting synthetic evidence without that map is where credibility breaks. Presenting it with the map is where the method earns its place in the insights stack.

Sources

  • Sawtooth Software Conjoint Analysis Papers — Sawtooth Software
  • ESOMAR Global Market Research Report — ESOMAR
  • GRIT Business and Innovation Report — Greenbook
  • BASES Concept Testing Norms — NielsenIQ BASES
  • Jobs to be Done: Theory to Practice — Ulwick, Strategyn
  • Out of One, Many: Using Language Models to Simulate Human Samples — Argyle et al., arXiv
  • American Community Survey — U.S. Census Bureau
  • AAPOR Standards and Best Practices — American Association for Public Opinion Research

Related Articles

  • Automated Concept Testing: How to Validate Product Concepts in Hours, Not Weeks — A guide to automated concept testing with AI personas: the 5-step workflow, scoring metrics, comparison to traditional tests, and when to validate live.
  • Price Elasticity Surveys in FMCG: How AI and Synthetic Research Are Changing the Game — How FMCG brands use surveys to derive price elasticity of demand, and how AI respondents and synthetic research accelerate and improve pricing decisions.
  • How to Run a Validation Study for AI Synthetic Consumer Research — A practical methodology for validating a synthetic consumer research panel against a live national survey: what to measure, how to design a fair benchmark, and how to present the evidence to skeptical stakeholders.