Message Testing on a Synthetic Panel: Rank, Not Score
Methodology · 9 min read
TL;DR: Message testing is the job a synthetic panel does best, because a written claim lives entirely in the channel a language model operates in. Nothing has to be seen, remembered or lived through. That also sets the limit. The panel ranks claims by what reads well to a model, and length, polish and framing move a synthetic result in ways they would not move a buyer. Pew Research Center reported synthetic samples missing human opinion by about 12 percentage points on average in September 2026, so absolute levels are not yours to quote. Normalise the claim set, use forced choice, randomise order, run replicates, and read the top group. Send that group to people.
Can a synthetic panel test marketing messages?
Yes, for rank order. A written claim is pure text, so a language model processes the whole stimulus with nothing missing. What comes back is a defensible ordering of a well-constructed claim set. What does not come back is a score you can quote, a predicted lift, or a number that holds against a human benchmark.
Most arguments about synthetic research are arguments about what the method loses. A persona cannot see an ad, cannot taste a product, and cannot recall a service failure that never happened to it.
Message testing is the case where almost nothing is lost. A benefit statement, a headline, a value proposition: each is a short piece of text, and text is the only material a language model has ever worked with. No part of the stimulus sits outside the channel.
That makes it the strongest honest use of a synthetic panel inside a marketing programme. It also makes the failure mode specific. When the stimulus is text, the model's own preferences about text become part of the measurement.
Why is message testing the easiest job on a synthetic panel?
Because the three things a persona cannot do are not required. A message test needs no sensory exposure, no episode the respondent lived through, and no behaviour to recall. It asks a reader to compare written propositions, which is the one research task where a model and a person start from the same material.
Set it beside the neighbouring jobs. An ad test measures reception and response at once, and reception depends on a viewing that a synthetic run never stages, which is why a persona cannot judge breakthrough. A customer experience study asks about an episode that did not happen.
A message test asks for neither. It puts eight to twenty written claims in front of a reader and asks which ones land. The work is comprehension and judgement about language, which is the layer the model is built on.
That is narrower than it sounds. Processing the stimulus is not the same as responding to it the way a buyer would.
What actually moves a synthetic message score?
Properties of the writing, more than properties of the proposition. Published evaluations of language models as judges document a preference for longer text and an order effect that changes the winner when positions are swapped. Models also drift toward whatever the prompt appears to favour. All three move a claim ranking.
The clearest evidence sits in the literature on models grading text. Zheng and colleagues, studying models as judges of other models' output, documented verbosity bias, a preference for longer responses, and position bias, where the order options are presented in changes which one wins.⁴
Sycophancy compounds it. Sharma and colleagues found that models trained with human feedback shift their stated judgements toward what the user appears to want.⁵ In a message test, the brief is where that signal leaks in. Name your favourite claim, or explain why the test is being run, and you have told the panel the answer.
Survey-specific work points the same way. Tjuatja and colleagues found models do not reliably reproduce human response biases, and on some wording manipulations move in the opposite direction to people.⁶ Dominguez-Olmedo and colleagues found model answers to survey items sensitive to option ordering and response formatting rather than stable across them.⁷
None of this makes a ranking worthless. It makes the stimulus set the part that has to be engineered.
How much does the persona layer actually add here?
Less than in most other study types, and the run should be planned for that. A 2026 sim-to-real evaluation of copy simulation reported a no-persona baseline outperforming persona-conditioned prediction of real audience response. Treat census grounding in a message test as a source of segment contrast, not as an accuracy upgrade on the overall order.
The paper states the result in its own title: a sim-to-real study where a no-persona baseline beats persona-based copy simulation.³ It is one study on one task, not a verdict on persona conditioning in general. It is a warning about this task.
The finding is plausible. The choice between two competently written claims in the same category runs on category logic the model already holds. Demographic conditioning adds variance without adding much information about which proposition is stronger.
Where the persona layer earns its place is contrast. The useful question is rarely whether claim C wins overall. It is whether claim C wins with one segment and loses with another, and whether the written reasons differ in a way you can act on.
Pew Research Center's September 2026 silicon samples work sets the scale for the absolute numbers. Across public opinion items, synthetic samples differed from human responses by around 12 percentage points on average.¹ ² Those were opinion questions rather than claim rankings, and the figure is the right order of magnitude to hold before anyone quotes a synthetic percentage outside the room.
How do you build a message set a synthetic panel can rank?
Six rules. Normalise length and register, write the whole set in one pass, strip brand and framing cues, use forced choice rather than ratings, randomise order per persona, and run replicates until the top group stops moving. The discipline sits in the stimulus, not in the analysis.
Normalise first. Every claim should sit inside a narrow word-count band and use the same register. A set where one option runs to thirty words and the rest to twelve is measuring length.
Write them in one pass, by one writer or one prompt. Claims drafted at different times carry different polish, and polish is one of the things being read.
Strip the framing. What a persona sees should be the claims and nothing about which one you hope wins. Sycophancy and acquiescence controls matter more here than in most designs, because a claim set is exactly the stimulus a model will try to be agreeable about.
Use forced choice. Asking a persona to pick the best and worst claim from a shown subset is standard best-worst scaling, which derives a ranking from discrete choices rather than from how a respondent uses a scale.⁸ It also sidesteps the scale behaviour above.
Randomise order per persona and run replicates. Order effects are documented,⁴ so one pass through one fixed order is not a result. Three runs is a practical minimum: two to see whether the order holds, a third to break a tie.
Keep the set small enough to show properly. Twenty claims in subsets is workable. Forty attributes with levels is a different instrument, and it belongs in a trade-off design.
What can you read off the result?
Three things: the rank order of the claim set, the size of the gap between the top group and the rest, and the written reasons personas give. Not the scores, not a predicted conversion rate, and not a one-place difference that moves the next time the set is run.
The table sets out what each output is worth. Comparisons inside one run carry weight; levels carried outside it do not.
PersonaHive runs synthetic panels grounded in national census data, country by country, across nine countries. Personas are built from aggregated public statistics and validated against real surveys. For a message test that is what makes the segment contrast meaningful: the split is drawn against a national distribution, not against a persona someone on the team wrote. New accounts get free credits without a card.
The smallest useful first run is a message test you already completed with people. Load the same claim set, see whether the synthetic ranking separates the same top group, and you have a calibration for your own category before the method is carrying a live decision.
When does a message test still need people?
Four cases. Any claim heading for regulatory or legal review, any message whose effect depends on delivery rather than wording, any final choice between two claims the synthetic run could not separate, and any market where the category vocabulary is thinly represented in the model.
The first case is a hard boundary rather than a preference. A claim you intend to make in advertising needs evidence that people were asked, which is the argument in what synthetic research cannot substantiate. A synthetic ranking is a design input, never the file.
The second is about where the message lives. A line that works because of how it is said, in a voiceover, on a pack, under three seconds of attention, has moved back into the reception layer. Test the words synthetically and the execution with people.
The third is the common one. Two claims finish adjacent across every replicate run. That is the panel reporting it cannot separate them, and the answer is a small human test on two options rather than a third synthetic pass.
The fourth is coverage. Category language in a smaller market may be sparsely represented, and a ranking written in that language inherits the sparsity. Pretest the set locally before reading the order as a finding.
What else do teams ask about synthetic message testing?
Five questions recur: how many claims belong in a set, how many personas to run, whether to show the brand, whether a synthetic ranking can be compared with an old human one, and what belongs in the methods note. Short answers follow, and each is a decision to make before the first run.
How many claims should a set contain? Between eight and twenty. Below eight the ranking has little to say. Above twenty the subsets get long and length effects grow.
How many personas? Enough that the top group stops moving between replicate runs, which is a stopping rule rather than a power calculation. The rule matters more than the number.
Should the brand sit inside the claim? Not in the first pass. Stored brand associations enter the judgement, so run blind first and add the name as a second arm if the name is a variable you care about.
Can a synthetic ranking be compared with an old human one? Yes, and it is the best calibration available. Re-run a completed message test, compare the two orderings, and record how far they agreed. That is evidence about your category, which no vendor benchmark can be.
What belongs in the methods note? Four lines: the claim set as fielded, the choice format, the number of replicate runs and whether the top group was stable, and one sentence stating that no human respondents took part.
Start with the claim set you already have a human result for. Run it, compare the orderings, and you will know what a synthetic ranking is worth in your category before it carries a decision.
Sources
- Silicon Samples and Synthetic Surveys: Can AI Stand In for Human Respondents? — Pew Research Center Data Labs, Pew Research Center
- AI survey samples poorly replicate human public opinion — Pew Research Center Data Labs, Pew Research Center
- Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation — arXiv
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — Lianmin Zheng and colleagues, Advances in Neural Information Processing Systems (NeurIPS 2023)
- Towards Understanding Sycophancy in Language Models — Mrinank Sharma and colleagues, International Conference on Learning Representations (ICLR 2024)
- Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design — Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkar and Graham Neubig, Transactions of the Association for Computational Linguistics
- Questioning the Survey Responses of Large Language Models — Ricardo Dominguez-Olmedo, Moritz Hardt and Celestine Mendler-Dunner, Advances in Neural Information Processing Systems (NeurIPS 2024)
- MaxDiff (Best-Worst Scaling) — Sawtooth Software