Brand Tracking on a Synthetic Panel: What Waves Mean
Methodology · 11 min read
TL;DR: A brand tracker reports a difference, not a level, so it only works when everything except the market is held still. A synthetic panel breaks that condition in a way most vendor material skips: the model underneath the panel moves on its own schedule, independent of anything happening in your category. Wave-over-wave movement therefore carries three things at once, real market change, model change, and instrument change, with nothing in the output labelling which is which. This guide separates them. It covers which tracker measures a census-grounded synthetic panel can honestly move, the control wave that makes model drift visible instead of invisible, the reporting format that survives a methods review, and the measures that belong on live fieldwork whatever the budget says.
Can a synthetic panel run a brand tracker?
Partly. A synthetic panel can track what its inputs carry, such as reaction to new messaging, a competitor's repositioning, or news events fed to each persona. It cannot track what only accumulated exposure produces, such as unaided recall built over years. The limit is not output quality. It is what the panel is made of.
Start with what a tracker is for. A brand tracker exists to report a difference. Nobody commissions one to learn that consideration is 34 percent. They commission it to learn that consideration moved from 34 to 37 after a campaign.
That makes the whole method conditional on holding everything except the market still. Same instrument, same sampling frame, same fieldwork conditions, wave after wave. Any of those moving contaminates the difference.
Most commercial trackers already use independent samples each wave rather than the same people, so a synthetic panel does not fail on the population condition. Composition can be held to the same census frame every wave, which is the part vendors usually worry about and the part that is least at risk.
The condition that breaks is the one nobody writes into the brief: the measuring instrument itself has to stay still. On a synthetic panel it does not, because the model underneath it is a product on someone else's release schedule.
What actually changes between two synthetic waves?
Three things, and only one of them is the market. The model underneath the panel changes as providers ship new versions. The instrument changes whenever wording, sampling, or configuration is edited. The market changes too. A raw wave-over-wave difference is the sum of all three, and nothing in the output labels which is which.
The model term is documented. Chen, Zaharia, and Zou tracked the same prompts against the same named model over several months and found substantial behaviour changes between snapshots, published in Harvard Data Science Review.¹ The name on the API stayed constant while the answers did not.
Bisbee and colleagues found the same instability in a survey context. Their examination of language model responses as substitutes for human survey data reported divergences that shifted across model versions, in Political Analysis in 2024.² A tracker built on that surface inherits the shifting.
The instrument term is easy to underrate. Gui and Toubia showed that prompt-design choices that look cosmetic can confound the estimated effect entirely.⁵ On a human tracker, changing an adjective in a question is a small risk. On a synthetic tracker it is a configuration change to the measuring device.
So when wave 2 reads four points above wave 1, you have one number and three unlabelled contributors. The rest of this guide is about labelling them.
Which brand tracking measures survive on a synthetic panel?
Measures driven by stimulus survive. Message resonance, claim preference, concept diagnostics, and reaction to a competitor's new positioning can be read wave over wave. Measures driven by accumulated exposure or actual purchasing do not. Unaided awareness, penetration, and share of requirement return the model's priors or nothing usable at all.
Sort the tracker questionnaire by what generates the answer, not by what the slide is called.
Stimulus-driven questions put the thing in front of the respondent and ask for a reaction. The persona has everything it needs. Exposure-driven questions ask the respondent to retrieve something from a lifetime of ambient contact with a category. A persona has no such lifetime, so what comes back is the model's prior over brand names, which is a fact about the training corpus rather than about your market.
Santurkar and colleagues showed that language model opinion distributions align with some populations far better than others, which is the same problem seen from the population side.³ Verasight's study of synthetic omnibus data reached a related conclusion at survey level: how well a model reproduces survey answers depends sharply on the topic being asked about.⁴
The practical version is a sorting rule. If a live respondent would have to remember something to answer, treat the synthetic reading as unusable. If a live respondent would only have to react to something you showed them, the reading is worth having.
How do you build a synthetic tracker that measures something?
Add a control wave. Alongside each new wave, re-field the previous wave's exact instrument on the same configuration. The difference between that wave's original result and its re-run is the system's own movement. Subtract it from the raw change and what remains is an estimate of market movement rather than a number mixing both.
Five steps, in order.
1. Pin the configuration and write it down. Model version, provider settings, panel definition, instrument, sampling. These are the six fields that make a synthetic study re-runnable, and a tracker needs them recorded per wave rather than once.
2. Freeze the instrument harder than you would on a human tracker. No rewording, no reordering, no adding a brand to a list mid-series. Additions go in a separate module that is never compared backwards.
3. Run the control wave. Every time you field wave N, re-field wave N minus 1 unchanged. You now have that wave measured twice on two different dates. The gap between the two readings of the same wave is your system movement, isolated.
4. Do the subtraction and publish both numbers. Take an illustrative case. Wave 2 reads 4 points above wave 1. The re-run of wave 1 reads 3 points above its original result. Market movement is therefore about 1 point, and the drift term is three times larger than the signal you were about to present. That is a finding, not a failure.
5. Anchor the level periodically. The control wave protects the direction. It says nothing about whether the level is right. That needs a small matched human sample and the estimator for correcting a synthetic reading against human completes.
The control wave costs one extra field per period. On live fieldwork nobody would pay to re-run last quarter's survey twice. On a synthetic panel the marginal cost of a re-field is low enough that the control is affordable, which is the one structural advantage this method has over the thing it is imitating.
When should a brand tracker stay on live fieldwork?
When the deliverable is a level rather than a direction. Equity numbers quoted to a board, penetration and share reconciled against sales, regulated claims, and any figure feeding a valuation or an incentive plan need fielded data with documented sampling. Synthetic tracking sits upstream of those measures, not in place of them.
The honest split is by what the number is used for.
Direction at speed is where a synthetic tracker earns its place. Did the new campaign line land better than the old one, did the competitor's repositioning move how our brand is described, is the message that tested well in June still the strongest in September. Those are readings you would never buy a live wave for, run at a frequency live fieldwork cannot match.
Levels that leave the building stay live. If the number appears in an annual report, an incentive scheme, a regulator submission, or a valuation model, it needs sampling you can document to a third party. That constraint sits alongside the other failure modes where synthetic evidence breaks.
There is a middle path worth knowing about. Kantar has published on synthetic data boosting in brand health tracking, where a reduced human wave is augmented rather than replaced.⁶ It is a different trade from the one described here: boosting keeps the human sampling frame and buys back sample size, while a synthetic tracker gives up the frame and buys frequency. Brand, Israeli, and Ngwe's Harvard Business School work on using language models for market research is the broader reference for where model-derived findings have held up against known results.⁷
Where PersonaHive fits is narrow and specific. Its personas are grounded in national census data on a country-specific basis, built from aggregated public statistics and validated against real surveys, across nine national panels covering the United States, Germany, France, Austria, Czech Republic, Hungary, Romania, Denmark, and Finland. That matters for tracking because a control wave only works if the panel is composed against the same frame each time, and census grounding per country is what makes a multi-market tracker comparable across markets rather than only within one. Every response ships with a written rationale, which is how you tell a genuine shift in reasoning from a scoring artefact when the control wave flags movement. Real-time news exposure tied to each persona's profile is the mechanism by which an outside event can reach a wave at all.
Smallest useful next step: before you design a tracker, measure your noise floor. Take the core question from your current tracker, run it twice on the same census-grounded panel with the identical configuration, and read the gap between the two runs. Any wave-over-wave movement smaller than that gap was never going to be real. A free PersonaHive account includes 250 credits, which covers the exercise.
Frequently asked questions about synthetic brand tracking
**Can I just pin the model version forever and avoid drift?**
Only until it is retired, and retirement dates are published rather than open questions. Run the old and new versions in parallel for one wave, record the offset, and treat the series as having a documented join rather than a silent one.
**Does a bigger panel make wave-over-wave movement more reliable?**
No. More personas tighten the interval around a single wave's estimate. They do nothing to the drift term, because drift is a property of the model, not of sample size. A large panel makes an unreliable difference look precise.
**Can a synthetic tracker replace a live tracker to cut cost?**
Not on the same measures. It can replace the wave frequency you cannot afford, which is a different purchase. Teams that get value from this run a reduced live tracker for levels and a high-frequency synthetic read for direction.
**How do I present a synthetic tracker to a methods reviewer?**
With three numbers per wave rather than one: the raw movement, the control-wave movement, and the corrected difference. A reviewer who can see the drift term will argue with the size of it. A reviewer who cannot see it will reject the whole series, and will be right to.
**What if the control wave and the tracker move by the same amount?**
Then the wave contains no evidence of market change, and the correct report says so. This is the case a synthetic tracker is best at catching and the one a raw wave-over-wave chart is guaranteed to miss.
**Does news exposure make personas aware of real events between waves?**
It makes events reachable, which is not the same as making them absorbed the way a person absorbs a year of advertising. Treat news exposure as a way to test reaction to an event, not as a substitute for the exposure history behind unaided measures.
Sources
- How Is ChatGPT's Behavior Changing Over Time? — Lingjiao Chen, Matei Zaharia, James Zou, Harvard Data Science Review
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models — James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, Jennifer M. Larson, Political Analysis, vol. 32, no. 4, pp. 401-416 (2024), Cambridge University Press
- Whose Opinions Do Language Models Reflect? — Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, Tatsunori Hashimoto, Proceedings of the 40th International Conference on Machine Learning (ICML 2023), PMLR vol. 202
- Can Large Language Models Replicate Survey Data Across Topics? — Verasight research team, Verasight
- The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective — George Gui, Olivier Toubia, arXiv preprint 2312.15524
- A definitive guide to synthetic data boosting in brand health tracking — Kantar, Kantar
- Using LLMs for Market Research — James Brand, Ayelet Israeli, Donald Ngwe, Harvard Business School Working Paper 23-062