Synthetic Research Reproducibility: Six Fields to Record
Methodology · 12 min read
TL;DR: Synthetic research reproducibility is a narrower question than most teams treat it as. It asks whether the same instrument, re-run tomorrow, returns the same reading. It does not ask whether that reading is correct, which is a validity question covered separately. The distinction matters, because a study can be perfectly reproducible and wrong, or genuinely valid and impossible to re-run. This guide stays inside the reproducibility lane. It separates the three kinds a research lead can honestly claim, sets out the six fields every synthetic study should record, and gives an instrument-stability test that shows whether the measurement itself is holding across model versions. Providers publish retirement dates for the exact versions studies run on, frontier models drift between releases, and inference is not deterministic even at temperature zero. At PersonaHive a study you cannot re-run is treated as a study you cannot defend.
Why does a synthetic study stop replicating?
Because the instrument is a model, and the model moves. Providers publish retirement dates for the exact versions studies run on, frontier models drift between releases, and inference is not deterministic even at temperature zero. A synthetic result is a measurement taken through a component that changes underneath you, on a schedule you do not control.
Three mechanisms are at work, and teams usually notice only the third.
Retirement is the blunt one. Model providers publish shutdown dates for specific API versions, and once a version is retired the exact instrument that produced your result no longer exists.⁴ A tracker fielded on a retired version cannot be re-run at all, only re-approximated.
Drift is the quiet one. Even where a name stays the same, behaviour moves across releases. A 2026 study of response drift across frontier models documents the pattern directly.³ Your March panel and your September panel may share a label and not much else.
Nondeterminism is the one that surprises people. Setting temperature to zero does not make a model deterministic in production, because inference kernels are not batch-invariant and results shift with server-side batching.² Two identical requests can return different answers for reasons that have nothing to do with your study.
Social scientists reached this conclusion first. Barrie, Palmer and Spirling argue that results produced through closed, changing models carry a replication problem no methods section alone can solve.¹
None of this makes synthetic research unusable. It makes an unrecorded synthetic study unusable.
How is reproducibility different from validity?
Reproducibility asks whether the same instrument produces the same reading on a re-run. Validity asks whether that reading matches reality. They are independent axes, they fail for different reasons, and they need different fixes. A study can be perfectly reproducible and wrong, or genuinely valid the first time and impossible to re-run. This guide covers the first axis only.
Confusing the two is the single most common mistake in synthetic research procurement, and it lets vendors answer one question when the buyer asked the other.
Reproducibility is a property of the run: the model version, the settings, the panel definition, the instrument text, the randomization scheme, the date. If those are recorded and can be re-created, the study is reproducible in the sense that matters. Whether the result is correct is a separate audit.
Validity is a property of the finding: whether the topline matches a real benchmark, whether the subgroup direction is right, whether the response spread looks human. Peer-reviewed work shows synthetic responses can be aggregate-plausible and subgroup-wrong, which is a validity failure that a perfect run record does not touch. That territory belongs to a separate field guide on where synthetic personas break.
The practical consequence: solve them separately. A run record cannot make a bad instrument valid, and a validation study cannot rescue a run you cannot describe. The rest of this guide stays on the reproducibility side.
What does synthetic research reproducibility actually mean?
It means one of three different things, and only two are achievable. Exact reproducibility, the same personas returning the same answers, is not available in production inference. Distributional reproducibility, the same shape of result inside a stated interval, is. Decision reproducibility, the same conclusion in the same direction on the anchor question, is the one worth contracting on.
Most arguments about synthetic reproducibility are really arguments about which of the three is being promised.
Exact reproducibility is off the table, and it is worth saying so plainly rather than discovering it in a stakeholder meeting. Inference kernels are not batch-invariant, so the same request can return a different answer depending on how it was batched on the server at that moment.² Temperature zero does not fix this.
Distributional reproducibility is achievable and testable. Fix the panel size and the interval before the run, then check that a re-run lands inside it. Work on uncertainty quantification for language-model survey simulation gives the useful framing: report a simulated result as an interval, and state how many human respondents the simulation is worth, rather than presenting a point estimate.⁵
Decision reproducibility is the level that matters commercially. The question is not whether persona 47 gave a 4 both times. It is whether the anchor question that routed the decision still points the same way.
Test at the level you will act at. Anything stricter costs money and proves nothing.
Which six fields should every synthetic study record?
Six: the model and its exact version identifier, the provider and inference settings, the panel definition and census reference, the instrument as fielded, the randomization scheme or seed, and the run date. Write them into the study file at the moment you field it. Any one of the six missing makes a later re-run uninterpretable.
Record these six when you field the study, not when someone asks six months later.
Model and version. The exact version identifier, not the family name. A note saying the study used a large language model is not a record.
Provider and inference settings. Temperature, top-p, system-level configuration, and the endpoint. These move the result and are usually the first thing nobody wrote down.
Panel definition and census reference. The country, the national statistical office, and the tables the panel was composed against. PersonaHive publishes the reference office and census source for each country it ships, which makes a panel definition restatable rather than described.
Instrument as fielded. Full question text, option lists, and scale labels, in the order the personas saw them. Small wording differences move synthetic answers more than human ones, so an instrument recorded loosely cannot be rebuilt.
Randomization scheme or seed. Option order should be randomized per persona, and the scheme has to be recorded for a repeat run to mean anything.
Run date. Drift is time-indexed. A result without a date cannot be compared to anything.
The 2025 ICC/ESOMAR International Code requires disclosure of methods, data sources and limitations, which puts most of this list in the report rather than an appendix.⁸
Smallest useful next step: open the last synthetic study your team ran and fill in all six fields from what was written down. Whatever you cannot fill in is your reproducibility risk.
How do you test whether the instrument is still stable?
Re-run one anchor question on the current model version, holding the panel definition, the instrument text, and the randomization scheme constant. Compare the re-run reading to the original at the level you will act at. This is a measurement-stability check on the instrument, not a check on whether either reading is correct. The correctness question belongs to a validation study.
An anchor re-run is cheap. A full replication is not, and it is rarely the right spend.
Pick the anchor before you need it. The anchor is the single question whose answer routed the original decision, not the question with the most interesting chart. Field it again on the same recorded panel definition and the current model version.
Read it at the level you contracted on. If you promised decision reproducibility, check that the direction on the anchor still holds. If you promised distributional reproducibility, check the re-run lands inside the interval you fixed before the original run.⁵ Do not silently raise the bar between runs.
Interpret movement as an instrument signal, not a truth signal. If the anchor reading shifts across a version change, that tells you the measurement is version-sensitive. It does not tell you which reading, if either, matches reality. That is a different question, answered by calibrating against a human anchor, not by re-running the model.
Treat wording fragility as an early warning. Rupprecht and colleagues ran 167,400 simulated interviews across nine models and found that small perturbations to questions and option lists moved answers substantially, with larger models holding their original answer far more often than small ones.⁶ An instrument that shifts under paraphrase will also shift under a version change, so pretest the wording before you contract on stability.
Set a tolerance in advance. Decide what movement in the anchor reading would prompt a re-validation before you see the re-run, and write it next to the original result. Tolerance chosen after the fact is not a test.
Every PersonaHive response ships with a written justification generated before the rating. That does not make a run reproducible, but it makes a divergence readable, and comparing rationales rather than ratings shows whether the instrument is reasoning differently or only landing on a different number.
What should you demand from a synthetic research vendor?
Four commitments in writing: the exact model version behind each run, notice before that version changes, an exportable run record you can archive, and the panel construction reference for each country. A vendor who cannot name the model version behind a study cannot help you defend it later, and no accuracy claim substitutes for that.
Procurement language does more here than a technical conversation.
Name the version. Ask the vendor to state in writing the model and version identifier behind any study they deliver. A vendor who will not put that in a report is asking you to defend a result you cannot describe.
Ask for change notice. You need to know before the model under your tracker changes, not after your wave-over-wave comparison stops making sense. Providers publish retirement dates, so this is a scheduling question.⁴
Ask for an exportable run record. The six fields should leave the platform as a file you archive next to the results. A record that lives only in the vendor's account makes your audit trail depend on your subscription.
Ask for the panel construction reference. For census-grounded panels that means the country, the statistical office, and the tables used. PersonaHive publishes these per country across the markets it ships, which lets a reviewer restate the panel definition without asking anyone.
These are reproducibility commitments. Validity commitments, such as how the vendor benchmarks against human panels and where they say synthetic evidence should not be used alone, are a separate section of the same evaluation. Both belong in the platform RFP, and the two should not be collapsed into one line item.
Barrie, Palmer and Spirling make the point for academic work, and it carries into commercial research. Where the instrument is a closed model that changes, the burden shifts onto documentation.¹ A vendor should carry most of that burden.
What else do teams ask about reproducing synthetic research?
The recurring questions are about tracker migration when a model version is retired, disclosure obligations, whether human panels have the same problem, and whether a bigger panel makes a run more reproducible. Short answers follow. The through-line is that reproducibility on a synthetic panel is a documentation discipline first and a statistical one second.
Should we keep trackers on a fixed model version forever? Only until it is retired, which is a published date rather than an open question.⁴ Run both versions in parallel for one wave, report the offset, and migrate on your schedule rather than the provider's.
Do we have to disclose that responses were model-generated? Yes. The 2025 ICC/ESOMAR International Code requires disclosure of methods, data sources and limitations, and the model version belongs in that disclosure.⁸
Does any of this apply to human panels? Some of it. Human panels drift too, in composition and in conditioning. The difference is that a human panel does not get retired on a published schedule while your tracker is running.
Does a higher persona count make a run more reproducible? Partly. More personas tighten the interval around a distributional result, which is a distributional-reproducibility gain. They do not protect against a version change, because that moves the whole panel at once, and they do nothing for validity, which is a separate audit.
Is replication studies work relevant here? Yes, as a template. The Greenbook replication study that tested synthetic data against academic benchmarks asks whether known results come back, which is the right shape for an instrument-stability check on your own anchor question.⁷
Then do it properly once. The next synthetic study you field should be one a colleague could re-run from the record alone. A free PersonaHive account includes 250 credits, enough to field a short anchor question and start the record on something real.
Sources
- Replication for Large Language Models: Problems, Principles, and Best-Practice — Christopher Barrie, Alexis Palmer, Arthur Spirling, Working paper
- Defeating Nondeterminism in LLM Inference — Thinking Machines Lab, Thinking Machines Lab
- Response drift across frontier large language models — arXiv preprint 2607.20454, arXiv
- Deprecations — OpenAI API documentation, OpenAI
- How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective — Chengpiao Huang, Yuhang Wu, Kaizheng Wang, arXiv
- Prompt Perturbations Reveal Human-Like Biases in LLM Survey Responses — Jens Rupprecht, Georg Ahnert, Markus Strohmaier, arXiv
- Testing Synthetic Data Against Academic Benchmarks: A Replication Study — Greenbook, Greenbook
- ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics — ICC and ESOMAR, International Chamber of Commerce