Saturation Scores: How to Determine Sample Size for Synthetic Persona Research
Methodology ยท 11 min read
TL;DR: Saturation Score is a methodology for deciding when a synthetic persona study has produced stable insight, measured by how little new information each additional interview contributes. Unlike traditional power analysis, which fixes sample size up front from variance and effect-size assumptions, saturation is observed in flight: you keep adding personas until the marginal gain in new themes, new claims, and changed segment estimates falls below a threshold. For most synthetic studies this lands between 80 and 400 interviews depending on audience breadth, topic complexity, and the decision risk being supported.
What is a Saturation Score?
A Saturation Score is a running measure of how much new information each additional synthetic interview adds, expressed as the percentage drop in novel themes, claims, or segment-level estimate movement over the last batch of responses. The study is considered saturated when that percentage falls below a pre-set threshold and stays there.
Qualitative researchers have used the idea of saturation for fifty years: keep interviewing until you stop hearing anything new. In live qualitative work this is a judgment call by the moderator. In synthetic research it can be measured directly, because every interview produces structured output that the platform can score in real time.
The Saturation Score formalizes the judgment. After each batch of personas, usually 10 to 25 at a time, the platform compares the new responses against the cumulative pool and reports three signals: how many new themes appeared, how many new claims or attributes were mentioned, and how much the segment-level estimates moved. When all three signals fall below a threshold for two consecutive batches, the study is saturated and additional interviews stop earning their cost.
The practical implication is that sample size is no longer a guess made before the study begins. It is an observation made during the study, with a clear stopping rule.
How is saturation different from traditional power analysis?
Power analysis sets sample size before the study using assumed variance and a target effect size, optimizing for the ability to detect a quantitative difference. Saturation observes information gain during the study and optimizes for stability of insight. They answer different questions and are best used together, not as substitutes.
Traditional quantitative research relies on power analysis. The researcher specifies the smallest effect worth detecting, an expected variance, a confidence level, and a power target, and the formula returns a required sample size. For a typical concept test with a 5-point purchase intent scale and a 0.5-point detectable difference, that lands around n = 300 to 400 per cell.
Power analysis is the right tool when the goal is to estimate a population parameter or detect a difference between groups within defined error bounds. It is the wrong tool when the goal is to understand what people think, why they think it, and what variations of the idea exist in the audience.
Saturation answers the second question. The unit is not statistical power, it is information completeness. A study can be saturated at n = 80 if the topic is narrow and the audience is homogenous, or still unsaturated at n = 600 if the audience spans multiple sub-cultures and the question taps deep belief structure.
The mature program runs both. Power analysis sizes the quantitative cells when the deliverable is a number with a confidence interval. Saturation Score sizes the qualitative and exploratory work when the deliverable is a complete understanding of the response space. Synthetic research makes saturation cheap enough that it can be the default sizing method for almost everything upstream of a final go/no-go.
How is a Saturation Score actually calculated?
Three components are tracked after each batch: theme novelty (percentage of themes in the new batch not seen before), claim novelty (percentage of distinct claims, features, or attributes that are new), and estimate stability (the largest movement in any segment-level metric across the last batch). The composite score is the weighted average of all three, expressed as remaining information gain.
There is no single industry-standard formula, but the version used by most rigorous synthetic platforms looks like this.
After each batch of n personas, the platform extracts three signals from the cumulative response set.
Theme novelty. Open-ended responses are clustered into themes, durable patterns of reasoning the personas use to explain their reaction. Theme novelty is the count of themes that first appeared in this batch, divided by the total themes in the batch. A batch that introduces no new themes scores 0; a batch where everything is new scores 1.
Claim novelty. Structured attributes, features mentioned, objections raised, comparisons drawn, price points cited, are extracted and deduplicated. Claim novelty is the proportion of claims in the new batch not present in the prior cumulative set.
Estimate stability. For every reported metric (top-two-box purchase intent, mean appeal score, segment-level willingness to pay), the platform records the absolute change between the cumulative estimate before the batch and after. The largest such change is the instability signal; its complement is the stability signal.
The Saturation Score itself is the weighted average of theme novelty, claim novelty, and instability, expressed as remaining information gain on a 0 to 100 scale. A score of 100 means every batch is still introducing substantial new material. A score under 10, sustained across two consecutive batches, is the conventional stop signal.
What sample sizes does saturation typically produce?
Most synthetic studies saturate between 80 and 400 personas. Narrow questions on homogenous audiences saturate fastest (60โ120). Broad questions on heterogeneous audiences with multiple sub-segments take longer (250โ500). Highly novel categories or audiences with thin survey baselines may not saturate cleanly and should be flagged.
The advantage of running saturation as the sizing rule is that it adapts to the study. Three patterns recur in practice.
Narrow question, homogenous audience. A single concept tested against US category buyers aged 25 to 44 typically saturates at 80 to 120 personas. The response space is small; the demographic spread is moderate; new themes dry up within four to six batches.
Broader question or segmented decision. A multi-segment concept test, a positioning evaluation across attitudinal clusters, or a pricing study with three price points typically saturates at 200 to 400 personas. Each segment needs its own coverage, and the cumulative set has to stabilize within every reported segment before the overall study is saturated.
Novel category or fragmented audience. A genuinely new concept (e.g. a product category that does not yet exist), or an audience with thin survey baselines, may keep producing new themes well past 500 interviews. The right response is not to keep running until the budget is exhausted, it is to flag the topic as low-baseline, narrow the audience to a tractable starting segment, and pair the synthetic read with primary qualitative work.
The planning rule is to budget for the upper end of the expected range, monitor the Saturation Score in flight, and stop the moment the threshold is hit. Most studies finish well below the budgeted ceiling.
What thresholds and stopping rules are reasonable?
A conventional rule is to stop when the Saturation Score drops below 10 percent remaining information gain and stays there for two consecutive batches. For high-risk decisions, tighten to 5 percent across three batches. For exploratory discovery work, loosen to 15 percent across one batch. The threshold should be set before the study begins, not negotiated after the results land.
Saturation thresholds are decision-coupled. The more consequential the decision, the lower the acceptable residual information gain.
Exploratory / discovery work. Use a 15 percent threshold across a single batch. The goal is to map the space, not to certify completeness. One stable batch is enough to brief follow-on work.
Screening and iteration (the default). Use a 10 percent threshold across two consecutive batches. This is the right setting for most concept screening, message testing, packaging evaluation, and feature prioritization. Two stable batches in a row is strong evidence the response space has been mapped.
Directional read informing material spend. Tighten to a 5 percent threshold across three consecutive batches. This is the setting for studies that will be used to justify a media plan, a development brief, or a launch-readiness gate.
Final validation. Saturation is not the right sizing method for this tier. Use power analysis on a quantitative live or hybrid panel.
The critical discipline is locking the threshold before the study runs. Adjusting the threshold after the fact, usually to declare saturation early so the team can stop spending, undermines the credibility of every future read.
How do you report a saturated synthetic study to stakeholders?
Report four things: the saturation threshold used, the batch number at which the study saturated, the final sample size, and the residual information gain in the last batch. Pair the headline scores with confidence indicators by segment and a brief note on any segments that did not reach saturation.
Stakeholders need to know that the sample size was earned by the data, not by the budget. A defensible saturation report includes five elements.
Threshold and rule. State the threshold (e.g. "saturation defined as Saturation Score below 10 over two consecutive batches") and confirm it was set before the study began.
Saturation point. Report the cumulative sample size at which the rule was satisfied (e.g. "saturated at n = 180 after 9 batches of 20 personas").
Final sample. Report the total personas in the study, which is usually one batch beyond the saturation point to confirm stability.
Residual gain. Report the Saturation Score for the final batch, broken into theme novelty, claim novelty, and instability components. This is the equivalent of a confidence interval for a saturated study.
Segment-level coverage. Flag any segment whose internal Saturation Score is materially higher than the overall, with a note on whether the segment-level findings should be treated as directional only.
This reporting format gives reviewers everything they need to interrogate the study and is the format most likely to survive a procurement or methodology review at an enterprise research function.
Where does saturation fall short?
Saturation cannot detect a quantitative difference between two cells with statistical power, cannot certify rare-event incidence, and cannot substitute for live validation on regulator- or court-bound claims. It is a sizing method for insight completeness, not for hypothesis testing.
Three honest limits keep the method credible.
Quantitative comparisons. If the deliverable is a number, willingness to pay, market share, expected lift, sized to a defined confidence interval, saturation is the wrong tool. Use power analysis on a quantitative cell, synthetic or live, sized to the effect you need to detect.
Rare events. Saturation rewards stability. If the question depends on a low-incidence behavior (e.g. only 3 percent of users would convert), the study may saturate on the majority response well before the rare behavior is reliably observed. Pre-screen for incidence or oversample the rare segment.
Regulated or adversarial decisions. Claims defended in front of a regulator, in court, or in a published peer-reviewed context still belong to traditional fielded research with documented sampling. Synthetic studies can de-risk the front end; they should not be the citation on the claim itself.
Used inside its scope, saturation gives synthetic research the methodological discipline that turns a fast read into a defensible read.
What is the bottom line for research leaders?
Adopt Saturation Score as the default sizing rule for synthetic studies, lock the threshold before each study, report the saturation point and residual gain alongside every result, and keep power analysis for the quantitative cells that need it. The combination gives your team both the speed of synthetic and the rigor stakeholders expect.
The cost of an additional synthetic interview is small enough that the old sample-size argument, too small to be safe, too large to be affordable, does not apply. The new discipline is to size every study to its own information curve, stop when the curve flattens, and report the curve.
Research leaders who implement Saturation Score as a program standard get three durable benefits. Studies finish faster, because the stopping rule is observed rather than guessed. Costs fall, because no study runs longer than the data justifies. And stakeholder trust rises, because every read ships with an explicit sizing rationale anchored in the data itself.
That is the version of synthetic research that complements traditional methods rather than competing with them, fast where it should be fast, rigorous where rigor matters, and honest about both.