Who Sees Your Concept on a Synthetic Panel

Enterprise · 9 min read

TL;DR: Is it safe to test confidential concepts with AI personas? The honest answer has two halves. A synthetic panel shows your unreleased work to no respondents at all, which removes the several hundred strangers a human concept test discloses it to. In exchange, the brief passes through a research vendor and a model provider, so your protection stops being a click-through promise and becomes a contract. That is usually the better trade, and it is only better if you read the contract. Six controls decide the answer: training exclusion, retention window, subprocessor list, processing terms, internal access, and a documented deletion route. Both US and EU trade-secret law make this part of your own record, not just a vendor question.

Is it safe to test a confidential concept on a synthetic panel?

Safer on one axis and riskier on another. A synthetic panel shows unreleased work to no respondents, which removes the largest group of people who could leak it. In exchange the brief travels through a research vendor and a model provider. The question stops being about human discretion and becomes a question about contracts.

Almost everything written about synthetic research argues about accuracy. Procurement asks something else first: who else will see this before we launch it.

The two methods fail on opposite surfaces. A concept test with 400 people discloses your unreleased packaging, price or positioning to 400 strangers, and controls it with a tick-box non-disclosure agreement nobody reads. A synthetic run discloses it to nobody outside the processing chain, and controls it with whatever terms sit in that chain.

Neither removes exposure. They move it from many weak custodians to one accountable counterparty. That is usually the better trade, and only if the counterparty is accountable in writing.

This is a separate argument from the privacy argument for synthetic personas, which concerns respondents' personal data. Here the confidential material is yours.

Who actually sees your brief in a synthetic study?

Four parties, at most. Your own team, the research platform, the model provider behind it, and any subprocessor either one uses for hosting or logging. No respondent is in that list, and no respondent device renders your stimulus. The length of the chain, not the number of personas, sets your exposure.

Walk the chain in order and ask what each link keeps. Your team uploads the stimulus. The platform stores the project and sends prompts to a model. The model provider processes them and, depending on configuration, retains them briefly for abuse monitoring. Subprocessors host the infrastructure underneath.

The major model providers now address this in their commercial terms. OpenAI states that business data sent through its API platform is excluded from model training by default, documents a retention period with deletion, and offers a zero-retention option for eligible endpoints.¹ Anthropic's commercial terms likewise commit that customer inputs and outputs are not used to train its models.²

Those commitments are the floor, not your position. You contract with the research platform, not with the model provider, so your protection is the weaker of the two agreements. A vendor whose own terms are silent on training does not inherit the model provider's promise on your behalf.

How does that compare with a human concept test?

The exposure changes shape rather than size. A human test spreads your stimulus thinly across hundreds of people you cannot identify or sue. A synthetic run concentrates it in two or three corporate counterparties you can audit. Enforceability improves. Visibility into logs gets worse.

The table sets the two side by side for a standard 400-complete concept screen.

Does a concept test weaken your trade-secret protection?

It can, because protection is conditional on your own conduct. Both the US and EU definitions require the holder to have taken reasonable steps to keep the information secret. How you ran your pre-launch research is part of that record, which makes the method choice a legal fact rather than a preference.

Under 18 U.S.C. Section 1839, information qualifies as a trade secret only where the owner has taken reasonable measures to keep it secret.³ Directive (EU) 2016/943 uses near-identical wording: the information must have been subject to reasonable steps, under the circumstances, by the person lawfully in control.⁴

Neither text names market research. Both make your research design evidence. Disclosing an unreleased formulation to several hundred recruited strangers under an unverified tick-box is a weaker measure than disclosing it to one processor under signed terms with a stated retention window, and a dispute years later gets argued on exactly that comparison.

This is not legal advice, and the answer turns on your jurisdiction and your own documentation. The practical point holds regardless: keep the method, the counterparties and the retention settings in the file beside the concept.

Can your concept end up in another company's answer?

Not through training, if your inputs are contractually excluded from it. The research on models reproducing their training data concerns material that was in the training set. Exclude your inputs and that pathway closes. The risks that remain are mundane ones: retention, logs, access, and your own team.

The fear traces to memorisation research. Carlini and colleagues showed in 2020 that verbatim sequences could be extracted from a language model's training data,⁵ and a 2023 follow-up showed the attack scaling to aligned production models.⁶

Read what both papers are about. They recover data the model was trained on. If your brief never enters a training set, there is nothing of yours in the weights to recover. That is why training exclusion is the first control to verify, and the only one that answers this question.

What the papers do not cover is where real leaks happen: a retention window nobody checked, a shared workspace that outlived the project, a screenshot in a group chat, an unapproved consumer chatbot.

Which controls should you verify before uploading a confidential stimulus?

Six, and all six are answerable in writing. Training exclusion, retention window, subprocessor list, processing terms, internal access control, and a documented deletion route. Ask them of the research platform, then ask which of its answers also bind the model provider underneath it.

Two of the six are standard instruments rather than vendor policy. A data processing agreement meeting Article 28 of the GDPR sets out what a processor may and may not do with material you send it,⁷ and the European Commission's Standard Contractual Clauses cover the case where that processing happens outside the EU.⁸ Ask for both by name.

These sit inside the wider evaluation covered in the enterprise RFP checklist for AI research platforms. Pulled out on their own, they are the pre-upload shortlist.

When is a synthetic panel not the safer option?

Four cases. When the material cannot leave a jurisdiction and no regional hosting is offered. When the stimulus is physical and has to be handled. When the output has to be filable evidence about real people. And when your research vendor's own terms are weaker than your panel provider's.

The fourth case is the one teams miss. A traditional agency is not unprotected: the ICC/ESOMAR International Code binds member researchers to protect client confidential information as a professional duty.⁹ The difference between the two routes is the mechanism, not the presence of an obligation, so a thinly contracted synthetic vendor can be the riskier choice.

The third case is a hard boundary rather than a preference. A study destined for a regulator or a challenge has to show that people were asked, which is the argument in what synthetic research cannot substantiate.

The smallest useful next step is a two-stage upload. Run the first pass on a de-identified proxy stimulus, with the real names, numbers and dates stripped out, to learn the instrument and the platform. Put the real concept in once the six controls have come back in writing.

What else do insight and legal teams ask?

Five questions recur once the chain is visible: whether a vendor agreement reaches the model provider, whether self-hosting solves it, whether stripping details is enough, whether an unknown brand name skews the result, and what belongs in the methods note. Short answers follow.

Does our agreement with the research platform cover the model provider? Not by itself. Contracts bind the parties that sign them. Ask for the subprocessor list and written confirmation that the training exclusion and retention terms flow down to each one.

Does a self-hosted or open-weight model remove the problem? It removes the model provider from the chain and adds your own infrastructure to it. A real improvement for the most sensitive material, a cost increase everywhere else. It changes who holds the logs, not whether logs exist.

Can we strip identifying details instead of contracting for them? Partly, and it is worth doing for a first pass. A concept with the brand, price and launch window removed is still testable for appeal. It stops being testable once the decision depends on those specifics.

Does the model knowing nothing about our new name distort the reading? Yes, and in your favour here. An unreleased name carries no stored prior, which is the mechanism behind blind versus branded stimulus testing. A confidential concept is already close to a blind stimulus.

What belongs in the methods note? Three lines: the counterparties the stimulus passed through, the retention setting in force on the run date, and the date the project and its logs were deleted.

PersonaHive runs census-grounded synthetic panels across nine countries, and new accounts get 250 credits without a card. Spend that allowance on a de-identified proxy of the concept you are not ready to show anyone, and run the contract review in parallel rather than after.

Sources

  • API Platform Data Privacy — OpenAI
  • Commercial Terms of Service — Anthropic
  • 18 U.S. Code Section 1839: Definitions — Legal Information Institute, Cornell Law School
  • Directive (EU) 2016/943 on the protection of undisclosed know-how and business information (trade secrets) — Official Journal of the European Union
  • Extracting Training Data from Large Language Models — Nicholas Carlini and colleagues, arXiv
  • Scalable Extraction of Training Data from (Production) Language Models — Nicholas Carlini and colleagues, arXiv
  • Regulation (EU) 2016/679 (General Data Protection Regulation), Article 28 — Official Journal of the European Union
  • Commission Implementing Decision (EU) 2021/914 on standard contractual clauses for the transfer of personal data to third countries — European Commission
  • ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics — ESOMAR

Related Articles

  • Synthetic Personas, Privacy, and Ethics: No PII, No Consent Debt, No Re-Identification Risk — Synthetic personas remove three privacy risks that live respondent panels carry: personal data processing, consent management, and re-identification. Here is the compliance argument in plain terms, with the GDPR references that matter.
  • The Enterprise RFP Checklist for AI Consumer Research Platforms: 50 Questions, Scoring Rubric, and Red Flags — RFP checklist with 50 evaluation questions, a weighted scoring rubric, and red flags for selecting an AI consumer research or synthetic persona platform.
  • Claim Substantiation: Where Synthetic Data Stops — Regulators want evidence about people. What the FTC, NAD and ASA require to back an ad claim, and the three jobs a synthetic panel does before that study.

PersonaHive

  • Home
  • Pricing
  • Use Cases
  • Blog
  • Glossary
  • FAQ
  • Validation Report
  • AI Persona Platforms
  • Persona Authenticity
  • Why Traditional Research Breaks Down
  • Market Research Tools Guide
  • Customer Insights Platform
  • Brand Research Platform
  • Synthetic vs Traditional Panel
  • AI Focus Groups vs Synthetic Personas
  • About
  • Security
  • Recognition and Reviews
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Acceptable Use
  • Cookie Policy