Back to blog

When to Use Synthetic Users in Product Research

You are deciding whether to remove a feature, redesign a familiar workflow, or change a plan. A synthetic-user exercise could return feedback today. But would that feedback answer the product question—or merely sound like it does?

The debate often collapses into “synthetic users replace research” versus “synthetic users are useless.” Neither framing helps a PM choose a method. The responsible choice depends on the evidence required, the consequence of being wrong, the reversibility and blast radius of the next step, and how much the team already knows.

Use synthetic users when a product decision is early, reversible, and only needs directional input, such as pressure-testing assumptions, generating objections, or preparing real-user research. Combine simulation with real participants when uncertainty or consequences increase. Skip simulation and go directly to real research for behavioral, usability, accessibility, pricing, safety, compliance, market-sizing, or high-impact decisions.

That input is model-generated, not observed customer evidence. The framework below turns this boundary into three practical paths: use synthetic users, combine them with real research, or skip them.

Synthetic users produce directional input, not customer evidence

What “synthetic user” means here

A synthetic user is an LLM-generated response produced from a supplied profile and scenario context. The profile might include known demographics, goals, beliefs, prior research, or product context. The prompt and model also shape the answer.

More relevant grounding can make an answer more relevant to the scenario you framed. It cannot turn generated text into an observed human response. Nielsen Norman Group defines synthetic users as AI-generated profiles that express simulated thoughts, needs, and experiences, and recommends using their outputs for hypotheses and research preparation—not final decisions (NN/g).

What the output can and cannot support

Can support Cannot establish
Plausible objections to investigate What customers will actually do
Alternative hypotheses Causal or behavioral validation
Research questions and scenario variants Market size or segment prevalence
Early risk-language review Usability or accessibility proof
A reversible next test Willingness to pay, safety, or compliance

Confidence label: Directional · Model-generated · Unvalidated

Methodological limitation: Fluent LLM output can resemble a polished interview transcript. Fluency is not evidence that a person experienced, believed, or would act on the response. A 2026 preprint comparing GPT output with human results from 12 first-click tests found material behavioral mismatches and warned that persona prompting can increase believability without fixing fidelity (Kuric, Demcak, and Krajcovic). Treat coherence as readability, not confidence.

Choose one of three evidence paths

Start with four gates:

  1. Evidence needed: Are you asking for ideas, reported attitudes, observed behavior, task performance, prevalence, or assurance?
  2. Consequence: What harm follows if the answer is wrong?
  3. Reversibility and blast radius: Can you cheaply undo the next step, and how many customers could it affect?
  4. Uncertainty: Are you challenging known assumptions, or trying to discover something absent from your inputs?

Then choose a path:

Decision factor Outcome 1: Use synthetic users Outcome 2: Combine with real research Outcome 3: Skip synthetic users
Question type What might we have missed? What should we prepare, then verify? What will people do, experience, or need?
Evidence needed Ideas and hypotheses Directional preparation plus human evidence Behavior, performance, prevalence, lived experience, or assurance
Consequence Low Meaningful but manageable High harm or high stakes
Reversibility Next step is cheap to undo Simulation step is reversible; decision awaits humans Choice is hard to reverse or broad in impact
Permissible output Questions, risks, messages, or scenarios Research plan plus clearly separated human findings Real-participant evidence or qualified review
Next step Run a reversible test with a real-evidence checkpoint Collect human evidence; it governs any conflict Recruit, measure, test, or seek specialist review now

Decision flow: begin with the evidence required; any mandatory escalation gate overrides the lower-risk paths.

flowchart TD
    A{What evidence is required?}
    A -->|Behavior, performance, prevalence, or assurance| S[Skip simulation: real research or qualified review]
    A -->|Ideas, objections, or research preparation| H{Could being wrong cause material harm?}
    H -->|Yes| S
    H -->|No| R{Is the next step reversible?}
    R -->|Yes: this step only needs direction| U[Use synthetic users for directional input]
    R -->|Yes: decision still needs human evidence| C[Combine simulation with real research]
    R -->|No| S
    G[Any mandatory escalation gate] -. Override .-> S

Outcome 1 — Use synthetic users

Use simulation alone for this step only when all of these conditions are true:

  • The task is hypothesis generation, objection generation, scenario comparison, or research preparation.
  • The immediate action is low-consequence and reversible.
  • Nobody will present the output as customer truth, prevalence, validation, or prediction.
  • Existing profile and customer inputs are relevant enough to frame the scenario.
  • The team has named the point where it will gather real evidence.

The permissible deliverable is a prioritized list of questions, risks, messages, or scenarios to investigate. The output should lead to a reversible next action, such as rewriting a research question or preparing two prototype variants—not a ship decision.

Outcome 2 — Combine synthetic users with real research

Choose this path when simulation can improve preparation but the product decision still needs human evidence:

  1. Run a synthetic pressure test to expose assumptions and missing questions.
  2. Convert the output into an interview guide, prototype task, recruiting criterion, or alternative explanation.
  3. Collect evidence from real participants.
  4. Where simulated and human findings conflict, treat the human findings as authoritative.
  5. Document what changed after the real research.

This fits a consequential but reversible redesign concept, feature-removal communication, or pricing-change message review. Simulation prepares the study; it does not approve the change.

Outcome 3 — Skip synthetic users and go to real research

Skip simulation when the core question can only be answered through real people, observed behavior, representative measurement, or qualified review. This is not an anti-AI position. It is method fit: plausible model output cannot establish customers’ task performance, spending behavior, or lived access needs; supply legal or regulatory assurance; or support population estimates.

For a broader method-selection workflow after choosing a path, see how to test product changes before launch.

Mandatory gates that require real participants or qualified review

These gates are non-negotiable. If any applies, synthetic output may help prepare the work, but it cannot provide the required evidence.

Escalation gate Required route
Behavior or adoption Use analytics, experiments, field evidence, or real participants to learn what people click, complete, retain, abandon, or adopt.
Usability and comprehension Observe real users completing representative tasks and explaining their understanding.
Accessibility and lived experience Involve people with relevant access needs and qualified evaluators. W3C recommends combining evaluation with users with disabilities and standards-based conformance work (W3C WAI).
Pricing proof Use real customers and appropriate pricing or behavioral research for willingness to pay, price sensitivity, conversion, prevalence, and revenue impact.
Market sizing or representativeness Use a defensible sample and statistical method for population estimates, segment share, frequency, or inference. Synthetic profile variety is not sampling.
Novel discovery Speak with people to uncover needs, language, workarounds, contexts, and experiences missing from existing inputs.
Safety, privacy, legal, regulatory, or compliance Route assurance, approval, and risk determinations to qualified specialists and, where relevant, real participants.
High harm, vulnerable populations, high blast radius, or hard-to-reverse choices Escalate regardless of how coherent or unanimous the simulation appears.
Thin, stale, biased, or untraceable grounding data Gather or repair the source evidence. Do not simulate confidence from missing context.
Conflict with real evidence Real observed or reported evidence governs. Investigate the discrepancy; never average it with synthetic output.

Synthetic users must never be the sole ship/no-ship evidence for any case in this table.

Product-change pressure tests that fit the framework

The following examples are illustrative. They are not customer stories, validation studies, or claims about genjury outcomes.

Illustrative example 1: Feature removal — generate perceived-loss hypotheses

Input: A precise description of the feature being removed, affected workflows, available usage data, and known segment context.

Permissible synthetic task and output: Ask for plausible perceived losses, workarounds, trust concerns, and unanswered questions. Turn the response into interview prompts and mitigation hypotheses, each carrying the confidence label directional, model-generated, unvalidated.

Mandatory real evidence: Review product analytics and interview affected customers before removal. Simulation cannot establish usage, dependency, workarounds in practice, or churn impact. The reversible next step is to improve the research plan or prototype a mitigation—not remove the feature.

Illustrative example 2: UX redesign — prepare habit and comprehension checks

Input: The current and proposed workflow, known user habits, change rationale, and segments likely to be disrupted.

Permissible synthetic task and output: Brainstorm where expectations, learned habits, terminology, or perceptions of control might break. Convert those possibilities into prototype tasks, recruiting priorities, and neutral questions.

Mandatory real evidence: Run usability sessions and accessibility evaluation with relevant people. A simulated critique cannot justify saying the redesign “tested well.” Use the dedicated guide to prevent product redesign backlash when planning rollout, communication, and rollback.

Illustrative example 3: Pricing or plan change — pressure-test messaging, not willingness to pay

Input: The current and proposed offer, which entitlements change, known customer contexts, and the draft explanation.

Permissible synthetic task and output: Generate plausible fairness, transparency, loss, and communication objections. Use them to create alternative explanations and questions for real pricing research.

Mandatory real evidence: Real customers and appropriate quantitative evidence are required for willingness to pay, conversion, prevalence, and revenue conclusions. The pricing change backlash checklist continues into communication, review, segmentation, rollout, and monitoring.

Use synthetic users to prepare better real-user research

Improve the interview guide

Translate simulated objections into neutral questions. “Some users will feel cheated” should become “How, if at all, does this change affect the value you receive?” Add disconfirming questions, and never quote a synthetic answer as customer voice.

There is a sequencing risk here. Product practitioner Radhika Dutt argues that hypotheses formed before open human discovery can prime teams to hear confirming evidence and miss surprises (Radical Product). Use simulation to challenge established assumptions or pilot a guide—not to pre-script what an unfamiliar audience will say.

Improve recruiting and scenario coverage

Use the exercise to identify known segments, access needs, behaviors, or contexts that must appear in the human study. Do not invent a segment because a model produced a compelling persona, and do not mistake profile diversity for sample representativeness.

Improve prototypes and test plans

Generate alternative task wording, edge cases, explanation variants, and competing hypotheses. Carry them into real task-based studies, concept tests, or interviews. Keep simulated suggestions separate from observations gathered during the study.

Reconcile simulation with human findings

Record which hypotheses the human evidence supported, contradicted, or left unresolved. When evidence conflicts, real findings govern. Investigate whether the simulation lacked context, reproduced a stereotype, followed leading wording, or surfaced a low-frequency possibility. Do not publish an “accuracy rate” without a valid evaluation design and sufficient data.

Synthetic users are not AI-moderated real-participant research

The two methods can both involve an LLM, but the source of the evidence is different.

Method Who produces the evidence? What AI does Evidence status
Synthetic users A model generates the participant response Produces a simulated response from supplied context Directional, model-generated input
AI-moderated real-participant research A real person answers or performs tasks Moderates, probes, transcribes, or helps synthesize Human evidence, subject to study quality

Dscout’s documentation, for example, describes an AI moderator that dynamically questions participants while recording their screens, webcams, audio, and responses (Dscout). The participant remains human; the moderator is automated.

Adding an AI moderator does not make a real participant synthetic. Giving a model a detailed persona does not make it a real participant. Neither method is automatically rigorous: objectives, recruiting, question design, consent, analysis, and interpretation still determine study quality.

A responsible synthetic-user pressure-test workflow

Frame a bounded question

Ask for objections, assumptions, missing questions, or scenario variations. Do not ask, “Should we launch?” or “How many customers will churn?” A better prompt is: “What plausible trust concerns should we investigate before removing this control?” Define the reversible decision the output may inform.

Document grounding and gaps

Record the profile and context inputs, where they came from, their dates, and what is missing. Also record the model and prompt used. If source material is stale, biased, or untraceable, stop and improve the evidence base instead of generating more polished output.

Generate and challenge hypotheses

Ask for conflicting reactions, alternative explanations, disconfirming possibilities, and reasons the initial framing could be wrong. Fluent consensus is not confidence. The Market Research Society’s 2026 AI toolkit warns that LLM-generated open responses can be uniform and polished, losing the messiness and richness of human language (MRS).

Label the output and choose the next evidence step

Mark every artifact directional, model-generated, unvalidated. Route each material risk to the right method: analytics, interviews, usability or accessibility testing, an experiment, staged rollout, or specialist review. End the exercise with one of the three outcomes, a named owner, and a real-evidence checkpoint.

If a live experiment is next, define hypotheses and guardrails before exposure. The broader guide shows how to pressure-test product changes before launch without confusing preparation with proof.

Common questions about using synthetic users

Can synthetic users replace customer interviews?

No. They can help pilot an interview guide, expose leading questions, and generate topics to investigate. They cannot report lived experience or surprise you with an undocumented workaround as a real customer can. If the decision depends on needs, context, behavior, or consequences, interview real participants. Never present a simulated transcript or quote as customer evidence.

Can synthetic users validate a product idea?

No. They can pressure-test assumptions around an idea and help produce competing hypotheses or clearer prototype tasks. Validation requires evidence matched to the claim, such as real interviews, observed task performance, behavioral data, or a well-designed experiment. A synthetic response can help decide what to investigate next; it cannot prove demand or authorize launch.

Are synthetic users useful for pricing changes?

Only for directional preparation. They can generate plausible objections about fairness, transparency, packaging, grandfathering, or message clarity. They cannot establish willingness to pay, price sensitivity, conversion, objection prevalence, or revenue impact. Treat every output as a question for real customers and appropriate quantitative research, not as pricing proof.

How are synthetic users different from AI-moderated interviews?

With synthetic users, the model produces the participant response. In an AI-moderated interview, a real person provides the response while AI asks questions, follows up, records, transcribes, or assists with analysis. The latter is human evidence, although its quality still depends on recruiting and study design. Automating the interviewer does not automate the participant.

Use the lightest responsible evidence method

Use synthetic users when the immediate job is directional, low-consequence, and reversible. Combine them with real research when simulation can sharpen preparation but the decision still requires human evidence. Skip them when the question concerns behavior, usability, accessibility, pricing proof, representativeness, novel discovery, safety, compliance, or material harm. When real and synthetic findings conflict, real evidence governs.

To explore the bounded simulation workflow when access opens, join the genjury waitlist. Limitation: genjury simulation is directional preparation, not customer validation, behavioral prediction, or a replacement for real-user research.

About the author

Malte Hedderich is the founder of genjury, a customer response simulator for product teams at B2C software companies.

  • Experienced AI Engineer with a multi-year track record shipping production AI systems at enterprise scale.
  • Direct day-to-day collaboration with PMs in large B2C company setups.