Back to blog

Synthetic Users vs Real User Research: When AI Simulation Helps Product Teams and When It Should Not

When product teams compare synthetic users vs real user research, the important question is not whether AI can produce a convincing answer. It is: “Is this customer evidence or plausible model output?”

Synthetic responses can help before a risky redesign, feature removal, or pricing change by exposing assumptions worth testing. They cannot tell a team what customers actually experienced, did, or will do. The practical boundary is decision rights: simulation may challenge a hypothesis, while real research and safe behavioral evidence support consequential product decisions.

Synthetic Users vs Real User Research: The Short Answer

Synthetic users cannot replace real user research. They can help product teams pressure-test a proposed change by surfacing plausible objections, comparing scenarios, and improving research questions. Use real participants whenever you need lived experience, observed behavior, new discovery, representative evidence, or confidence for a high-stakes product decision.

Compare Synthetic users Real user research
Source Model output conditioned by training, prompt, and supplied context Statements/actions from actual participants
Lived experience None Present, with normal self-report limits
Observed behavior Simulated Available through usability, field study, analytics, or experiments
Best use Pressure-test known assumptions; prepare research Discovery, context, comprehension, behavior, validation
Decision right Hypothesis only Depends on method quality; can support real product decisions

Synthetic users generate responses from a model. Real user research collects statements or actions from actual people through methods such as interviews, usability studies, fieldwork, surveys, or experiments. That difference remains even when the synthetic response is detailed, internally consistent, or grounded in customer data.

Speed and scale are operational properties. They do not establish validity. A fast simulation can produce many hypotheses quickly, but volume does not turn those outputs into customer observations.

Real research is not automatically trustworthy either. Poor recruitment, leading questions, unrepresentative samples, weak tasks, and overconfident analysis can all mislead a team. The comparison is therefore not “fallible AI versus perfect research.” It is simulated output versus evidence from real participants, with the quality and authority of the latter depending on the method.

NN/g recommends using synthetic users to supplement research and generate hypotheses, not for final decisions (NN/g, Synthetic Users). A 2026 systematic-review preprint covering 182 publications similarly frames synthetic participants as a supplemental, heuristic-like approach rather than substitutes for humans (Kuric et al., preprint).

What Synthetic Users Are—and Are Not

A synthetic user is an AI-generated participant or profile that produces simulated responses. Its answer depends on the underlying model, its training data, the prompt, and any context supplied by the team. It may sound like an interview participant, but no person had the reported experience.

Synthetic participant vs AI-moderated interview

Three uses of AI are often grouped together even though they create different evidence:

  1. Synthetic user: AI generates the participant’s response. The output is simulated.
  2. AI-moderated interview: AI asks questions of a real participant. The participant’s answers are real research data, subject to the quality of recruitment and moderation.
  3. AI-assisted research: AI helps prepare a study, transcribe sessions, code responses, summarize findings, or retrieve existing research. The underlying evidence still comes from real people or recorded behavior.

The distinction matters when findings reach a roadmap or launch review. An AI interviewer may affect how a session unfolds, but it does not replace the participant. A synthetic user does.

Grounding improves relevance, not automatic validity

A generic prompt draws on model training plus the persona instructions provided. Adding first-party interviews, support themes, demographics, values, beliefs, or product context can constrain the response to information the team already has. This may make the output more relevant to a known segment.

Task-specific benchmarking against humans may also support a narrowly defined application. But grounding alone does not create lived experience, prove that the profile represents a population, or turn the profile into a reliable digital twin.

Grounded synthetic output is a model-generated interpretation of existing evidence, not a new customer observation.

This boundary is especially important when a system uses interview transcripts. The transcripts may be real evidence; the model’s later answer remains a simulation derived from that evidence.

Why Plausible Synthetic Answers Can Mislead

No lived experience or unknown unknowns

A model can reproduce the form of a customer story without using the product, paying the bill, working around a limitation, or facing the consequence of a decision. It has no memory-as-personal-history, physical environment, social pressure, or stake in the outcome.

That makes it weak at revealing unknown unknowns. Real participants hesitate, contradict themselves, abandon tasks, use unexpected language, and describe workarounds a team did not know to ask about. A simulation can recombine known patterns, but it cannot make those experiences its own.

Believability is not fidelity

Synthetic answers can be polished enough to earn more trust than their evidence warrants. Common failure modes include agreeing with the premise, producing generic needs, treating every benefit as important, overstating certainty, and compressing or distorting variation between people.

In NN/g’s practitioner evaluation, synthetic interviews often produced shallow or overly favorable answers and struggled to prioritize needs (NN/g, Synthetic Users). The 182-publication systematic-review preprint groups problems across the literature into cognitive misalignment, distortions, misleading believability, and overfitting or contamination (Kuric et al., preprint).

The danger is not only an obviously wrong answer. It is a plausible answer that confirms the team’s framing, smooths over variance, and is promoted from “possible objection” to “what customers think.”

Population similarity does not identify an individual

A peer-reviewed CHI 2026 paper by Wang and Siu studied 51 knowledge workers and four AI document-workflow concepts. Interview-informed agents showed promise at approximating population-level response distributions but were imprecise at reproducing the individual participants on whom they were based (Wang and Siu, CHI 2026).

That is conditional support for bounded, early concept screening—not broad validation of synthetic research. The study’s authors limit their conclusion to that population, domain, set of concepts, and agent design. Questions about an individual’s workflow, trust, or adoption barriers still required authentic accounts.

Results are task-, domain-, model-, and prompt-specific

A result from one concept test does not automatically transfer to pricing, accessibility, churn, or a different population. Model choice, prompt structure, supplied data, task format, and evaluation metric can all change the outcome.

A separate 2026 design-preference preprint compared simulations with 29 real preference tests covering 2,073 participants. It reported systematic discrepancies in visual preferences and found synthetic justifications tended toward generic properties, elaboration, and overpraising (Kuric et al., design-preference preprint).

Positive or negative findings therefore need a human benchmark for the same task and domain. Market sentiment is not a substitute. User Interviews’ 2026 vendor research surveyed 150 researchers and added five moderated interviews; it is useful as a snapshot of attitudes and governance concerns, not as validation that synthetic users reproduce customers accurately (User Interviews, vendor research).

When Synthetic Users Help Product Teams

Synthetic users are most defensible before a team claims to have learned something new about customers. They can broaden a pre-mortem, challenge a known assumption, or improve the design of the real evidence step.

Good pre-ship uses

A product team may use labeled simulation to:

  • Generate plausible objections to an already-defined change.
  • Challenge assumptions about messaging, user control, fairness, trust, or perceived loss.
  • Explore scenarios based on known segment descriptions.
  • Rehearse an interview guide and identify weak or leading questions.
  • Turn existing interviews or support themes into traceable hypotheses.
  • Support a pre-mortem decision to advance, revise, escalate, or stop for more evidence.

These uses ask, “What might we have missed?” They do not ask the model to approve the launch or declare what customers will do.

Operating rules

Record the profile’s provenance, model and version, prompt, run date, supplied context, and exact scenario. Label every output as simulated and non-representative. Keep generated reactions separate from interview transcripts, survey responses, support records, and analytics.

Never present synthetic quotes as customer quotes. Do not create synthetic percentages or imply that repeated generations estimate prevalence. If a simulated finding would change prioritization, pricing, rollout, or launch approval, name the real method that must test it.

When Real User Research Is Mandatory

Discovery and lived context

Use real people when the team needs to discover unknown needs, current behavior, spontaneous language, real tradeoffs, or the context surrounding a problem. Interviews, contextual inquiry, diary studies, and support conversations can reveal experiences the team did not already encode in a prompt.

Real contact also helps teams understand emotional weight and competing priorities. A coherent persona response cannot create that empathy or establish how important one problem is relative to another.

Usability, accessibility, and observed behavior

Use representative participants when the question involves task completion, hesitation, error, recovery, comprehension, assistive technology, or environmental constraints. A redesigned flow must be used, not merely described to a model.

Accessibility research particularly depends on people who use relevant assistive technologies and encounter the real interface in context. Analytics, field studies, usability sessions, and experiments can supply observed behavior; synthetic users cannot.

High-stakes and final decisions

Real evidence is mandatory for pricing and willingness to pay; health, finance, privacy, safety, or legal-sensitive changes; decisions affecting marginalized or vulnerable groups; consequential feature removals; and launch or no-launch calls.

It is also required for exact estimates of demand, conversion, churn, retention, or revenue. Choose the method according to the claim: pricing research for willingness to pay, usage data for current dependence, usability testing for comprehension, and guarded experiments or staged rollout for live impact. The broader guide to test product changes before launch maps these questions to appropriate methods.

The Product-Change Evidence Gate

The Product-Change Evidence Gate is a decision aid, not a validated methodology. It assigns authority according to the decision, required signal, blast radius, and available benchmark.

Gate 1: What decision will this support?

Internal challenge and research planning may use simulation. Roadmap prioritization, rollout scope, or launch approval requires real evidence proportionate to the consequence.

Gate 2: What signal is needed?

A plausible objection may be simulated. Lived experience, observed behavior, prevalence, accessibility, or willingness to pay must come from real people or behavior.

Gate 3: What is the blast radius?

A low-stakes, reversible change can be screened with simulation before monitoring. Broad, hard-to-reverse, trust-sensitive, pricing, privacy, or safety changes should escalate to real research and the appropriate specialist review.

Gate 4: Is this system validated for the exact task/domain?

Without a human benchmark for the same task and domain, simulation is brainstorming. A transparent benchmark grants confidence only for the bounded task tested. Opaque, vendor-wide accuracy or parity claims are not sufficient decision evidence.

Product question Simulation role Real method Decision right
What objections might this change trigger? Generate a challenge list from known context Targeted interviews or a concept test with real participants if findings matter Simulation may create hypotheses
Does the interview guide expose our assumptions? Rehearse questions and probes Pilot with real participants Simulation may improve preparation
What unmet needs do customers have? Organize known themes only Discovery interviews, fieldwork, support evidence Real research creates the evidence
Can users complete the redesigned flow? Flag possible friction for the test plan Representative usability testing Observed task performance decides
Which visual design do users prefer? Brainstorm evaluation criteria Human preference test with appropriate recruitment Real participant choices decide
Will customers accept the price? Stress-test fairness and explanation language Pricing research and customer interviews Real evidence supports pricing decisions
Will removing a feature cause churn? Enumerate dependence and loss hypotheses Usage data, affected-user research, safe rollout Behavioral evidence supports the decision
Is this launch safe? Add risks to a pre-mortem Research, specialist review, staged rollout, monitoring Accountable humans approve launch
flowchart TD
    A[What decision must the team make?] --> B{Only challenge hypotheses?}
    B -->|Yes| C{Low-stakes, reversible, grounded?}
    C -->|Yes| D[Use labeled simulation]
    C -->|No| E[Use real research first]
    B -->|No| F{Need lived experience, behavior, prevalence, or launch confidence?}
    F -->|Yes| E
    F -->|No| G[Use simulation to prepare real evidence]
    D --> H[Document uncertainty and hypotheses]
    G --> H
    H --> I[Validate decision-changing findings with real users or behavior]

A Responsible Pre-Ship Workflow

1. Ground in current evidence

Start with recent interviews, support themes, product analytics, prior studies, and clearly sourced segment knowledge. Separate first-party facts from team assumptions. Manual structured inputs are valid context, but they should be labeled as such.

2. Describe one change and risk question

Specify the affected users, current experience, proposed experience, and feared outcome. “How will users react?” is too broad. “What control or fairness objections might existing subscribers raise about this bundle change?” is usable.

3. Record simulated risks as hypotheses

Save each plausible risk with its source context and uncertainty. Merge duplicates, but preserve meaningful disagreement. Do not count generated responses as votes.

Evidence ladder: simulated possibility → documented hypothesis → real-user evidence → safe live behavioral evidence

Each rung grants more decision authority. Simulation begins the ladder; it does not skip to the end.

4. Route to the lightest reliable real method

Use interviews for motivation and lived tradeoffs, usability testing for comprehension and task performance, pricing research for willingness to pay, and analytics or support data for current behavior. Use an A/B test or staged rollout only when exposure is appropriate and guardrails are ready. The guide to test product changes before A/B testing covers that readiness step.

5. Advance, revise, escalate, or stop

Advance a reversible change when the evidence is sufficient. Revise when research identifies a tractable problem. Escalate when the blast radius or uncertainty requires stronger research, accessibility, legal, privacy, security, or safety review. Stop when the expected harm or unresolved assumption exceeds the team’s evidence.

Three Product-Change Examples

Hypothetical; not genjury customer evidence.

Change Simulation may do Cannot establish Real evidence
Familiar navigation redesign Surface known habit/control risks; improve test guide Task success or acceptance Representative usability test; staged rollout
Legacy feature removal Enumerate perceived-loss scenarios Actual dependence or churn Usage data, affected-user interviews, rollback criteria
Price/AI bundle change Stress-test fairness and transparency language Willingness to pay, conversion, or churn Pricing research, interviews, appropriate review, safe live evidence

In the redesign scenario, simulation helps prepare the study; real users reveal whether navigation habits transfer. See the guide to product redesign backlash prevention.

For a feature removal, usage data establishes current dependence before interviews explain it. For pricing, simulated fairness objections may improve communication questions, but the pricing change backlash checklist still requires real pricing and launch evidence.

Where genjury Fits—and Its Boundary

Disclosure: the author, Malte Hedderich, is the founder of genjury.

genjury is a customer response simulator for B2C software product teams. Teams define LLM-powered customer profiles using structured inputs, describe a product change, and receive aggregated simulated reactions. Profiles can be created from manual structured inputs; they are not necessarily based on real interviews.

The useful output is a set of plausible objections, risk patterns, research questions, and mitigation ideas to examine before the next evidence step. Reports should be treated as directional, simulated risk signals—not research findings, customer evidence, representative samples, exact predictions, validation, or launch approval.

This boundary applies even when a reaction sounds specific. Product teams remain responsible for checking decision-changing hypotheses with real participants, current behavior, specialist review, or controlled exposure.

If that bounded role fits your pre-ship process, join the genjury waitlist. Use the output to develop hypotheses, not to approve a launch.

What to Ask a Synthetic-User Vendor

A vendor should make its evidence boundary easier to understand, not hide it behind a broad accuracy claim. Ask:

  • What exact task was benchmarked?
  • What human comparison set served as the reference?
  • Which domain, product type, and population were included?
  • How were people recruited, and what was the sample?
  • Is performance reported at population, segment, or individual level?
  • What variance, disagreements, and failure cases occurred?
  • How is each metric defined, and why does it fit the intended decision?
  • Are prompts, methods, data provenance, model version, and evaluation date documented?
  • How are low-confidence or unsupported outputs labeled?
  • Can a finding be traced to a supplied first-party source?
  • Which decisions does the vendor explicitly prohibit customers from making with simulation alone?

A benchmark can support only the use it actually tested. Evidence from a visual preference task, for example, does not establish pricing, churn, accessibility, or launch accuracy.

Q&A: Synthetic Users vs Real User Research

Can synthetic users replace real user research?

No. Synthetic users generate model output rather than observations from people with lived experience. They can support desk research, challenge assumptions, or help prepare a study. Discovery, usability, accessibility, validation, and consequential product decisions still require appropriately recruited real participants or observed behavior.

When should a product team use synthetic users?

Use them early to pressure-test an already-defined change, generate plausible objections, compare known scenarios, rehearse an interview guide, or turn existing evidence into hypotheses. Label the output as simulated, document its provenance, and identify the real method required before any finding changes a decision.

Does grounding synthetic users in interviews make them accurate?

Grounding can make responses more relevant to the interviews and context supplied. It does not guarantee accuracy, representativeness, or fidelity to an individual. The grounded response remains a model-generated interpretation of existing evidence, so the team still needs task-specific benchmarking and real research for decision-changing claims.

How are synthetic users different from AI-moderated interviews?

A synthetic user generates the participant’s answer. In an AI-moderated interview, a real person answers questions asked by an AI system. The latter can collect real participant evidence, although recruitment, consent, moderation quality, and analysis still determine how useful that evidence is.

Can synthetic users predict whether customers will churn?

They may surface plausible reasons a feature removal, price change, or trust issue could contribute to churn. They cannot establish actual dependence, prevalence, or a churn rate. Use product analytics, support and cancellation data, affected-user interviews, and an appropriately guarded rollout to measure behavior.

Conclusion

Synthetic users and real user research have different decision rights. Simulation can surface a possibility and turn it into a documented hypothesis. Real participants provide lived experience and observed evidence. Safe live tests and staged rollouts measure behavior when exposure is appropriate.

The boundary is simple: use AI simulation to make the next evidence step sharper, never to relabel plausible output as customer proof. Teams facing a sensitive change can also review how to avoid user backlash. To use genjury within the hypothesis stage of this ladder, join the waitlist.

About the author

Malte Hedderich is a machine learning engineer and the founder of genjury. He builds AI and agentic software workflows and writes about machine learning and AI systems at hedderich.pro.

  • Machine learning engineer with experience in artificial intelligence and MLOps.
  • Master of Science in Business Informatics from the Technical University of Darmstadt.
  • Has shipped multiple SaaS and software products and works with LLM-powered, agentic workflows.