What Are Synthetic Users?
Product teams can now ask an AI-generated profile to react to a pricing change, redesign, or feature decision before launch. Here, “synthetic users” means the category; Synthetic Users is also the name of one company in it.
Synthetic users are AI- or LLM-simulated participants or profiles conditioned on context a team supplies. They can help generate hypotheses and screen risks before a product change. Their responses are generated, not observed: they cannot replace real-user evidence, represent a population, or predict how an individual will behave.
The useful distinction is simple: simulation can screen what to investigate, while real people and observed behavior decide questions that depend on lived experience or actual action.
What a synthetic user is—and is not
A synthetic user is a generated participant or profile produced by a model from instructions and context. A team might ask it to react to a changed price, explain possible objections to a redesign, or challenge the assumptions behind a feature removal. The result is a simulated response, not testimony from a person.
Implementations vary. A simple version may be a general-purpose LLM prompted with a short persona. A richer system may use documented research findings, structured customer attributes, product details, and a specific decision scenario. Nielsen Norman Group similarly describes synthetic users as AI-generated profiles that produce artificial research findings.
Keep two layers separate:
- Source material is the research, analytics, support evidence, profile data, and product context supplied to the system.
- Generated output is the model’s continuation from that material and its broader training.
Better source material may make a response more relevant. It does not turn the generated text into a new observation. A fluent, detailed answer can still be wrong, stereotyped, prompted into agreement, or silent about an absent population.
Synthetic user — is: an AI-simulated participant or profile that can generate hypotheses, objections, and questions to inspect.
Synthetic user — is not: a real participant, an independent research sample, a representative population estimate, or an individual digital twin. Generating more profiles does not by itself create independent observations or human consensus.
How synthetic users work
Inputs and grounding
A synthetic-user system conditions generation on some combination of:
- Existing research findings and documented customer evidence.
- Structured customer context, such as needs, constraints, habits, or relevant segment attributes.
- Product context, including the current experience and proposed change.
- A scenario, decision question, and affected user group.
- Task instructions, such as “identify objections” or “argue against this proposal.”
The input pack needs the same scrutiny as any other decision input. Record where each item came from, when it was collected, who it covers, and what is missing. Check whether personal or sensitive information may be used for this purpose, whether consent and privacy requirements permit it, and whether the model provider may retain the data. Name populations absent from the source material instead of assuming the model can fill them in.
Richer grounding can improve relevance within a tested task. In a 2026 revision of an empirical preprint, agents grounded in interviews or structured self-reports outperformed demographics-only agents on several social-science measures (Park et al.). That finding does not convert grounded output into observed product evidence; it shows that the content and depth of inputs matter.
Generation, interpretation, and validation
An LLM generates text conditionally: the model and version, prompt wording, and run conditions can affect what appears. In a peer-reviewed survey simulation study, small prompt changes and repeating the same prompt three months later produced materially different response distributions (Bisbee et al.).
Treat generated themes, objections, and edge cases as hypotheses. A reviewer should trace each one back to supplied evidence where possible, separate prompt echoes from genuinely useful combinations, and preserve disagreement rather than merging everything into one polished “user view.”
The final step is not more generation. Compare the hypotheses with existing research, analytics, support data, or domain evidence, then choose a fit-for-purpose validation method: interviews, usability testing, a survey, analytics, an experiment, specialist review, or staged rollout.
flowchart LR
A[Documented inputs] --> B[Model]
B --> C[Generated reactions<br/>GENERATED, NOT OBSERVED]
C --> D[Human synthesis]
D --> E[Validation method<br/>OBSERVED OR FIT-FOR-PURPOSE EVIDENCE]
E --> F[Decision]
C -. Evidence boundary .-> E
Accessible equivalent: Documented inputs enter a model, which produces generated—not observed—reactions. A human synthesizes those reactions, validates them with an observed or otherwise fit-for-purpose method, and only then uses the combined evidence in a decision.
What synthetic users can be useful for
Synthetic users are most useful before a team mistakes an unchallenged assumption for a launch-ready decision. They may surface possible objections, trust concerns, perceived loss, habit disruption, or misunderstood messaging. They can also prepare a research plan by generating competing hypotheses, scenarios, and recruitment gaps.
These uses apply to consequential changes without pretending to forecast their outcome. A team reviewing a price can use a pricing-change risk checklist; a redesign team can map habit and trust risks before launch; and a broader change review can examine the causes of avoidable user backlash.
| Job | Useful output | What it cannot establish | Next evidence step |
|---|---|---|---|
| Screen a pricing or packaging change | Possible fairness, value-loss, access, and messaging objections | Willingness to pay, acceptance rate, or revenue impact | Pricing research, customer interviews, experiment, or staged rollout |
| Challenge a redesign | Habit-disruption, control, trust, and navigation hypotheses | Whether people can complete tasks or will adopt the new design | Usability test with relevant users, followed by monitored rollout |
| Stress-test a feature launch or removal | Possible perceived loss, confusion, dependency, and support questions | Actual use, retention, cancellation, or support volume | Analytics review, interviews, prototype test, experiment, or staged release |
| Prepare research | Competing hypotheses, interview questions, scenarios, and recruitment gaps | Which needs are real or important | Recruit relevant participants and run the appropriate study |
| Check known conditions consistently | A repeatable challenge pass and requested counterarguments | Representative coverage or independent agreement | Audit missing groups, compare with observed evidence, and validate material gaps |
The useful output is therefore not “what users think.” It is a structured list of possibilities the team can inspect, reject, refine, or send to a stronger method.
The Screen / Pair / Escalate framework
Use one governing rule: the greater the consequence—or the more a decision depends on actual behavior or lived experience—the less weight simulation should carry.
| Decision question | Consequence | Evidence needed | Synthetic role | Required next step |
|---|---|---|---|---|
| What objections could a reversible copy change create? | Low and reversible | Plausible hypotheses | Screen | Label each output “hypothesis to inspect”; review and monitor |
| Which known risks might affect a meaningful feature change? | Moderate | Existing research, analytics, or support evidence plus hypotheses | Pair | Compare sources; prototype-test or research unresolved risks |
| Can target users understand and complete a changed flow? | Moderate to high | Observed comprehension and task performance | Escalate | Run usability testing with relevant participants |
| Will customers pay, adopt, stay, or change behavior? | Material business consequence | Observed willingness, choice, or behavior | Escalate | Use pricing research, analytics, an experiment, or staged rollout |
| How many people hold a preference or face a problem? | Population-level decision | An appropriately sampled quantitative estimate | Escalate | Run a designed survey or analyze suitable behavioral data |
| Could the change affect safety, legal rights, privacy, fairness, accessibility, trust, finances, or an individual outcome? | High or potentially harmful | Lived experience, domain evidence, and specialist judgment | Escalate | Involve affected people and the relevant legal, privacy, safety, accessibility, or domain specialist before launch |
Screen when the task is low-cost, reversible hypothesis generation. The output label should remain “hypothesis to inspect,” even when several generated profiles agree.
Pair when the decision matters but existing observed evidence can constrain interpretation. Pair simulation with research findings, analytics, support data, or domain evidence. If generated output conflicts with observed evidence, the observed evidence wins.
Escalate when the question concerns actual needs, comprehension, usability, willingness to pay, behavior, lived experience, or a population estimate. Real participants or observed behavior are mandatory in those cases. Specialist review is also mandatory for material safety, legal, privacy, fairness, accessibility, trust, financial, or individual-level consequences.
A launch or no-launch decision should never rest on generated reactions alone. Simulation can identify the next question; it cannot serve as the sole answer when customers bear the consequence.
Limits and validity
There is no universal “synthetic user accuracy” percentage. Validity depends on the task, population, grounding, model, prompt, comparator, and evaluation design.
A preprint by Park and colleagues used GPT-4o to build agents for 1,052 US adults from two-hour interviews, structured surveys, or both. Against the same people’s answers and two-week retest consistency, the agents completed held-out General Social Survey items, Big Five measures, economic games, and experimental tasks. On held-out survey items, grounded agents reached 82–86% of retest consistency versus 74% for demographics-only agents; differences among agent types on the economic games were not statistically significant. This does not validate simulated product feedback or individual predictions beyond those tasks (Park et al.).
Bisbee and colleagues prompted ChatGPT-3.5 Turbo with personas from 2016 and 2020 American National Election Study respondents, then compared feeling-thermometer responses about 11 sociopolitical groups with matched human data. Averages were sometimes close, but generated responses varied less, regression results often differed, and distributions changed with minor wording and over three months. This covers US political-attitude surveys, not product research (Bisbee et al.).
Santurkar and colleagues compared nine OpenAI and AI21 Labs models with human distributions across 1,498 questions from 15 Pew surveys and 60 US demographic groups. They found substantial misalignment—including for adults over 65 and widowed people—and demographic steering did not resolve it. The study was English-language, US-centric, multiple-choice, and not about product use (Santurkar et al.).
Sharma and colleagues found sycophancy across five AI assistants and four free-form tasks: responses could match a user’s stated view over a more truthful answer. Although not a synthetic-user validation study, it shows why agreeable output is model behavior, not customer consensus (Anthropic).
Output red flags
- Inputs or profile assumptions cannot be traced.
- An accuracy claim has no task, comparator, population, and evaluation design.
- Generated quotes are presented as customer evidence.
- Every profile agrees or uncertainty disappears.
- Minority populations or lived constraints are inferred rather than included.
- There is no independent validation plan.
- A high-stakes or individual decision rests on generation alone.
Fluent language can make weak evidence look finished. It does not give a model lived experience, make generated responses independent observations, or justify extending a study result to a new product decision.
Synthetic users versus adjacent methods
The word “AI” can describe who creates the evidence, who moderates its collection, or who helps analyze it. Those are different roles.
| Term | What it represents | Real participant? | Output | What it can establish |
|---|---|---|---|---|
| Synthetic user | An AI-generated profile that produces responses conditioned on instructions and context (NN/g) | No | Simulated reactions, themes, objections, or scenarios | Hypotheses and challenge coverage—not customer evidence |
| Persona | A research-based representation of a user group’s shared needs or behavior (GOV.UK) | No; the persona is an artifact grounded in prior research | A stable reference for design and product work | A synthesis of evidence already collected, not a new finding |
| Synthetic data | Artificial records generated to retain selected statistical characteristics of seed data (NIST) | No | A dataset | Fitness for a defined analytical or testing use after validation—not attitudes or lived experience |
| AI-moderated interview | An AI interviewer asks and adapts questions while a real person responds (User Interviews) | Yes | A real participant’s transcript or recording | Attitudinal evidence within the recruitment and study design; not behavior unless behavior is observed |
| Real-user research | Interviews, observation, usability tests, surveys, or other methods involving actual or likely users (GOV.UK) | Yes | Observations, statements, task results, or measurements | Needs, comprehension, behavior, or attitudes within the method and sample’s limits |
“AI-assisted research” is therefore only an umbrella. Always state whether a model generated the response, moderated a real participant, summarized evidence, or assisted analysis. A polished transcript is not enough to identify its evidentiary status.
A responsible pre-ship workflow
Use synthetic users as one bounded step inside a documented product-decision process. For a broader method-selection sequence, see how to test product changes before launch.
Define the decision and consequence. Name the proposed change, affected users, unknown, blast radius, and reversibility. Predefine what generated output may change: a research question, prototype, message, or risk register—not a launch decision by itself.
Document the inputs and gaps. Record each profile input, source, date, population, and owner. Check provenance, freshness, consent, privacy, and permitted model use. Explicitly list missing groups, behaviors, and contexts.
Generate competing hypotheses. Ask for objections, counterarguments, failure modes, and conditions under which a reaction would differ. Avoid asking only whether the proposal is good. Do not count multiple generated responses as a sample.
Review the output. Check traceability, prompt echo, disagreement, and stability across controlled reruns. Compare each material claim with research, analytics, support data, or domain evidence. Keep generated output in a separate evidence layer from findings involving real participants.
Choose and own the next step. Revise, run research, prototype-test, experiment, stage the rollout, or stop. Assign a validation owner and an escalation condition, such as an unresolved accessibility risk or a contradiction with observed behavior.
Operator note from the author: Malte Hedderich is an experienced AI engineer with a multi-year record shipping production AI systems at enterprise scale and working day to day with PMs in large B2C setups. His operator contribution here is an engineering control: log the input pack, prompt, model and version, run conditions, reviewer, validation owner, and escalation rule. Malte is also the founder of genjury and therefore has a commercial interest in the synthetic-user category.
The workflow is complete only when the team can show which statements were generated, which were observed, who will validate the unresolved risk, and what evidence can stop or change the release.
How to judge a synthetic-user output
Before moving any generated output into a product discussion, ask six questions. This review can also serve as an early gate when testing product changes before A/B testing.
Are the inputs traceable? Can a reviewer find the source, collection date, covered population, and owner for every material assumption?
Is the output context-specific? Does it respond to the actual change, affected group, and scenario, or merely produce generic product advice?
Is it new, or a prompt echo? Does the output combine documented evidence into a useful hypothesis, or simply restate the framing and adjectives supplied?
Does it show disagreement? Are counterarguments and conditions visible, and do controlled reruns reveal unstable themes that should not be treated as dependable?
What observed evidence supports or conflicts with it? Compare the claim with interviews, analytics, support records, usability findings, experiments, or domain evidence. Preserve conflicts.
What consequence dictates validation? Match the next method to the harm of being wrong. Higher consequence, actual behavior, lived experience, or population claims require escalation.
A strong-looking output can still fail this review. If its provenance, disagreement, or validation path is unclear, treat it as a prompt artifact—not a decision input.
Questions product teams ask
Can synthetic users replace customer interviews?
No. Customer interviews collect accounts from real people; synthetic users generate possible accounts from supplied context and model training. Use simulation to prepare competing interview hypotheses or improve a discussion guide. When a decision depends on actual needs, motivations, constraints, or experiences, recruit relevant people and follow the Screen / Pair / Escalate framework.
How accurate are synthetic users?
There is no defensible category-wide accuracy rate. Results change with the task, population, input depth, model, prompt, run, comparator, and definition of accuracy. Review the studies and boundaries in Limits and validity, then validate the exact output against evidence suited to your decision. Do not import a benchmark from survey questions into pricing, usability, or retention claims.
When are real users mandatory?
Real users are mandatory when you need evidence about actual needs, comprehension, usability, willingness to pay, behavior, lived experience, or the effects on a particular population. Include affected people and specialists for material safety, legal, privacy, fairness, accessibility, trust, financial, or individual consequences. Use the responsible pre-ship workflow to assign the validation owner before launch.
How are synthetic users different from AI-moderated interviews?
In an AI-moderated interview, the moderator is automated but the participant and their responses are real. With a synthetic user, both the participant profile and its response are generated. Choose an AI-moderated interview when you need testimony from recruited people and the research design supports automated moderation; use simulation only when a generated hypothesis is enough. See the adjacent-method comparison.
Synthetic users belong early in a decision, as a screen—not at the end as a verdict. Use interviews for needs and motivations, usability testing for task performance, surveys for appropriately designed estimates, analytics and experiments for behavior, and specialists for high-stakes consequences.
genjury is a pre-launch customer-response simulator designed to help product teams screen possible reactions before exposing real users. It is currently pre-launch; if that bounded use fits your workflow, you can join the waitlist.