What are synthetic users? A plain-English guide (2026)
Tookii · July 29, 2026 · 9 min read

The short answer
A synthetic user is an AI-simulated research participant: a language model paired with a stored description of who a person is, used to predict how that kind of person would react to your product, your pricing or your landing page. The term entered the industry in 2023, when a startup called Synthetic Users offered user research "without the users". Technically they're generative agents, software that simulates human behavior using a language model augmented with memories that shape its responses. They are measurably good at some questions, measurably bad at others, and worthless for the thing user research exists to do. Which of those you get depends almost entirely on what the persona was built from.
Where the term came from
The category has a clear origin. In 2023 a startup named Synthetic Users launched with the claim that you could run user research without recruiting anyone, using large language models to stand in for participants. The idea spread quickly, partly because the pain is real. Recruiting is slow, panels are expensive, and most teams ship without validating anything. It also spread faster than the discipline needed to use it well, which accounts for most of what goes wrong in the category.
Academic work followed the commercial claim rather than the other way round, which is unusual and worth knowing. A good deal of the research now cited by vendors was published after the products were already selling.
How synthetic users actually work
Three components, and every tool in the category is some arrangement of them.
A language model. The engine. Usually GPT, Claude or Gemini, and almost never trained by the vendor.
A memory. A stored body of text describing who this person is, retrieved as needed to shape responses. This is where the tools differ most, and it is the component that decides whether the output is worth anything. The memory might be a one-line demographic label, or a full interview transcript, or your actual site and customer data.
A reflection step. A mechanism that synthesizes those memories into higher-level inferences, so the agent behaves consistently rather than answering each question from scratch.
That is the whole architecture. Notice what is not in it: any connection to a real person who is alive today and might disagree with the output.
Synthetic users vs. the terms they get confused with
These get used interchangeably in marketing and they mean different things.
| Term | What it models | Built from | Typical use |
|---|---|---|---|
| Synthetic user | A type of person, sampled across a population | Personas, transcripts, market data | Testing a product or message against a segment |
| Digital twin | One specific individual or system, mirrored | Data about that one entity, often live | Predicting one named person or asset |
| AI panel | Real humans, with AI doing recruitment or analysis | Actual recruited participants | Traditional research, accelerated |
| AI persona | A character description in a prompt | A written brief, often invented | Design artifact, not a measurement |
| Traditional research | Nothing. You observe people | Real humans in real sessions | Discovery, and anything unmapped |
The distinction that matters most commercially is the last two rows against the first. An AI persona is a writing device. A synthetic user is supposed to be a measurement, which means it can be wrong, which means it can be checked. If a vendor cannot tell you what their personas are checked against, you are buying the writing device at the price of the measurement. It is the same trap that catches teams asking ChatGPT directly whether their idea is any good.
What synthetic users are genuinely good at
Three areas, all of them documented.
Reproducing attitudes when grounded in real data. Agents built from two-hour interviews with over a thousand people reproduced those individuals' survey answers with 85% normalized accuracy, scored against how consistently the same people replicated their own answers two weeks later. That is the high-water mark for the category, and it required real interviews as the memory.
Well-mapped cognitive patterns. Anchoring, framing and loss aversion are stable enough that models reproduce the direction and significance of documented effects, which makes a synthetic cohort useful for catching a pricing table that anchors everyone to the wrong tier.
Coverage and speed. A cohort can run forty persona-decision combinations against tonight's build. A recruited round of five users cannot, at any price.
The fuller version of this argument, with the boundary drawn precisely, is in when synthetic testing actually beats real users.
Where synthetic users fail
This is the part vendor pages skip, and it is not a small footnote.
They are too agreeable. Default user simulators are "overly cooperative, perfectly consistent, and highly forthcoming with information," while real users withhold details, push back on wrong assumptions, use ambiguous language and vary in patience. Researchers call this the behavioral gap, and it means an agent can pass a test that a real user would fail you on. It is the same agreeableness problem that makes ChatGPT a bad judge of your work, sold in a different product. The size of it is measurable: in one human evaluation, annotators rated conversations from a purpose-built realistic simulator as human 80.4% of the time, against 46.5% for the default one.
They overstate effects. A replication of 156 published psychology and management experiments through current models found effect sizes inflated two to three times, with reliability worst on race, gender and ethics.
They break on socially strategic behavior. Attitudes replicate. Negotiation under conflicting incentives does not.
They can't do discovery. This is the hard limit. Every question a synthetic user answers is one you already knew to type, and the finding that changes a roadmap is almost never in that set. A human hands it to you unprompted. No tool in this category solves that, including ours.
We catalogued the specific failure modes in four ways synthetic users lie to you.
Are synthetic users accurate?
The honest answer is that "accurate" is not a property they have. Accuracy is per-question.
The same study that produced the 85% figure also found those agents' correlation with their own humans dropping to roughly zero on a public-goods game. Same agents, same people, different question. So a vendor quoting a single accuracy number is telling you almost nothing, because the number is an average across tasks that behave completely differently.
There is also no independent benchmark in this category. Not one vendor's number has been reproduced by a third party, ours included. The only outside comparison anyone can cite is NN/g's 2024 test, which predates every current tool.
The practical consequence: treat every accuracy claim as a marketing artifact until you have tested it yourself. We wrote the longer version of this in how would you know if your synthetic users are right.
How to evaluate a synthetic user tool
One test, and it costs you an afternoon.
Take ten decisions where you already know the human answer. Things you have already tested, shipped, or watched people fail at. Run them through the tool before you subscribe. Count how many it gets right.
This works because it sidesteps every vendor claim and measures the only thing that matters: whether this tool, on your product, in your market, tracks reality. A tool that is genuinely good will help you set it up. A tool that is not will steer you toward a demo instead.
If you are comparing specific products, we ranked the category in the best synthetic user testing tools in 2026.
What separates a good synthetic user from a bad one
One variable dominates: what the persona was built from.
Agents grounded in real interview transcripts predicted individuals' responses 14 to 15 percentage points more accurately than demographic or persona-prompted agents running on the same underlying models. Same engine, different memory, materially different accuracy. And stuffing a persona with attributes irrelevant to the task drops performance by almost 30 percentage points, so more detail is not automatically better detail.
This is why "act as a 34-year-old product manager" produces such convincing nonsense. It is a writing prompt wearing the costume of a measurement. We wrote about what grounding actually buys you after a trial lawyer paid six figures for it, and about what the disciplined version of this category looks like once every failure above is treated as a choice you can reverse.
FAQ
What are synthetic users? AI-simulated research participants. A language model is paired with a stored memory describing who a person is, and that combination is used to predict how such a person would respond to your product, message or price. They are also called generative agents in the academic literature.
Are synthetic users accurate? It depends entirely on the question. Grounded agents reproduced real people's survey answers at 85% normalized accuracy, while the same agents dropped to near-zero correlation on strategic economic games. No independent benchmark exists for any commercial tool in this category, so treat vendor numbers as unverified.
What is the difference between synthetic users and digital twins? A digital twin models one specific individual or system and mirrors it, often with live data. A synthetic user models a type of person sampled from a population. The practical difference is scale: twins are built one at a time against a real counterpart, synthetic users are generated across a segment.
Can synthetic users replace user interviews? No, and this is the clearest limit in the category. An interview can go somewhere you didn't plan for it to go, because the person on the other side has their own agenda. A simulation stays inside the question you wrote.
Are synthetic users the same as AI personas? Not quite. An AI persona is usually a character description written into a prompt, which is a design artifact. A synthetic user is meant to be a measurement checked against something real. Many products sell the first while describing the second.
What do synthetic user tools cost? Published entry pricing in 2026 ranges from free tiers to roughly $50 a month for developer-oriented tools, $47 to $79 a month for budget interview simulators, and $2 to $60 per interview for enterprise products. Several vendors publish no pricing at all.
What can synthetic users not do? Discovery, socially strategic behavior, and anything culturally contingent. They also skew agreeable, which means they under-report friction a real user would hit.
How do I know if a synthetic user tool is any good? Run ten decisions you already know the human answer to and score the tool before you buy. No vendor's published accuracy figure substitutes for this, because none of them have been independently reproduced.
The one-line version
Synthetic users are a narrowing instrument. Sold as a replacement for research, they fail. Used as the step before it, they're the cheapest way to stop shipping guesses.
Free during beta. No demo call, no credit card.