Tookii
Let's talk
Synthetic users

AI user testing vs. real user testing: when synthetic users actually win

Tookii · July 1, 2026 · 11 min read

AI user testing vs. real user testing: when synthetic users actually win

Nobody argues a telescope beats a microscope.

Yet that's the shape of the entire AI-versus-real user testing debate. Two instruments, built for different questions, argued over as if one is about to replace the other. The vendors say recruiting is obsolete. The skeptics say AI feedback is fiction dressed as research. Both are answering the wrong question, which is a shame, because the right one has an unusually clean answer.

Get the instruments confused and it costs you in either direction. One team bets a quarter's roadmap on a synthetic panel's enthusiasm for a product no human asked for. Another spends six weeks and a five-figure recruiting budget to learn their pricing page confuses people, a finding a synthetic cohort would have handed them before lunch. Opposite failures, same root: somebody asked one instrument the other's question.

The short answer

Synthetic and real user testing answer different questions, so the useful comparison is which one your question needs. Synthetic testing narrows. It's measurably good at first impressions, at well-documented cognitive biases, and at coverage on a timescale recruiting can't match, which makes it right whenever you can already name the thing you're checking. Real users widen. Only a person volunteers the problem you never thought to ask about, and no simulation substitutes for that. The expensive mistake runs in both directions: betting a roadmap on a synthetic panel's enthusiasm, or spending six weeks of recruiting budget to learn that your pricing page confuses people. The sequence that works is synthetic first to kill the weak hypotheses, then real people on whatever survives.

DimensionSynthetic testingReal users
TurnaroundHoursWeeks
Practical sample sizeTens to hundredsAround five per round
Cost per roundSoftware pricingFour to five figures
Behavioural signalYes, when it drives the real productYes
Discovery of the unasked questionNoThe only source
Breaks down onSocially strategic and culturally contingent behaviorBudget, scheduling and access

The wrong question about accuracy

Accurate at what is the question, and the research is unusually specific about it.

When Stanford researchers built agents from two-hour interviews with 1,052 people, the agents reproduced their humans' General Social Survey answers with 85% normalized accuracy, scored against how well the same people replicated their own answers two weeks later. A mega-replication ran 156 published psychology and management experiments through current models: main effects reproduced in 73–81% of them. And in six canonical economic games, GPT-4 tracked a human dataset of over 108,000 people closely enough to pass a behavioral Turing test in several of them.

Most of that evidence postdates the piece everyone still cites against it. When the Nielsen Norman Group (NN/g, the closest thing UX research has to a referee) tested synthetic users in 2024 and concluded they were fit for desk research but not much else, they were testing text interviews with 2024-era models. Fair verdict on dated evidence. The tools have moved since. Some of the criticism still lands, and we'll get to it, but the blanket version hasn't survived the data.

The same models break on cue

The pricing pages skip the next part: the same studies draw the failure map.

GPT-4 sits inside the human range across most of those economic games, then diverges exactly where the game turns socially strategic: the Prisoner's Dilemma, and playing investor in the Trust Game. Stanford's interview-grounded agents, the 85% ones, crater on the same terrain. On the Public Goods game, their correlation with their own humans drops to roughly zero. The 156-experiment replication found effect sizes inflated two to three times, with reliability falling hardest on race, gender, and ethics. Persona-prompted models flatten the variance real groups actually have. And they shift opinion toward whatever framing the asker brings, the agreeableness problem we've written about before.

The failures cluster. Fast, well-mapped, near-universal cognition replicates. Socially strategic, culturally contingent behavior doesn't, and neither does anything that depends on the model volunteering what you didn't ask. (The full failure catalog is in Four ways synthetic users lie to you.)

An instrument that fails predictably isn't a broken instrument. It's a scoped one.

When synthetic users actually win

Start with what changed. Shipping got cheap. AI writes a bigger share of the code every quarter, prototypes that took a sprint now take an afternoon, and the backlog fills faster than any team can validate it. The expensive part of product work moved upstream, from can we build it to which of these twelve should we build, and why. That's a testing question, and a recruiting pipeline that delivers five opinions in three weeks was never built for its cadence.

This is the ground synthetic users were built for. Three capabilities, with receipts:

First impressions. A visitor forms a judgment of your page in about 50 milliseconds, and that snap judgment propagates: visual appeal drives perceived quality, which drives trust. The inconvenient part for human testing: nobody can articulate a 50-millisecond judgment. Ask a participant why the page felt off and you get a story their brain assembled after the fact. But the judgment itself is systematic. Systematic enough that models trained on UI features already predict where attention lands better than classic saliency baselines, validated against eye-tracking. Predictable perception is exactly what a model can be measured against, at a resolution self-report will never reach.

Known bias patterns. Anchoring, framing, loss aversion: decades of research treat these as stable regularities you can design around, and models reproduce the direction and significance of the documented effects. A synthetic cohort built to carry those biases can walk your checkout and flag that the plan-comparison table anchors everyone to the wrong tier, a mechanism a real participant experiences but can't name. Real users feel the bias. A well-built synthetic user points at it.

Coverage at speed. Human testing economics top out fast; the classic doctrine caps a round at five users, and even its critics won't get you to fifty. Persona populations built with guided search now reach behavioral coverage just short of a human reference set (0.602 against a 0.614 human benchmark in one retail study), which means forty persona-decision combinations against tonight's build, with a distribution instead of an anecdote by morning.

And those three compound into the jobs a roadmap actually needs done:

  • De-risking a decision before it costs a sprint. Two directions on the table: run both through the cohort tonight and let the weak one die before anyone builds it.
  • Prioritizing with a distribution instead of a debate. Five candidate fixes, forty persona runs each. Rank them by measured friction, not by whoever argued loudest in planning.
  • Testing when you have no users to recruit. Pre-launch, new market, new segment: you can't panel users who don't exist yet. A cohort grounded in your market signal is how the first hundred decisions get tested at all.
  • Market research before the build. Point the cohort at your positioning, your pricing page, your competitors'. The same instrument reads the market as well as the prototype.

None of these say "replace your user research."

When real users win, every time

Discovery. The question you didn't know to ask.

Here's what that looks like when it bites. A team building LinkedIn's AI features spent weeks polishing generated summaries: the length, the tone, the format. Then they watched real people use the product, and the feedback had nothing to do with any of it. Users skipped the summaries and asked for action items, who does what and by when. Weeks of tuning, aimed at a target nobody wanted hit.

No synthetic user could have caught that. Every question the team knew to ask (is the summary too long? is the tone right?) was a question about summaries. The answer lived outside all of them. A model answers what you put in front of it. A human volunteers what you never thought to raise. That single behavior, volunteering the unasked, is what real research is for, and no simulation substitutes for it. The gap was never in the users. It was in the questions.

So the rule is plain. When you can already name the thing you're checking, synthetic answers fast, at scale, tonight. When the problem is that you can't name it yet, only a person will hand it to you.

Which question are you asking?

Synthetic-shaped questionsHuman-shaped questions
Do people understand what this page is?Why are people churning after week two?
Which of these two variants reads clearer?What problem are they actually hiring us for?
Where does attention land in the first five seconds?What would make them switch from the incumbent?
Does the pricing table anchor users to the wrong plan?What does this feel like the first time, for real?
Which of these five roadmap bets is worth building first?What almost made them not sign up?
How does our offer read next to the two incumbents?What are they doing the moment before they open the app?
Which onboarding change gives week-two retention the best shot?What didn't we think to ask?

Left column: answers by tonight, at whatever scale the decision needs. Right column: no model will give you an answer you should trust, so go find a human. The columns also trade in opposite directions. The left one keeps narrowing your bets; the right one keeps widening what you know to bet on.

The fine print

Two honest complications.

First, the split leaks a little. There's a fair case for simulation in exploratory mode, early pilots where surfacing possibilities matters more than avoiding false positives. "Simulate broadly, validate the winners with humans" is a legitimate workflow in its own right. Synthetic discovery exists, as a hypothesis generator that never returns a verdict.

Second, and no pricing page will tell you this, every "synthetic wins" result above comes from measured, grounded setups. Interview-grounded agents beat demographic-prompt agents by 14–15 points. Rich backstories improve match to human response distributions by up to 18%. Stuffing a persona with irrelevant attributes drops task performance by almost 30 points. "Act as a 34-year-old product manager" gets you neither instrument: not the human's unpredictability, not the model's measurable fidelity. That gap between a prompt and an instrument is why Tookii builds personas from your real market signal and treats accuracy as something you measure per decision, not something you claim in a tagline.

Synthetic first. Humans for what's left.

The two instruments belong in sequence.

Run the synthetic cohort before the sprint, before the recruiter, before the debate hardens into a roadmap. Twenty hypotheses go in; the grounded cohort retires the weak seventeen: the variant nobody parses, the tier nobody picks, the positioning that loses to the incumbent on first read. That's the instrument's real job — a hypothesis reducer that narrows the field until what's left has earned human attention.

Then spend the recruiting budget on the survivors. Five real people, every session pointed at the questions no model can volunteer answers to.

Run it in that order and the instruments sharpen each other: the synthetic pass arrives grounded in your market, and the human sessions stop burning their first twenty minutes on questions a cohort could have retired overnight. Skip the order and you're back at the top of this post, betting a roadmap on an AI panel's enthusiasm, or paying five figures to relearn what a model already measured.

Most of what's on your roadmap this quarter has never been tested against anything. That's tonight's job, and it finishes before sprint planning turns those guesses into commitments.

FAQ

Is AI user testing accurate? For some questions and not others, and the boundary is now reasonably well mapped. Grounded agents reproduce attitudes and documented cognitive patterns at useful accuracy, and they break on socially strategic behavior. No commercial tool in the category has independent third-party validation, so test any vendor's claim on decisions where you already know the human answer.

Can AI replace user testing? Not the discovery part, which is what user research exists for. What AI replaces is the narrowing work that happens before you recruit: the twenty candidate answers you'd otherwise burn sessions eliminating one at a time.

When should I still recruit real users? Whenever you can't name what you're looking for: churn you don't understand, a market you haven't served, or the first run of an unfamiliar flow. Also whenever the decision is large enough that being wrong costs more than a panel does.

What does AI user testing cost compared with a panel? Published entry pricing across the category runs from free tiers to roughly $50 a month for developer tools, and $2 to $60 per interview at the enterprise end. A traditional round of moderated sessions is typically four to five figures. The comparison by tool is in the best synthetic user testing tools in 2026.

Run your first test.