Tookii
Let's talk
Synthetic users

Everyone's simulating their users. Most are doing it wrong.

Tookii · June 13, 2026 · 5 min read

Everyone's simulating their users. Most are doing it wrong.

Imagine a focus group that never sleeps, costs nothing, and answers in ten seconds.

That's the pitch. It's a good pitch. Good enough that a lot of sharp teams walk straight into the catch sitting behind it.

The appeal is honest. Real user research is slow, expensive, and a nightmare to schedule, so a tool that spins up your users on demand (answering surveys, sitting interviews, reacting to a mockup, all before lunch) is hard to refuse. In the past year, a wave of new tools has arrived selling the same promise as a feature: user research "without the users." Synthetic participants for surveys, interviews, and concept tests, generated on demand. The want is real, and it's rational.

The pain is real too, and I've felt it. Recruiting once through a well-known research panel, I got matched with a "Silicon Valley CTO" who joined the call from what looked like a clay hut on the edge of a savanna. Nice guy. Not a CTO. The panel paid per finished session, so somewhere down the chain, someone had every reason to be whoever the screener was looking for.

So I started checking. I'd ask what the weather was doing where they were, or what the local time was, small things you can't fake on the spot. More often than I'd like, the answer didn't match the profile. This isn't just my bad luck: participant fraud is a well-documented problem in research panels. Pay people to be a certain kind of user and some of them will simply become that user for half an hour. That's what the synthetic pitch sells against: skip the recruiting, skip the theater, get your users on demand.

Then you make a call based on what they told you, and you meet the catch.

The catch nobody puts on the landing page

Your synthetic users are too agreeable to be real.

Real people are messy in ways that matter. They meet your product at the end of a long day, before their coffee, mid-argument, on a train with one bar of signal. A moment you don't control and never designed for. They're tired, distracted, impatient. They have habits, baggage, and somewhere better to be. They get confused and quietly blame themselves; they get annoyed and loudly blame you. They withhold the thing you needed until you ask three times, push back on your premise, or just go quiet and close the tab.

A synthetic user is none of those things. It's never tired, never distracted, never having a worse day than yesterday. It's "overly cooperative, perfectly consistent, and highly forthcoming," so where a real person hesitates, gives up, or pushes back, the simulated one helpfully hands you the answer you were hoping for. The research has a name for the distance between the two: the behavioral gap. And it isn't cosmetic. A product can score beautifully against the cooperative version and come apart the moment a frustrated, skeptical, one-word-answer human shows up.

The behavioral gap: a default synthetic user is forthcoming, consistent, patient, and agreeable; a real user withholds, pushes back, loses patience, and goes quiet.

You can catch it red-handed. When Nielsen Norman Group tested synthetic users against three of their own real studies and asked whether they'd finished every course they'd signed up for, the synthetic users cheerfully said yes: the flattering answer, not the human one. Their verdict across many research tasks was blunt: the responses are too shallow to be useful. Independent work finds the same fingerprint: simulated users come out conspicuously more polite and more eager than the humans they stand in for, and systems tuned against them "appear robust in benchmarks while failing disproportionately for real users".

The second catch is what makes the first one dangerous: it all reads beautifully.

The output is fluent, structured, confident. It looks exactly like research. That's the trap. Convincing and accurate are different things, and on a screen of clean prose they're nearly impossible to tell apart. One study that interviewed AI-generated personas found they "appear credible" and then, under any real probing, reveal "a notable lack of clear-cut opinions": plausible on the surface, hollow underneath. Sounding like a user was never the same as behaving like one.

The third catch is quieter, and we'll come back to it later in this series: a naive simulation gives you an average, not a crowd. Ask a hundred real people and you get a spread: disagreement, outliers, the strange 8% who turn out to matter. Ask the model a hundred times and it collapses toward one bland median; in one study LLMs ran flatter than humans on nearly every measure, covering just three of five answer options across a hundred tries. A market isn't a median. The people who break your assumptions live in the tails the model quietly sands off.

None of this means it doesn't work

Here's where most takes stop: "use real users instead," or the gentler "supplement, don't replace." Both are correct, and both quit right before the useful part. They tell you to be careful without telling you what careful actually looks like.

The catch is a craftsmanship problem, not a dead end. The same research that documents the behavioral gap also points at the way out: drop the off-the-shelf simulator that agrees with everything. Ground the personas in the real distribution of human behavior, build the friction back in (the impatience, the skepticism, the brevity) and check the result against actual people before you trust it. Done that way, the field's own verdict isn't "give up." It's that simulation is a genuinely promising research method, on two conditions the cheap version skips: diversity and verification. Treat it as something you build and check, not something you buy and believe.

That's the entire difference between a toy and an instrument. Not whether it sounds human. Whether someone did the work to make it behave like one. And then checked.

So

Synthetic users aren't the problem. The agreeable, ungrounded, straight-out-of-the-box kind are. And they're the kind most people are running.

So the question was never whether you can simulate your users. You can. It's whether the ones you built will ever tell you something you didn't want to hear.

Run your first test.