Tookii
Let's talk
Validating ideas

How would you know if your synthetic users are right?

Tookii · June 21, 2026 · 4 min read

How would you know if your synthetic users are right?

Someone shows you a synthetic user that sounds exactly like your customer. One question decides whether it's worth anything: how would you know if it's right?

Not does it sound right. It will. Sounding right is the one thing these systems are guaranteed to do. Right is a different claim, and it's the one you're actually betting the roadmap on.

Here's the trap most teams fall into. They can't answer the question, so they reach for a number someone handed them. Vendors are happy to supply one: you'll see persona accuracy quoted as a correlation as high as 95%. It sounds like proof. It isn't. Not because it's faked, but because a single headline number is the easiest thing to put on a slide and the least informative thing about whether you can trust these personas for your decision. Right isn't an adjective you attach with one statistic. It's a measurement, and it takes three checks.

1. Check it against real humans

There's no internal test for "right." A synthetic user can't certify itself; the word only means something relative to real people doing the real task. The research is blunt about it: you validate by replicating a slice of the work with actual humans and comparing the outputs directly. Where you can't, you're not measuring validity. You're asserting it.

The unglamorous part is that you need real signal in the loop: a handful of interviews, support tickets, NPS verbatims, a small panel. Something human to check against. Even the practitioners most bullish on synthetic users agree the only thing that validates one is real human research. No human anywhere in the loop isn't confidence. It's vibes with a progress bar.

2. Compare the distribution, not the average

This is where that 95% hides its sins. A model can nail the average and flatten everything around it. And the spread is usually the point. Real people disagree; they cluster, they throw outliers, they contain the strange 8% who churn for reasons the median never sees. LLMs are notoriously thin here: their responses show far lower variance than real interpersonal variance, and they routinely fail to match the critical moments of a human distribution: variance, quantiles, skew. Serious validation compares the whole shape against humans, not just the center.

So a correlation can read 0.95 and still miss every person who matters. Matching the mean is the trick that looks like accuracy. Matching the distribution is accuracy.

Two ways to "match" humans: two response curves with the same average, but one collapsed into a tall narrow spike (the synthetic users) and one spread wide across the full range (real people). Same mean, different people.

3. Name the boundary

A validated synthetic user is valid for a decision and a moment. Not in general, and not forever. Validity doesn't travel across tasks: the same model can sit inside the human range on five economic games and break on the sixth. And it doesn't travel across time: models are trained on the past, so previous alignment doesn't guarantee current applicability: they miss shifts in attitude, language, and what people have started to care about.

So the honest output of validation isn't "these personas are accurate." It's "these personas track real people on this question, right now. And here's where they stop." Confidence you can't draw a boundary around isn't confidence. It's a guess that hasn't failed yet.

What this actually costs

Three checks is more work than a dashboard wants you to believe, and none of them are exotic: they're just the difference between a tool you trust and a tool that flatters you. A number on a slide is not the same as confidence. Confidence is knowing which of the three you ran, and what each one told you.

That's the bar. It's higher than "sounds like my user," and it's the only one that protects the decision underneath.

Three questions, not one

The next time something sounds exactly like your customer, you've got three questions instead of one. Did you check it against a real person? On the spread, not the average? And for this decision, right now?

A synthetic user is never right just because it's convincing. Convincing is free. Right is what's left after you've checked. Skip that, and the most lifelike persona on your screen is just a confident guess you've chosen to believe.

Run your first test.