Why you can't validate an idea by asking ChatGPT
Tookii · June 8, 2026 · 7 min read

You pasted your idea into ChatGPT and asked it the hard questions. Size this market. Poke holes in the model. Be brutally honest.
And it was. The answer came back measured, structured, a little critical even. A couple of real risks named, a competitor you hadn't thought of. It read like analysis. So you trusted it.
That's the trap. Not the model being too nice and telling you what you want to hear. That happens, but you can usually feel the flattery coming. This is the opposite: you ask a hard question, you get back something that looks exactly like rigor, and you have no way to tell whether it is.
Being right, for a language model, isn't a switch that's either on or off. It's something it does per task: sometimes, on some questions, and not on others, with no tell in the output for which one you just got.
The instrument that works until it doesn't
Here's who gets burned. Not the founder who asked ChatGPT to cheerlead. The one who asked it to think — got back a clean, skeptical, professional-looking answer with no smiley edges on it, and bet the roadmap on the one decision the model happened to whiff.
There's a whole cottage industry betting you won't notice: validate your idea in 120 seconds, the 30-minute startup test, paste your pitch and get a verdict. And the models are good enough, often enough, that it feels like it's working. One reviewer ran three ideas through ChatGPT; it told him to build all three. Including one where it had just scored the competition a 9 out of 10, an almost un-enterable market. It saw the wall. It said climb it anyway. The analysis and the conclusion came from the same fluent paragraph and never spoke to each other.
The danger was never that the model is unreliable. It's that it's unreliably reliable: right enough, often enough, that you quietly stop checking.
Right on Tuesday, wrong on Wednesday
Validity is a property of the task, not the model. This is the whole post, so it's worth proving instead of asserting.
The cleanest demonstration is a Turing test built out of economics. Researchers sat GPT-4 down to play six canonical economic games and compared it against a human dataset spanning more than 50 countries and over 108,000 people. The model did well. In several of those games it lands inside the human range. It passes a behavioral Turing test, indistinguishable from a person by its choices. Then it sits down to play the Prisoner's Dilemma, and the resemblance falls apart; it diverges most sharply there, and again as the investor in the Trust Game.
Read that again, because the structure of it is the point. Same model. Same fluency. Six tasks, and the fidelity is high on most and broken on a specific few, not randomly but as a function of which game it's playing. Nothing about the model "being good at economic games" survived the move from one game to the next.
So the question "is GPT-4 a good stand-in for a human?" has no answer. It has six answers, and they disagree.
It doesn't transfer
The natural defense is to lean on a track record. It nailed the last three calls. That's a reason to trust the fourth. It isn't. Validity doesn't accumulate across tasks the way a reputation does.
When researchers mapped where LLM simulations hold and where they snap, they found "systematic boundary conditions": not noise, but a hard edge that moves task to task. The same agents that successfully replicated human behavior in ultimatum games and in the Milgram obedience experiments failed to reproduce the Wisdom of Crowds. Three for four, and the fourth wasn't a near miss. It was a different kind of question the model had no claim on.
This is why "the model is good at understanding people" is a sentence that means nothing. There's no general competence to point at. GPT-4 beats humans at detecting irony and at analogical reasoning, and underperforms them elsewhere in the same breath. The only true statement available is the narrow one: it matched humans on this exact task, this exact way, this once. Everything broader than that is a guess wearing the last result's clothes.
Why you can't see it coming
Here's the part that turns a limitation into a trap. When the model whiffs, the answer doesn't look any different.
A wrong answer and a right one arrive in the same confident prose, the same clean structure, the same tone of having-thought-about-it. There's no tremor in the output when fidelity drops. And you only know it drops at all because someone went and measured it. In one study, researchers ran LLM-simulated users against real ones and found the simulation tracked humans decently on some slices of the work and was badly off on others: a 45.2% overall success rate, with calibration error that ran unevenly across task difficulty rather than as a flat, predictable discount. Nothing in the output flagged the slices where it was wrong. It can't. It doesn't know.
That's the whole problem in a sentence: the failure is invisible at exactly the moment you'd need to catch it. You find out the model was wrong about your users the way you always find out: when real users show up and disagree, after you've built the thing. The instrument has no error bars, so you supply them yourself, in production, at the worst exchange rate available.
Where the line actually falls
None of this makes the model useless, and being precise about the boundary is the difference between skepticism and cynicism.
Ask it something with ground truth sitting in the prompt (does this code run, is this clause contradictory, what breaks if this number doubles) and it's on solid ground; the answer is checkable and the checking is cheap. The trap is the other kind of question: what will real people do, where the ground truth lives out in the world, in humans you haven't talked to yet.
For that second kind, there's a discipline, and it's not "trust the model" or "don't." The mistake was never using a model to help you think. It was handing a general-purpose model your decision and calling its answer validation. Validation is something you measure, not something you ask for. So you stop asking is the model valid (an unanswerable question) and start asking is it valid for this specific decision, and what's the evidence. That means checking it task by task, against real human data, for the exact use you're about to make of it, which is exactly what researchers who do this seriously already do. They don't grant the model general credibility; they earn confidence one task at a time, and they say so out loud. The discipline is real. Most people just skip it, because skipping it feels identical. Right up until it doesn't.
How you actually run that check (what "validate against real humans" looks like in practice) is its own subject. We'll get there.
So
A detailed answer feels like validation. It's not. It's a language model doing what it does best: generating plausible, fluent, confident text. And it reads exactly the same whether it nailed your market or invented it.
The better it sounds, the worse this gets. The more thorough the answer, the harder it is to walk away from a bad idea.
So "is the model right about my idea" was never the question. It can't be right or wrong in general, only about one specific call, and the output won't tell you which one you're holding.
Valid for what — and who checked?