Tookii
Let's talk
Building personas

The lawyer who beat Meta didn't use AI out of the box. Neither should your personas.

Tookii · July 12, 2026 · 8 min read

The lawyer who beat Meta didn't use AI out of the box. Neither should your personas.

Earlier this year, a Texas jury found Meta and YouTube negligent in the first social-media addiction case in the United States to reach a jury verdict. Mark Lanier tried it, and he has been open that AI ran through the whole thing.

Most of the coverage stopped at "lawyer uses AI, wins." The more useful detail sits a line further down. He worked with the platform to build a custom license, six figures a year, "tailored to incorporate his 42 years of trial experience into the AI's context."

The short answer

Lanier used the same AI models available to anyone: ChatGPT, Claude and Gemini, through a platform called Boodlebox. What he did differently was pay for a custom license that loaded 42 years of his own trial experience into the model's working context. That's the variable worth copying, and the academic literature has measured it directly. Agents grounded in real interview transcripts predicted people's survey answers 14 to 15 percentage points more accurately than demographic or persona-prompted agents running on the same underlying models. The model is rarely the variable. What you put in front of it's.

What the AI actually did

Four jobs, none of them glamorous:

JobWhat it looked like
Nightly reviewEach day's court transcripts fed to several models for an assessment of how the day had gone
PhrasingFinding sharper, more vivid ways to put an argument to a jury
Reading the roomFeeding the jury's written questions into the models during deliberations to gauge where the panel stood
Document reviewOvernight passes through discovery, ready before court the next morning

The reported efficiency gain was roughly three to one, about thirty hours of work in ten.

Worth stating plainly, because it is the part that gets skipped: none of this simulated a juror. The twelve people who returned that verdict were twelve people. The AI helped a lawyer prepare for humans. It never stood in for them.

The part everyone skipped

A six-figure annual license is a strange line item when ChatGPT costs $20 a month and does, superficially, the same things. He has not said publicly why he judged it worth the money, so what follows is our reading rather than his. The mechanics are not mysterious.

An ungrounded model answers from the average of everything it has read. Ask it how to open a product-liability case and you get a competent, generic, faintly familiar answer, because it is reconstructing the center of mass of every trial-advocacy text ever written. That average is a reasonable place to start and a terrible place to finish. It is nobody's actual strategy. It is certainly not the strategy of a lawyer with four decades of specific, hard-won, idiosyncratic judgment about what works in front of a particular kind of jury.

Loading that judgment into the context changes what the model is doing. It stops generating plausible trial strategy and starts generating his trial strategy, evaluated against his standards. Same weights, same API, completely different instrument.

The research says the same thing, with numbers

This is one of the better-measured findings in the whole simulation literature, and it is unusually clean.

When Stanford researchers built agents from two-hour interviews with over a thousand people, those agents reproduced their humans' General Social Survey answers with 85% normalized accuracy, scored against how consistently the same people replicated their own answers two weeks later. Agents built on demographics or persona descriptions, using identical models with no access to the interviews, landed 14 to 15 points lower. Broken out, interview-grounded agents scored 0.83 against 0.71 for demographic agents and 0.75 for persona agents.

The failure direction is documented too, and it is worse than most people assume. Persona prompting does not reliably help at all: a systematic study of 162 personas found that adding one "does not necessarily improve an LLM's performance on objective tasks" and can hurt it, with the effect largely unpredictable. Stuffing a persona with attributes irrelevant to the task drops performance by almost 30 percentage points, and this holds even for the largest models.

So the picture is not "more detail is better." It is that real context helps, invented context is noise, and noise actively costs you.

Why this matters if you are testing a product, not a case

Swap the courtroom for a checkout flow and nothing about the mechanics changes.

The standard way teams use AI for user feedback is to type "act as a 34-year-old product manager evaluating this landing page" and read what comes back. That is the vanilla-ChatGPT version of what Lanier declined to do. The model has no access to your market, your customers, the objection your last twelve churned accounts all raised, or what your competitor's pricing page says two clicks away. It answers from the average of every product manager it has ever read about, which is to say from nobody.

The output looks fine. That is the trap. It reads articulate, it is confidently phrased, and it is unfalsifiable, because there is no real person it was ever supposed to match.

Ungrounded personaGrounded persona
Built fromA demographic label in a promptYour site, your data, your competitive context
Answers fromThe average of everything the model has readEvidence specific to your market
Measured againstNothingA held-out benchmark you can score
Documented accuracy gapBaseline+14 to 15 points on the same model
Failure modePlausible, generic, unfalsifiableWrong in ways you can detect

The last row is the one that matters commercially. An ungrounded persona cannot really be wrong, because it was never anchored to anything checkable. A grounded one can be measured, which means it can also fail, which is the only condition under which a result is worth anything.

What grounding actually requires

Three things, and Lanier's license had all of them.

Real source material. Not a description of your users, but artifacts they produced or interacted with: your live pages, your sales calls, your support tickets, your churn notes. Interviews worked in the Stanford study because they were transcripts of actual humans, not summaries of imagined ones.

Relevance discipline. Since irrelevant attributes cost up to 30 points, more context is not automatically better context. A persona carrying six facts that bear on the decision beats one carrying sixty that do not.

A measurement you can lose. Lanier had one built in: verdicts. Product teams usually have none, which is how "the personas seemed insightful" becomes a purchasing criterion. Pick ten decisions where you already know the human answer, run them, and score before you trust.

This is the design Tookii is built around, and the reason our personas start from your URL, your uploaded data and your competitive context rather than a dropdown of demographics. Accuracy gets measured per decision against a committed benchmark, and prompt changes that score worse than the version they replace do not ship. Whether that is worth your time is a question you should answer with the ten-decision test above, not with our say-so. It is the same test we recommend running on every tool in the category, us included.

FAQ

Does grounding an AI in your own data actually improve its output? Yes, and the effect is measurable rather than aesthetic. Agents grounded in real interview transcripts predicted individuals' survey responses 14 to 15 percentage points more accurately than demographic or persona-prompted agents running on the same underlying models. Same models, different context, measurably different accuracy.

Is a custom-context AI setup really better than just using ChatGPT? For general questions, no. For work that depends on judgment specific to you, your market or your customers, the gap is large enough that a trial lawyer with a landmark case on the line paid six figures a year to close it. Generic models answer from the average of everything they have read, which is a good starting point and a poor final answer.

What is the difference between persona prompting and grounding? Persona prompting assigns a label ("act as a 34-year-old PM"). Grounding supplies real evidence about the actual people and context involved. The distinction isn't academic: persona prompting alone doesn't reliably improve results and can degrade them unpredictably, while irrelevant persona attributes cost up to 30 percentage points of task performance.

Can AI predict how real people will react? Within limits that are now reasonably well mapped. Grounded agents reproduce attitudes and well-documented cognitive patterns at useful accuracy, and break down on socially strategic behavior and anything that depends on volunteering what you didn't ask. The fuller version is in our breakdown of when synthetic testing beats real users.

The transferable lesson

He did not win because he had better AI. He had the same models as everyone else, priced at $20 a month.

He won partly because he spent real money making those models answer from his context instead of the internet's average, and then checked the output against a standard that could actually fail him. That is available to anyone testing anything, and almost nobody does it, because ungrounded output looks good enough to skip the step.

Start testing free

Free during beta. No demo call, no credit card.

Run your first test.