Tookii
Let's talk
Guides

How to test your mobile game's FTUE: a 7-step guide (2026)

Tookii · September 23, 2026 · 10 min read

Say your tutorial completion rate is 85%. Your D1 is 22%. Both numbers are true. Only one of them is about your tutorial.

A forced tutorial gets finished. That's what forced means. The player taps the glowing finger, the arrow moves, the step counter ticks over. None of it tells you whether they understood a single thing they were shown. Understanding shows up later, at the first moment the game stops holding their hand — and by then most FTUE tests have already stopped watching.

This guide covers how to test the part that matters: where a new player stops understanding your game.

The short answer

To test a mobile game's first-time user experience (FTUE), write down what a new player must understand by the end of it, then track each tutorial step against that list. Test with people who have never played your game or its close clones. During the session, ask "what do you think you do now?" rather than "was that clear?". Sort every problem into one of three kinds: the player didn't see it, saw it but didn't get it, or got it and didn't care. Then re-test on every build, not once a quarter. Human players are the judges of fun. First-time testers, human or AI, are how you find where understanding breaks.

Why the first session carries so much weight

D1 retention is the closest number you have to a verdict on onboarding. GameAnalytics' 2026 benchmarks, built from more than 16,000 games, put the median D1 at about 22% and the top quarter at about 30% (GameAnalytics, 2026). Their reading of what a low D1 means is blunt: it "often points to confusing tutorials or unappealing initial gameplay" (GameAnalytics, 2025).

The losses start fast. A 2016 analysis on Game Developer put it at 20% of installs gone within two minutes of first launch (Game Developer, 2016). The number is a decade old. The shape hasn't changed: the first minutes are where you lose the most players for the least reason.

How to test your FTUE in 7 steps

1. Write down what "understood" means before you test

Before anyone plays, list the 5 to 8 things a new player must know when the FTUE ends. The goal of a level. The core loop. What the main currency is for. Which button actually matters. This list is your answer key, and without it a playtest produces opinions instead of findings.

The idea isn't new. Researchers testing tutorials with vision-language models used the same move: ask questions about each tutorial screen, compare the answers to what the developers expected a player to take away, and treat the gaps as the confusing scenes (Level Up Your Tutorials, 2024). Expected answers first. Then test.

2. Track the tutorial funnel, then distrust it

Fire an event at each tutorial step and build the funnel. You'll see where players drop out, which is useful. What the funnel can't show is comprehension. A forced step converts at nearly 100% whether the player understood it or not.

Two places to look instead:

  • The first free choice. The first moment after the tutorial where the player decides what to do. If they don't do the thing you just taught, the tutorial didn't land.
  • Legible but not understood. A goal that reads "Collect 12" is perfectly readable. Twelve of what, out of how many, by when? Every word on screen can be clear while the meaning isn't there.

3. Recruit people who have never played it

Your team can't test this. Neither can your community beta, mostly. The FTUE only happens once per person, so every tester you use for it is spent. Screen out anyone who has played your game or its close clones. Playtest vendors commonly recommend around 10 players per player profile for FTUE work, because onboarding problems surface quickly at small sample sizes (Antidote FTUE playbook).

Watch for one specific skew: experienced testers. Someone who has played 200 match-3 games will glide past a confusing step a real newcomer gets stuck on.

4. Ask "what do you think you do now?"

"Was that clear?" gets a yes. People are polite, and most don't know what they missed. Ask them to predict instead: what happens if you tap that, what's the goal of this level, what is that counter for. A wrong prediction is a finding. A confident wrong prediction is a better one.

Not all confusion is bad, either. A study of player confusion found that designers tend to treat it as a defect, while players can experience it as part of the fun (Volden et al., 2025). The useful split is productive confusion, the puzzle the player wants to solve, versus blocking confusion, the moment they stop knowing what the game wants from them. Only the second one costs you D1.

5. Sort every problem into one of three kinds

Most FTUE reports are a list of complaints. Sort them, because each kind has a different fix:

KindWhat it looks likeLikely causeFix direction
Didn't see itPlayer never looks at the element, or taps around itPlacement, contrast, competing animationMove it, isolate it, time it
Saw it, didn't get itPlayer looks at it, then does the wrong thingUnlabeled icon, missing context, unclear goalLabel it, show it in use, teach it at the moment of need
Got it, didn't carePlayer understands and leaves anywayWeak first reward, slow route to the fun partShorten the route, move the first payoff earlier

The third row isn't a comprehension problem, and no amount of tutorial rewriting fixes it. Knowing which row you're in saves a sprint.

Apple's own guidance for game onboarding points the same way: teach the core loop in short steps at the moment it's needed, let players skip, and hold back ratings prompts, notification requests and purchases until onboarding is done (Apple, Onboarding for Games). Nielsen Norman Group found that people forget front-loaded tutorials and recommends help in context (NN/g, 2023).

6. Re-test every build, not every quarter

Here's the trap in steps 1 to 5: done properly, they're slow. A human playtest round typically takes about 48 hours to return results (PlaytestCloud), and each first-timer can only be used once. So most studios test the FTUE at milestones, and the builds in between ship on faith. That's exactly when an innocent change to the HUD, a new pop-up or a reordered reward quietly breaks the first session.

AI first-time players are the way to close that gap, with two conditions.

They have to see what the player sees. Most test-automation tools find buttons by reading the app's element tree. Games mostly don't expose one. On one Android game we measured, the game screen exposed 4 elements to automation tools, against 33 on the phone's home screen. That's evidence that games can expose almost nothing, not that every game always does, but it's why an AI player has to work from the pixels on screen, the way a person does.

They have to stay newcomers. This is the part most AI testing gets backwards. A more capable agent is a worse stand-in for a confused new player. In a study comparing LLM and human walkthroughs of an app interface, the models completed more tasks than humans, took more direct paths, and identified fewer potential failure points (Zhong et al., 2026). Simulated users more broadly tend to cooperate through errors and rarely show real frustration, a kind of "easy mode" that inflates success (Mind the Sim2Real Gap, 2026). An AI tester that plays well will walk straight past the step that loses your players. The useful version is held to a first-timer's knowledge on purpose.

Tookii (tookii.ai), the synthetic user testing platform, runs FTUE tests for studios this way: AI first-time players with different player profiles, playing your build from the screen, reporting where each one stopped understanding. It's in early access, run with studios directly.

7. Read your store reviews by version

After launch, your players run the FTUE test for you, at scale, and write up the results. Tag reviews by app version and look for onboarding language: "confusing," "don't understand," "tutorial," "what do I do." A spike right after an update is a regression report you didn't have to pay for.

Which method answers which question

MethodBest atTurnaroundCan't tell you
Moderated sessionsWhy a player got confused, in their wordsDays to weeksAnything at scale
Unmoderated playtest panelsWatching 10+ real first-timers~2 days per roundEvery build; panels skew experienced
Tutorial funnel analyticsWhere players drop outLiveWhether finishers understood
Soft launchReal D1/D7 on real installsWeeksWhy the numbers are what they are
AI first-time playersWhere understanding breaks, on every buildHoursLong-term fun (pair with human playtests at milestones)

Where AI testers fit

AI first-time players are strongest at one job: finding where a newcomer stops understanding, on every build, before real players ever meet it. That's the part of FTUE testing studios most often skip, because it's the part that's hardest to repeat. A few calls still sit better with people. Grading difficulty is one: simulated users tend to read moderately hard tasks as easier than people find them (Lost in Simulation, 2026), which is why a well-built AI tester reports where understanding broke instead of scoring how hard a level is. Whether the game is fun enough to come back to is the other, and that's what milestone playtests and soft-launch numbers are for. The setup that works runs both: AI first-timers on every build, people at the milestones.

FAQ

What is FTUE in mobile games? FTUE (first-time user experience) is everything a new player meets from first launch until they understand the core loop: title screen, tutorial, first levels, first rewards. On mobile it usually has a few minutes to work.

How many playtesters do I need for an FTUE test? Around 10 first-time players per player profile is common practice, since onboarding problems show up at small sample sizes. Each tester can only experience your FTUE once, so recruit fresh people for every round.

What's a good D1 retention for a mobile game in 2026? GameAnalytics' 2026 benchmarks put the median D1 at about 22% and the top 25% at about 30%. Different sources define retention differently, so compare against one source, not a blend.

Can AI replace human playtesters? AI first-time players take on the most repetitive part of FTUE testing: checking every build for where a newcomer stops understanding, in hours instead of days. Human playtests stay valuable at milestones, for judging fun and long-term appeal. The strongest setup runs both.

How long should a mobile game's FTUE be? As short as it takes to teach the core loop and deliver the first real reward. Teach one thing at the moment it's needed, let players skip, and hold every prompt and purchase until onboarding is done.


We run FTUE tests for studios with AI first-time players. Early access is open to mobile game teams. Book a demo.

Your tutorial taught your players something. Do you know what it was?

Run your first test.