Skip to content
E
Egmatic
How to Run a Playtest: Script, Questions, Signal vs Noise
playtestinggame testingplaytest sessionindie game devuser researchplayer feedback

How to Run a Playtest: Script, Questions, Signal vs Noise

Run each playtest with a script: a short brief, silent observation while the tester thinks aloud, and questions about what just happened, not opinions.

Vladislav KovnerovOctober 2, 202616 min

A playtest session is forty-five structured minutes: a short brief, half an hour of watching one tester play while thinking aloud, and ten minutes of questions about what just happened. The facilitator's discipline is mostly negative. Do not explain the game, do not defend a choice, do not help past the first stuck moment. You are there to record what a player does; what a player says they would do is a different and far cheaper currency.

Recruiting the testers is a solved problem we covered separately in how to recruit playtesters. What the community threads keep showing, in the reads we run every week, is the gap one step later: developers who have assembled five strangers, a build, and a call, and then spend the call explaining their design decisions while the strangers nod politely. The session is an instrument, and this article is its manual: the preparation, the script, the questions that produce fixes, and the filters that separate a finding from an opinion. A disclosure before we start: this blog belongs to Egmatic, a no-code 2D editor in pre-alpha, and what that has to do with playtesting is near the end.

Quick answer

Part of the sessionWhat you doThe one rule
BeforeFreeze the build, write two or three objectives, prepare a note sheetIf you cannot name what the session should learn, cancel it
Brief, 2 minutesSet expectations: the game is being tested, not the tester"You cannot do anything wrong here"
Play, about 30 minutesTester plays and thinks aloud; you stay silent and take timestampsPrompt attention, never give answers
Debrief, 10 minutesQuestions about the last half hour, not about the game in generalAsk about events, not verdicts
After, same dayWrite the friction log while the memory is freshFix the top three, then test the fixes on new people

What a session is, and what it is not

The shape comes from usability research, and it has three parts: a facilitator, a tester, and tasks. The facilitator gives the tester something to do in the game, observes what happens, and asks follow-up questions to draw out detail. That is the entire machinery. Jakob Nielsen's summary of the core method runs to three things: recruit representative users, give them representative tasks, and shut up and let the users do the talking.

Two things a session is not. It is not a focus group: opinions volunteered by people who have not played are surface impressions, and the research record on them is blunt. Watch what people actually do, do not believe what they say they do, and definitely do not believe predictions of what they may do in the future. It is also not quality assurance: a tester who reports a crash has done you a favor, but bug collection is a side effect, not the goal. The goal is to watch a stranger meet your game with no one from your side in the room.

A session also has a known failure mode that has nothing to do with testers: teams leave playtesting too late and then squeeze impossible questions into it, a pattern games user researchers see constantly. Five testers in week ten cannot tell you whether the core idea is worth finishing. Whether a mechanic is fun before anything is built on it is a separate, earlier question with its own process, covered in validating game mechanics before you build.

Before the session: one page of preparation

Everything below fits on one page, and it should exist before the first tester arrives.

  1. Two or three objectives, phrased as questions. "Can a new player reach the first boss without help?" and "Where do players first stall in the crafting menu?" are objectives. "Get feedback" is not. If nothing in the build is uncertain enough to phrase as a question, the session has nothing to collect.
  2. A frozen build. Freeze it the night before and test the freeze on your own machine. A session that dies to a crash at minute three does not come back; you lose the tester's patience along with the data.
  3. Tasks, written down. One opening task ("start the game and play as you would at home") plus one or two specific tasks tied to the objectives. Tasks are what keep the session moving when the tester looks to you for direction.
  4. A note sheet. One line per incident: timestamp in words, what the tester was trying to do, what they did, what they said. Blank columns ready, because you will be writing during the session, not after.
  5. The offer and the channel. Know before the session what the tester gets, a key, credits or cash, and where follow-up questions go if any are needed.

The build has to reach the tester somewhere. A web build behind one link is the lowest-friction carrier, and hosting a playable web demo covers that side; the session discipline below is the same whatever carries the build.

The script: minute by minute

A forty-five-minute session, written out. Shorter first-impressions tests shrink the middle row, nothing else.

ClockPhaseWhat you say or do
0:00WelcomeThank them, confirm the time, start recording notes
0:01Expectations"The game is being tested, not you. You cannot do anything wrong. Rough edges are exactly what I need to see."
0:02Think-aloud ask"Play as you would at home. Say whatever crosses your mind, even if it seems obvious. If you go quiet, I'll nudge you."
0:03Opening task"Start the game and play however you like. I'll stay quiet."
0:03 to 0:35ObservationSilence. Prompt only when the monologue stops: "What are you trying to do right now?" Log every pause, backtrack and muttered remark with a timestamp.
StuckThe waitCount to ten. Then: "What are you expecting to happen?" Point at nothing.
0:35DebriefThe questions below, about the last half hour
0:45CloseThanks, the offer, what happens with their input, invitation to the next round

The script's value is not ceremony. A tester who hears "you cannot do anything wrong" in the first minute relaxes, and a relaxed tester misuses your game honestly. A facilitator holding a printed script does not improvise explanations at minute four, which is the single most common way sessions are ruined.

During the session: silence is the technique

The play portion runs on the think-aloud protocol: the tester plays while continuously saying what crosses their mind, and you keep them talking. Usability research calls this the single most valuable method in the field, and the reason shows immediately: a running monologue surfaces misconceptions as they happen. A tester who says "I guess this opens the door" while pressing the wrong object has just handed you a redesign, in their own words.

Keeping the monologue alive takes three neutral prompts, and only these:

  • "What are you trying to do right now?"
  • "Tell me what you're thinking."
  • "What did you expect to happen just then?"

Each returns attention to the game without planting an interpretation. The failure mode to avoid is the leading prompt, "Did you see the button?", which answers its own question and converts an observation into coaching. The method is forgiving of amateur facilitation in most respects; the one unforgivable move is putting words into the tester's mouth, because everything after that moment is contaminated.

What you write during the session matters more than what you say. Note the tester's own words for anything they react to, quoted, because "seemed confused by inventory" decays into nothing while "said 'why can't I stack these' twice" survives to the next morning. Timestamp roughly. Mark task outcomes as reached, reached with struggle, or abandoned. And note silences too: a mechanic a tester walks past without a word is data about interest, the quietest kind you will collect.

The questions that earn their place

Questions about the last half hour produce fixes. Questions about the game in general produce opinions. The table keeps them apart.

Ask thisBecauseNot thisBecause
"What are you trying to do right now?"Surfaces intent, so you see where intent and game disagree"Do you like it?"Invites a verdict the tester cannot actually know
"What did you expect to happen just then?"Captures the misconception at the moment it fires"Is that clear?"Nobody answers no honestly to the developer's face
"Where did you first feel stuck?"Locates the earliest break in your onboarding"Was it too hard?"Difficulty is a conclusion, not an observation
"What would you do next if I weren't here?"Reveals the unaided next step"Would you play again?"A prediction of future behavior, the least reliable data there is
"Which part would you show a friend?"Finds what carries the game"What would you add?"Invites design work the tester never signed up for

In the debrief, two more earn their keep: "Walk me through the last thing you did before we stopped," which replays the ending while it is warm, and "Which part would you apologize for?", which testers answer with surprising honesty and which points straight at the friction you already half-knew about.

Signal vs noise: the three filters

Raw session output is noise-heavy: forty opinions, six incidents, one tester who wanted a different game entirely. Three filters clean it.

First, behavior beats statement. The rule from usability research holds for games unchanged: watch what the tester did, treat what they said as commentary on it. A tester who calls the menu "fine" while opening the wrong screen three times has told you the menu is not fine, in the channel that matters. Verdict questions are noise by construction, which is why the question table above refuses them.

Second, repetition beats volume. One tester's impassioned critique is a hypothesis; three testers hitting the same wall is a finding. This is arithmetic the small-study research already did for you: five testers surface roughly 85 percent of a product's usability problems, and the recommended pattern is not one bigger study but three rounds of five, fixing between rounds. The corollary for a solo developer: do not act on one session's opinions. Log them, fix the top three, and spend the next session on new people, which is also where the recruiting guide's advice about small matched batches pays off. Testers who come back for a second round are the compound interest of the whole practice, and keeping a small room that does return is its own playbook, covered in building a game community before launch.

Third, a feature request is a symptom. "You should add a dash button" is not design input; it is a player solving a friction problem they experienced, with the tools they have. Diagnose the friction underneath, by asking what the tester was trying to do when the wish occurred, and decide the fix yourself. Acting on requests as stated makes the loudest tester your designer and quietly drops the four who hit the same wall in silence.

SeverityWhat it looks like in the sessionExampleReaction
BlockerTester cannot continue without your helpStuck in the first room, no path forwardFix the same day
Major confusionLong pause, wrong expectation spoken aloudTries to talk to an NPC with the attack buttonFix this week
FrictionPause, backtrack, muttered remarkReopening the menu twice to compare itemsNext build
TasteA preference with no observed struggle behind it"I'd prefer a darker palette"Log, decide alone

Remote and recorded sessions

The script survives the move to a screen share almost untouched: same brief, same prompts, same silence, with the note sheet replaced by timestamps against the recording. Discord screen share is the usual free carrier, and the web-demo hosting guide above covers getting the build itself one click away.

Recorded panels are the industrial version: strangers play unmoderated while software records their session and voice, and you watch later. PlaytestCloud is the established provider, and its free trial is a single-session playtest with two players, each recorded for up to an hour, which is enough to watch real strangers meet your tutorial once before paying for anything. Recorded sessions trade the live debrief for scale; the questions that work live become questions you answer by re-watching.

The weakest form is the one most developers default to: sending a link and asking for thoughts afterwards. Unmoderated written feedback contains no behavior at all, only self-report, which the filters above rank last. It is still worth having, as long as nobody mistakes it for a session.

Common mistakes

  • Explaining the game before play. Every sentence of orientation deletes a piece of the tutorial's test data. The brief sets expectations; it does not teach.
  • Helping at the first stuck moment. The stuck moment is the product. Wait the ten seconds out.
  • Asking "is it fun?" Fun is inferred: whether the tester kept playing, returned after a break, or asked when they could play again. Asking collapses it into politeness.
  • Testing an unfrozen build. The morning's fix is the afternoon's crash, and the session burns on it.
  • One session and done. A redesign ships new problems of its own; the second round exists to catch them.
  • Taking notes only of complaints. What a tester skipped without a word is the interest data; write that down too.

How Egmatic fits

The work around the session is mostly publishing work: freezing a build, getting it one click from a stranger, collecting what happened, and shipping the fix as the next build testers can reach. That road from scene to released build is the layer Egmatic is built around: a no-code 2D editor and engine whose ship layer carries the game to players instead of stopping at the canvas. The AI inside works as a craft amplifier under your direction: you decide what changes after a round of sessions, it executes, and every file stays yours, which matters when the fix needs to be in testers' hands the same week.

We will state the stage plainly: Egmatic is pre-alpha, and we do not announce dates we cannot keep. The waitlist at egmatic.com is where the ship layer takes shape, and it is where you can tell us what your testing loop is missing, because in pre-alpha the requests arriving now are the ones that steer the build.

Fix what players actually do.

A playtest is only as good as the session around it. Egmatic is a no-code 2D editor with a publishing layer that carries test builds to players without code or a backend. Reply on the waitlist with what your testing loop is missing: requests sent during pre-alpha are the ones that steer the build.

No spam. Unsubscribe anytime.

Conclusion

The session is not a conversation; it is an instrument. A two-minute brief that removes fear, half an hour of silence while one tester thinks aloud, ten minutes of questions about what just happened, and a friction log written the same day. Silence is the technique, repetition is the signal, and a feature request is a symptom with the diagnosis left to you. Run three rounds of five on frozen builds, fix between rounds, and the process pays for itself in the first session where a stranger gets stuck exactly where you always secretly suspected they would.

Sources

  1. Nielsen Norman Group — Thinking Aloud: The #1 Usability Tool, Jakob Nielsen, 15 January 2012: "thinking aloud may be the single most valuable usability engineering method"; the method's three steps (recruit representative users, give them representative tasks, shut up and let the users do the talking); facilitator prompting to sustain the monologue; findings remain robust to imperfect facilitation except when the facilitator puts words into the participant's mouth
  2. Nielsen Norman Group — First Rule of Usability? Don't Listen to Users, Jakob Nielsen, 4 August 2001: watch what people actually do; do not believe what they say they do; definitely do not believe what they predict they may do in the future; preference expressed without use reflects surface features
  3. Nielsen Norman Group — Why You Only Need to Test with 5 Users, Jakob Nielsen: about 31 percent of problems found by a single tester, roughly 85 percent by five, recommendation of three rounds of five over one study of fifteen
  4. Nielsen Norman Group — Usability (User) Testing 101, Kate Moran, 1 December 2019, last reviewed 15 July 2026: session anatomy of facilitator, tasks and participant; facilitator observes behavior, listens for feedback and asks follow-up questions
  5. Games User Research — How To Run A Games User Research Playtest, Steve Bromley, updated 5 June 2024: end-to-end playtesting process from objectives through unbiased data collection to reporting; teams running playtests too late and getting low-impact studies
  6. PlaytestCloud — Pricing: free trial of a single-session playtest with two players, each recorded for up to 60 minutes
  7. Community reads run by our team, September 2026: weekly playtest threads on r/gamedev and r/playmygame, including "What did your first real playtest reveal"; recurring reports of sessions spent explaining design choices instead of observing play

Related Posts