Skip to content
E
Egmatic
How to Playtest Your Game Before the Art Exists
playtest without artgraybox testingplaceholder artplaytestinggame prototypingindie game dev

How to Playtest Your Game Before the Art Exists

Yes, you can playtest a game with no art. Run a graybox build: plain shapes, counters instead of visuals, tasks that test the loop, not the look.

Vladislav KovnerovOctober 5, 202614 min

A game with no art is still a game, and it can still be tested. The community threads we read every week keep proving the point from both directions: one developer asks when it makes sense to invite friends into a build that has no art, and a playtester in a crafting thread explains what happens when the interface cannot explain itself: "if I don't know what your parameters do, I will pick randomly and see what happens." The first question is about timing. The second is about preparation. Both have real answers, and neither answer is "wait for the art".

This article is about running playtests on a build that consists of gray boxes, flat colors, and numbers where the pictures will go: what such a build can legitimately tell you, how to prepare it for someone else's hands, and how to read a session in which there is nothing to look at. A disclosure before we start: this blog belongs to Egmatic, a no-code 2D editor in pre-alpha, and what that has to do with graybox testing is near the end.

Quick answer

Your buildQuestions it can answerNotes to log but not act on
Paper prototype, no engineDoes the rule produce decisions at all?Anything about controls, pacing, or feel
Graybox: shapes, no artDoes the loop repeat? Do controls survive a stranger? Where does difficulty bite?"It looks boring", mood, art critique
Placeholder pass: free assets, flat colorsDo roles read? Does first-time guidance work?Style judgments, animation polish
First-pass art on one screenDoes the feel land? Does the promise match the loop?Whole-game balance conclusions

The sequence matters because each stage answers a different question and the earlier stages are cheaper to change. Testing whether an idea is worth a build at all is the subject of validating game mechanics before you build, and the full paper-to-engine pipeline lives in how to test game ideas fast. This article picks up at the moment the graybox build exists and you want other people in it.

What graybox testing means

The vocabulary is worth fixing first, because the terms get used loosely and the differences are exactly the differences between test stages.

TermWhat it isWhere it sits
Paper prototypeThe rules of the game played on paper, no engineBefore any build
Graybox (also greybox, whitebox, blockout)A playable build made of plain shapes and placeholder assetsThe main subject here
Placeholder artSimple or free assets standing in for final art, built to be replacedThe pass that readies a graybox build for outside hands
First playableThe first version with the major gameplay elements workingThe milestone a graybox build matures toward
Vertical sliceOne segment of the game at final qualityAfter the loop is proven, before full production

Two of these are documented practice, not folk terms. Game development's milestone vocabulary defines the first playable as the first version containing the major gameplay elements in working form, often grown directly out of the pre-production prototype, with alpha arriving when the game is feature complete and assets are still only partially finished. And placeholder art has its own name, programmer art, precisely because it is made when an asset is needed immediately and is intended to be replaced before the project ships. In other words, the industry's default assumption is that a game is playable, and testable, long before it looks like anything.

The readiness ladder: what feedback is valid at each stage

The core skill of graybox testing is matching the question to the coarseness of the build. Ask a graybox build about mood and you will collect noise. Ask it about the loop and you will collect fixes.

Rung one: the loop, tested alone. One mechanic, one room, boxes. The question is narrow: does the chain of actions the player repeats hold attention without any dressing? This is the stage our mechanics-validation guide operates at, and its discipline carries over: every polished asset you add is a confound, because it gives testers something to react to that is not the thing you are testing.

Rung two: the loop in sequence, still no art. Several rooms or screens, a goal, a fail state. Now you can test pacing, difficulty curves, and whether controls survive being learned without you in the room. Feel, in the mechanical sense of timing and response rather than polish, is also a graybox question, and how to tune it is its own article in how to make your game feel good. Behavior at this rung is high-quality signal: where testers stop, what they try next, what they ask aloud.

Rung three: the placeholder pass. Free asset packs and flat colors give every shape a visible role. This is the rung at which friends and strangers enter safely, and it is the cheapest rung at which to catch readability problems, because a misread silhouette is a design finding, not an art problem.

Rung four: one screen with real art. Mood, feel, and the match between what the game promises and what it delivers. This is vertical-slice territory, and conclusions drawn here about the whole game's balance do not transfer backward: a slice with finished art plays differently from the graybox build the balance was tuned on.

The rule that holds the ladder together: feedback about what a build does not contain yet is noise. Log it, because "it looks empty" repeated across rounds is data about promise, but do not act on it by drawing art early. The production logic is arithmetic: art made for a loop that still changes is rework, and the placeholder exists to absorb exactly that risk.

What the emptiness tells you

An artless build is not a degraded test instrument. Handled right, it is a sharper one, because with the pictures gone, the game has to communicate through structure, and every place it fails to is a finding you would otherwise have paid an artist to discover.

Counters are the art of a graybox build. Damage numbers, timers, attempt counters, a visible score: put a number anywhere the final game will communicate with a picture. Counters make the loop's grammar visible to the tester and give you something concrete to compare across sessions: if attempt three of a jump section takes eleven tries, that is a difficulty finding with a number on it.

Color is your first readability test. Assign one color per role, enemy, hazard, pickup, exit, and never reuse it. When a tester avoids a helpful box because it shares a color with a hazard, you have found a communication problem while it costs nothing to fix. Ask testers to say what they think each shape is; every mismatch between their guess and the shape's job is a finding, collected before a single asset exists.

Cheat keys respect the tester's time. Keys that skip to the section under test, toggle invulnerability, or reset a room turn a thirty-minute favor into a focused fifteen minutes. Testers are donating attention; spending it on replaying content you already trust wastes the donation and dilutes the data.

Unexplained options collect no data. The playtester quoted above, picking crafting parameters at random, was not failing as a tester: the screen was failing as a test instrument. An options menu whose parameters have no visible meaning measures nothing except patience. Either label what each option does in plain language or conclude that the system is not ready to be tested, which is itself a real answer about the build's stage.

Silence is a channel. Note what testers walk past without touching. A system no one engages with across three rounds is telling you something the exit survey never will.

Preparing the build for other people's hands

A graybox build going to testers needs preparation that has nothing to do with graphics.

  • Stability outranks polish. A crash in minute three ends the session and teaches you nothing. Freeze a build that ran clean to the end locally, and hand out that exact build.
  • Fifteen to thirty minutes of content, no more. That is the session length volunteers can actually give, the same range our session guide uses, and it forces you to decide what this round is about.
  • Write the tasks down before anyone arrives. Two or three concrete goals, in the tester's language: "get to the second room", "make one sword", "survive sixty seconds". Vague invitations produce vague sessions.
  • Put the brief inside the build. A first screen with the goal, the controls, and one line of context replaces the explanation you will be tempted to give in person, and what testers do with that screen is itself first-impression data.
  • Choose a carrier that costs nothing to enter. A web build behind one link removes installation from the session; how to host a playable web demo covers the options. When the rounds get bigger and more formal, Steam's Playtest format gates access behind a request button, and when the question becomes whether strangers should see the game at all, that transition has its own article in when to turn your playtest into a Steam demo.
  • Recruit for the round you are running. Five testers matched to the game's audience, the pattern our playtester recruiting guide details. The numbers behind small rounds are well established in usability research: a single tester surfaces about 31 percent of problems, five testers roughly 85 percent, and three rounds of five beat one round of fifteen because you fix between rounds.

Running the session with nothing to look at

The session itself runs on the script our playtest session guide already documents, so only the graybox-specific parts belong here.

The first is that thinking aloud carries more weight than usual. The method was borrowed from software usability, where Nielsen called it perhaps the single most valuable usability engineering method, and in a graybox session it does double duty: the monologue tells you what the tester is trying, and, uniquely at this stage, what the tester thinks the shapes are. Prompt for narration with the same neutral prompts the method prescribes, and add one graybox question of your own: "tell me what you believe that is" while pointing at a shape. The answers are your readability data.

The second is that the graybox session makes the oldest rule in usability easier to follow: watch what people do, not what they say they would do. There is no art to opinionate about, so opinions arrive thinner and behavior stands out more. A tester stuck at a gap for four attempts is a finding regardless of what they politely say afterward.

The third is the discipline of the log. One line per incident: timestamp, what the tester was trying, what they did, what they said. After the round, sort the log into behavior findings to act on and style notes to park. The sorting step is where most graybox sessions are won or lost, because the raw log of an artless build always contains a chorus of "looks unfinished", and the finding is never that.

Common mistakes

  • Waiting for the art to test. Every loop problem you find after the art pass costs an artist's time to fix; found in graybox, it costs an afternoon. The placeholder exists to absorb this risk.
  • Asking the graybox build art questions. Mood, promise, visual style: log the answers, act on none of them, and schedule the question for a build that can answer it.
  • Polishing one screen instead of roughing the whole loop. A beautiful first room in an untestable game is a vertical slice arriving two milestones early.
  • Explaining the build before the session. Every explanation you give is a finding you delete. The brief in the build either works or it does not, and both outcomes are data.
  • Testing unlabeled systems. Screens full of unexplained parameters generate random clicks, and random clicks are not feedback. Label the options or pull the screen from this round.
  • Acting on single-tester passion. One impassioned critique is a hypothesis; three testers hitting the same wall is a finding. Fix the top three, then recruit new people.

How Egmatic fits

The work this article describes is a loop: build rough, put it in front of someone, watch, change one thing, ship the next build to the same link. Most tools support the first half and leave the shipping half to the developer. Egmatic is a no-code 2D editor and engine being built the other way around: the scene you graybox is already running in the engine that ships, and the ship layer exists to carry builds to players, testers included, without code and without a backend standing between a fix and the next session.

The AI inside works as a craft amplifier under your direction: you decide what changes after a round, it executes, and every file stays yours. That division of labor matters most exactly here, in the weeks when the loop is changing every day and the bottleneck is the distance between "the tester said this" and "the tester is playing the fix". We will state the stage plainly: Egmatic is pre-alpha, we announce no dates we cannot keep, and the waitlist at egmatic.com is where you can tell us what your own test loop is missing, because in pre-alpha the requests arriving now are the ones that steer the build.

Put the build in front of people, art later.

A graybox build one click away from a tester is a playtest program, not a milestone. Egmatic is a no-code 2D editor with a ship layer that carries test builds to players without code or a backend. Reply on the waitlist with what your test loop is missing: requests sent during pre-alpha are the ones that steer the build.

No spam. Unsubscribe anytime.

Conclusion

Test the game that exists, not the game you will eventually draw. A graybox build answers the questions that decide whether the game is worth drawing: whether the loop repeats, whether controls survive strangers, where difficulty bites, whether shapes communicate their jobs. Prepare the build like an instrument, counters for numbers, colors for roles, labels for options, cheat keys for time, and the emptiness becomes the sharpest part of the test. Run small rounds of five, watch behavior, log everything, act on repetition, and add art only when new testers stop finding new loop problems. The art pass then lands on a target that has stopped moving, which is the cheapest gift a developer ever gives an artist.

Sources

  1. Wikipedia — Video game development, milestones section: the first playable as the first version containing the major gameplay elements in working form, often based on the pre-production prototype; alpha as feature complete with assets only partially finished
  2. Wikipedia — Programmer art: placeholder assets made when an immediate need exists and intended to be replaced before the project is published
  3. Kenney — Prototype Textures and the surrounding asset packs: placeholder-ready 2D assets released under Creative Commons CC0
  4. Nielsen Norman Group — Thinking Aloud: The #1 Usability Tool, Jakob Nielsen, 15 January 2012: "thinking aloud may be the single most valuable usability engineering method"; neutral prompting to sustain the monologue
  5. Nielsen Norman Group — Why You Only Need to Test with 5 Users, Jakob Nielsen: about 31 percent of problems found by a single tester, roughly 85 percent by five, three studies of five over one of fifteen
  6. Nielsen Norman Group — First Rule of Usability? Don't Listen to Users, Jakob Nielsen, 4 August 2001: watch what people do; do not believe what they say they do or predict they will do
  7. Community reads run by our team, September 2026: the weekly threads behind our feedback digest, including the developer asking when to invite friends into a build with no art and the playtester comment "if I don't know what your parameters do, I will pick randomly and see what happens" (23 September 2026)

Related Posts