
Build a Game With an AI Assistant: What Actually Works
AI assistants draft assets, dialogue, and balance tables well under your review, and fail at whole games from one prompt. Here is the split that works.
An AI assistant helps you build a game when it drafts and you decide. It hurts you when the roles reverse. The reliable split in 2026 is unglamorous and repeatable: assistants draft placeholder assets, dialogue, item text, store copy, and balance tables; they explain unfamiliar formats and review your event logic for missed cases. They do not invent a shippable game from one prompt, they hold no opinion about whether your jump feels right, and they cannot keep an art style coherent across two hundred files. Developers who treat the assistant as a fast junior collaborator ship sooner. Developers who treat it as the designer ship nothing they can maintain.
A disclosure before the drafting starts: this blog belongs to Egmatic, a no-code 2D editor in pre-alpha being built AI-native, and what that means in practice is near the end. This article is the operator's side of AI, what to delegate, how to prompt it, what to keep. The politics of the AI content flood on Steam, visibility, and disclosure fatigue, is a separate subject covered in how human-made indies stay visible among AI-generated games. If the no-code part is new to you, the visual editor guide is the primer, and low-code versus no-code draws the line this article assumes.
Quick answer
| Task | Verdict | The catch |
|---|---|---|
| Placeholder and concept assets | Works | Style drifts across a set; final art needs one human hand |
| Dialogue, item text, store page copy | Works | Review line by line against a tone document |
| Balance and content tables | Works | You define the schema, and you keep the numbers |
| Explaining formats, reviewing event logic | Works | It reads fast but confirms nothing; you verify |
| A whole game from one prompt | Fails | A weekend prototype at best; unreviewed logic breaks silently |
| Game feel, timing, difficulty curves | Fails | The assistant has never played your game |
| A consistent final asset set | Mostly fails | Regeneration is cheap; coherence is not |
The pattern behind the table: delegation works where output is cheap to review against a format you already control, and fails where reviewing is the whole job.
Two studies, and what they measured
Every argument about AI and development speed eventually quotes one of two studies, so it is worth knowing what each actually measured.
The optimistic one is a controlled experiment from 2023, published by Peng, Kalliamvakou, Cihon, and Demirer. Ninety-five freelance developers were asked to implement an HTTP server in JavaScript; half had GitHub Copilot. The Copilot group finished in 71 minutes against 161 for the control group, 55.8 percent faster. Note what the task was: bounded, well specified, greenfield, with success obvious at a glance.
The sobering one is a randomized controlled trial run in 2025 by METR, an AI safety research organization. Experienced open-source developers worked on their own mature repositories, the code they knew best, with AI tools switched on or off at random. The AI group finished 19 percent slower, with a confidence interval from 2 to 39 percent slower. The part worth framing: before the trial the developers predicted AI would speed them up by about 24 percent, and afterward they believed it had, by about 20 percent. The feeling and the stopwatch disagreed.
Both results are true at once. Bounded tasks with clear specifications speed up. Systems work slows down, because reading and verifying someone else's output is itself the job, and that cost compounds. Game development is mostly systems work with islands of bounded tasks: your item tables, your dialogue passes, your placeholder art are islands; your core loop, your art direction, your difficulty curve are the mainland. Delegate the islands. Keep the mainland.
What works in practice
Placeholder and concept assets
Assistants are strongest at the start of the art pipeline, where quantity beats quality: enemy silhouettes to test a combat idea, palette studies, prop variations to fill a scene before committing real art time, tileable texture drafts. You explore twenty directions in the time a hand-drawn pass gives you two.
Two limits keep it honest. Style drift is structural: each generation re-rolls small decisions, so fifty generated sprites read like fifty artists unless you anchor every request to reference images and iterate on the shortlisted few. And the output is a starting point, not a deliverable; the workflow that survives contact with a real project is generate broadly, select ruthlessly, and repaint or composite the finalists yourself. The ownership reasons for that last pass are below, and they are not cosmetic.
Text at volume
Dialogue barks, item descriptions, tutorial lines, achievement names, store page variants: text is the one medium where review is genuinely cheap, because you can judge twenty lines in a minute. Two habits separate the developers who get value here from the ones who get mush.
Keep a tone document, three adjectives that describe your game's voice, two sample lines that nail it, a short list of forbidden words, and paste it into every session. And ask for variants rather than verdicts: five options per line, then pick and edit. The first output a model offers is its most probable continuation, which is another word for its most average one; the fourth or fifth option, edited, is where a voice lives. Translation drafts work the same way, with the same rule, and a native speaker still has to read them before anything ships.
Balance and content tables
This is the strongest fit for a no-code workflow, because your items, enemies, and levels already live in tables. Assistants fill schemas fast. Give them columns, constraints, and a count: name, slot, cost, damage, cooldown, rarity, one-line note; twenty healing items for a run-based roguelike; cost between 3 and 7; draw names only from this vocabulary list. You get a spreadsheet in seconds, delete the boring rows, and tune the numbers yourself after watching a playtest, because the model cannot know your game collapses after the second boss.
One check stays non-negotiable: hallucinated references. Ask loosely and the assistant will invent an item that interacts with a mechanic you do not have. The schema must describe what exists, not what could exist, and anything referencing a mechanic gets verified against the actual event sheet before it enters the game.
Reading, reviewing, explaining
Paste an event sheet into a chat and ask which player actions have no handler. Ask what happens if the timer reaches zero while the pause menu is open, or if the shop closes during a purchase. Assistants are strong reviewers precisely because review is reading with a checklist, and they never get tired of your project the way a friend eventually does. They also explain unfamiliar formats well: an export structure, a save file layout, a data format you inherited. Treat every answer as a lead to verify, never as confirmation, and the reading pass pays for itself.
The same strength applies to playtest notes. Raw feedback is long, repetitive, and emotional; an assistant will cluster it into themes in seconds. Insist that quotes stay verbatim, then read the clusters against the raw notes yourself. The mechanics of getting good notes in the first place are covered in how to run a playtest; the assistant only helps once the notes exist.
Prompt patterns that survive contact with an editor
- Schema in, table out. Never ask for "some items" or "a feel"; ask for rows with named columns, bounded values, and a fixed vocabulary. Formats discipline both sides of the conversation.
- Variants, not verdicts. Request several options, choose one, edit it. A model asked once will answer once, and its first answer is its most average one.
- Ground it in the editor's real vocabulary. The characteristic failure of assistants wired to editors is confident invention: event names, conditions, and actions that do not exist. The fix is to paste the actual list of conditions and actions your editor supports and forbid anything outside it. This matters more as wiring gets tighter: the Model Context Protocol, which Anthropic open-sourced in November 2024 and which now sits under the Linux Foundation's Agentic AI Foundation, standardizes how assistants call external tools, and developers around no-code editors have started building these bridges; in our own customer interviews, the most engaged developer we have spoken to was building exactly this, a bridge between an AI assistant and Construct 3. Whether the bridge is a protocol or a copy-paste ritual, the rule is identical: the assistant may only use names from the list you gave it.
- Small batches you can review in minutes. Review debt is the METR tax. A batch you fully check costs less than a pile you plan to check later, because later never comes before the breakage does.
- Keep a decisions file. What you accepted, what you changed, what you threw away. It is your authorship record for the ownership and disclosure questions below, and it is the training your own judgment gets from the process.
What still fails
The one-prompt game
In February 2025 Andrej Karpathy named the thing: vibe coding, a way of working where, in his words, you "fully give in to the vibes" and "forget that the code even exists". For a weekend toy, that is a fine way to spend a Sunday. A shipped game is not a weekend toy; it is a maintenance project with players, and the parts nobody reviewed are exactly the parts that break silently in month three. The METR trial is the measured version of this sentence: the developers slowed down were the ones maintaining real systems, and they enjoyed it while it happened.
Feel and judgment
Coyote time after a ledge, a few frames of input buffering, a difficulty curve that respects the player: these come from playing the game, and the assistant does not play. It can analyze, which is worth having: ask for ten ways players might break your shop menu, and it will find four you missed. But the questions "is this fun" and "does this feel right" have no chat answer. The craft layer those verdicts come from is its own discipline, covered in how to make your game feel good. Delegate analysis, never the verdict.
Consistency at scale
A demo tolerates fifty assets that almost agree. A product does not, and coherence across a large set is precisely what per-generation randomness fights. Regeneration is cheap; a hundred files that agree is expensive. Plan the unifying human pass, one artist or one very disciplined you, before the asset count gets there, because the fix costs more than the generation ever did.
Ownership
The United States Copyright Office settled its position in a report on copyrightability released January 29, 2025: material generated entirely by AI is not copyrightable, and prompts alone do not make you the author. What can be protected is your selection, your arrangement, and your modifications, with the AI-generated portions disclaimed in a registration. For a shipping indie the translation is direct: the draft is not the asset. The pass where you repaint, retune, and choose is both what makes the work defensible and what makes it worth shipping. Ownership, like coherence, is a workflow, not a footnote.
What the store asks you to say
Steam has run AI disclosure through the Steamworks Content Survey since January 2024: you declare pre-generated content made before the game runs and live-generated content made while it runs, the disclosure appears on the store page, and players can report content they believe breaks the rules through the in-game overlay. In January 2026 Valve rewrote the form so the question is what players actually see or hear: in-game assets, store page art, marketing materials. AI used behind the scenes as a development tool, code help, brainstorming, drafting that never reaches the player, no longer requires declaring.
The practical habit costs nothing and pays twice: keep the decisions file from the loop above, and disclosure becomes a five-minute form instead of an archaeology project the night before submission. The visibility politics of carrying an AI label, and whether it helps or hurts a store page, is the subject of the AI-generated games on Steam piece linked earlier.
The loop that works
- Direct. One sentence of intent before the chat opens: what the game needs next, and why. If you cannot write the sentence, the assistant cannot help, and the game is telling you to playtest instead.
- Draft. One bounded task at a time, schema attached, variants requested.
- Review. In the editor, not the chat window. Paste the output into the scene, the sheet, the build, and judge it where players will meet it.
- Own. Touch it up, tune it, log the decision. The pass that makes it yours is the same pass that makes it shippable.
- Ship. Put the build at a link, watch one playtest, and feed what you learned into step one. The assistant amplifies this loop. It is not the loop.
Common mistakes
- Delegating the game instead of the next task. "Build my roguelike" produces an unreviewable pile; "twenty healing items in this schema" produces a table you can use in ten minutes.
- Accepting first outputs. The first draft is the average of everything the model has read; your game is not average, or it should not be.
- Reviewing in the chat window. Assets and logic look right in a chat and wrong in a scene. The editor is the only place a verdict counts.
- No record of what was AI-drafted. You will reconstruct it from memory the week a store form or a copyright question asks, and memory is not a source of truth.
- Letting the assistant invent editor vocabulary. An event that does not exist, confidently proposed and pasted, is a bug with good grammar. Ground every request in the real list.
- Trusting the feeling of speed. The METR developers felt 20 percent faster while being 19 percent slower. The stopwatch is the only reviewer that cannot be flattered.
How Egmatic fits
Egmatic, the no-code 2D editor this blog belongs to, is being built around exactly this split. The editor is AI-native, and the AI inside is a craft amplifier under your direction: you decide what the game needs next, it drafts events, assets, and text inside the same project, and every file stays yours to edit, export, or walk away with. Nothing ships because a model decided it should. Around the editor sits the ship layer this article kept bumping into: the scene you edit is already the build, carried to players at a link with analytics attached, without code and without a backend of your own. We will state the stage plainly: Egmatic is pre-alpha, we announce no dates we cannot keep, and the waitlist at egmatic.com is where you can tell us the first task you would hand to an assistant, because requests arriving during pre-alpha are the ones that steer the build.
Delegate the drafting. Keep the verdicts.
Egmatic is a no-code 2D editor and engine where AI drafts events, assets, and text under your direction, and the ship layer puts the build at a link with analytics. Join the waitlist and tell us the first task you would hand off: requests sent during pre-alpha are the ones that steer the build.
Conclusion
The working division of labor between you and an AI assistant is unglamorous: it drafts against your schemas, you own every verdict. Assets arrive as concept passes you select and repaint, text arrives as variants you edit against a tone document, balance arrives as tables you tune after a playtest, and analysis arrives as leads you verify in the editor. Whole games from one prompt remain a demo genre, feel and judgment remain yours because the assistant never plays, and consistency and ownership both reduce to the same habit: a human pass over everything that ships. Measure your own loop rather than trusting the feeling of speed, keep a decisions file for the day a store form or a copyright question asks what the machine made, and whether you build in code, low-code, or the drag-and-drop editors now carrying top-grossing mobile games, the split holds. The developers shipping with assistants in 2026 are not the ones who found a magic prompt. They are the ones who stayed the director.
Sources
- METR - Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (July 2025): randomized controlled trial, 19 percent slower with AI, confidence interval +2 to +39 percent, predicted 24 percent speedup and perceived ~20 percent speedup
- Peng, Kalliamvakou, Cihon, Demirer - The Impact of AI on Developer Productivity: Evidence from GitHub Copilot: 95 developers, HTTP server task, 71.1 vs 160.9 minutes, 55.8 percent faster
- U.S. Copyright Office - Copyright and Artificial Intelligence, Part 2: Copyrightability (January 29, 2025); release covered by the Library of Congress newsroom
- Steamworks documentation - Content Survey, including the AI disclosure section; policy introduction covered by Video Games Chronicle, January 2024
- Game Developer - Valve tweaks and clarifies AI disclosure rules for Steam (January 16, 2026): disclosure narrowed to player-facing content
- Anthropic - Introducing the Model Context Protocol (November 2024), since placed under the Linux Foundation's Agentic AI Foundation governance
- Wikipedia - Vibe coding: term coined by Andrej Karpathy in February 2025
- Community reads, EGM customer development interviews (EGM-373): the most engaged prospect identified was a developer who built a Model Context Protocol bridge between an AI assistant and Construct 3
Related Posts
AI-Generated Game Flood: How Human-Made Indies Stay Visible
A third of 2026 Steam releases carry an AI disclosure. Here is what the flood takes from you, what it does not, and what still gets human-made games seen.

Best 2D Engine by Genre: Rhythm, Roguelike, Platformer
Rhythm games need audio-clock precision, roguelikes need iteration speed, platformers need deterministic physics. Pick by genre, not by feature list.
Best Game Dev Community Platforms Compared
Reddit brings reach, Discord your core players, engine forums the best answers, itch.io honest feedback on early builds. What each is for, and the cost.