An AI-assisted game project needs more than a growing list of completed tasks. It needs a reliable way to decide whether each change made the playable game better. A script can compile while a jump feels worse, a menu becomes unusable on a phone, or a previously working save stops loading.
This workflow treats generated changes as proposals that need evidence. It works for a small Unity prototype or a browser game, and it does not depend on a particular coding tool. The useful output is a tested build plus an understandable decision record, not a confident completion message.
/01
Capture a playable baseline before assigning work
Choose one short route through the current game: launch, enter a level, use the main mechanic, fail or finish, and restart. Record the build identifier, device, input method, and any existing problems. Keep the previous working build available. If the starting point does not run, repairing that baseline is the first task; adding features will only make the diagnosis less clear.
For an imagined arena game, the baseline might be a two-minute match with one enemy type. The initial evidence could be a screen recording and a short note that restarting sometimes duplicates the player. That is a useful starting point because the next change can be judged against the same route. It is an example of a review method, not a claimed benchmark.
/02
Give each change a small, observable contract
A request such as improve movement contains too many possible interpretations. Specify the player problem, the permitted files or systems, the expected behavior, and the checks that will demonstrate it. Separate required work from ideas that should remain untouched. A task should be small enough that one reviewer can understand why every changed file is present.
For example, ask for a configurable jump buffer while preserving the existing landing behavior. Require tests for a press just before landing, a press outside the buffer window, and holding the button across a restart. This makes it possible to reject an implementation that looks plausible but changes an unrelated mechanic.
- Name the player-visible problem and one acceptance scenario.
- List the systems that must not change during this task.
- Require a before-and-after comparison using the same build route.
/03
Review the change as a fresh task
After implementation, inspect the diff before opening a new feature request. Ask whether configuration files, assets, dependencies, and generated files were changed intentionally. Pay special attention to a fix that disables a check, removes an error path, or silently substitutes sample data. Those changes can make a test appear successful without improving the game.
Use automated tests for rules that can be asserted repeatedly, such as score transitions or save migration. Use a playable build for timing, readability, input, and the overall loop. Neither replaces the other. A review record should say exactly which checks ran, which failed, and which still require a device or account that was unavailable.
/04
Measure performance before asking for optimization
Attach a representative capture to a performance task instead of guessing which subsystem is slow. Unity's Profiler can collect application data from the Editor or a target device. For browser games, Chrome's Performance panel provides a recorded timeline to inspect activity during a specific interaction. Select the tool that matches the runtime you are actually shipping.
Set the experiment around a visible moment: entering a crowded room, opening inventory, or restarting after defeat. Repeat that moment on the same device and record the environment. Change one suspected cause, then compare the result. Keep an optimization only when its measured benefit and maintenance cost make sense for your target players.
/05
End the loop with a clear release decision
Do a short regression pass on the original route after the new acceptance scenario passes. Check that the first session, repeat session, failure path, and restart still make sense. Keep unresolved issues visible with an owner and a concrete next check. Avoid turning a long list of minor improvements into a reason to overlook one broken essential action.
The final note can be simple: what changed, what evidence supports it, what remains uncertain, and how to return to the prior build. AgentGuild's Games Kit provides focused audit, playtest, and release workflows for organizing this kind of work. The kit does not replace running the game or making the final shipping decision.
Keep this in mind
Key takeaways
- Start every work cycle from a known playable build and a repeatable player route.
- Keep implementation tasks bounded and review the resulting diff independently.
- Combine automated checks, measured runtime captures, and hands-on playtesting before release.
Sources & further reading
Consult the source documentation for the version and platform you are working with.