AI can generate a playable game. But can it prove the game works?

by•

Most AI game tools optimize for the moment a build launches. That is not the same as proving a game is shippable.

A generated Godot project can compile and still fail because:

- the first-run window is blank,

- keyboard focus is wrong,

- controller input works on one OS but not another,

- the core loop stalls after an unexpected state,

- a repair fixes one path and breaks the next iteration.

This is the problem we are tackling with DeviLudo. After the Design and Development agents produce an exported game, the Test agent does not inspect source or trust the game's own success flags. It operates the real window through OS-level keyboard, pointer, and virtual gamepad input on macOS, Windows, and Linux.

Each target runs deterministic journeys plus three adaptive playthroughs. A separate read-only oracle checks state and progress; videos, screenshots, action traces, visual diffs, and the shortest successful path become replayable regression evidence. Product failures can return to the agent pipeline for bounded repair before SteamPipe delivery.

The question I keep coming back to: for AI-generated software, should “working” mean it compiled, it opened, or an independent agent completed the core user loop in the target environments?

Game developers: which failure do you see most often only after export—input, timing, rendering, or state progression?

6 views

Add a comment

Replies

Be the first to comment