Codex Build Arena - See what more Codex thinking actually buys you

by
Codex Build Arena is an open, visual benchmark of complete apps produced from the exact same one-shot prompt and starter repository. Compare GPT-5.6 Luna, Terra, and Sol across effort levels through live builds, side-by-side previews, engineering reviews, token usage, credits, duration, and source size. No cherry-picking, follow-up prompts, or hidden repairs. Use the results and decide which model and effort earned the cost.

Add a comment

Replies

Best
Hey Product Hunt 👋 On my first day using Codex, I had a basic problem: I could choose a model and an effort level, but I could not see what those choices actually bought me. I was mostly hoping that newer meant better and more effort meant better work. So I built Codex Build Arena. Every captured run receives the same one-shot product brief, the same clean starter repo, the same tools, and no follow-up guidance. The final app is published as a live interactive build, together with its tokens, estimated credits, duration, source size, and a frozen engineering review. The current arena includes GPT-5.6 Luna, Terra, and Sol, plus GPT-5.5, across Light, Medium, High, and Extra High effort. You can explore one build, compare two full screen, or analyze quality against cost. A few things surprised me: • More effort did not automatically produce a better result. • The most expensive run was not always the strongest. • Some lower-cost configurations were excellent starting points. • One output could use fewer total tokens while producing more finished source. The project is open source and all benchmark outputs are committed so anyone can inspect the evidence. I would love to know: which build would you keep, and what should the next one-prompt benchmark test? Built in public with Codex.
What did GPT-5.6 Sol, Terra and Luna unlock that made your launch possible?
GPT-5.6 Sol, Terra, and Luna made the benchmark itself possible. I could give each model the same one-shot product brief and starter repository, then capture complete working builds instead of isolated code snippets. Their different speed, reasoning, and output profiles created a real arena: Luna showed how far a fast run could go, Terra balanced quality and cost, and Sol tackled the most demanding decisions. Codex also helped me build the public comparison experience, run repeatable reviews, diagnose UI issues, and publish every artifact openly. Without three capable models with distinct tradeoffs, this would be another subjective model claim. With them, people can inspect the live apps, engineering scores, tokens, credits, and duration, then decide which level of intelligence and effort actually earns its cost.