Launching today

Pretty Bench
Human-ranked arenas for AI-generated work
5 followers
Human-ranked arenas for AI-generated work
5 followers
Pretty Bench is a public, human-preference benchmark for AI-generated images, video, text, and other creative work. Put models head to head, collect real votes, and see which systems people actually prefer. Explore transparent rankings, compare model outputs, and help build a more useful picture of AI quality beyond benchmark scores.









I used GPT-6 Astra as an active engineering partner while building Pretty Bench. It helped with site design and UX, model-ingestion scripts, image and video asset processing, Supabase database and storage integration, and the operational work needed to populate and verify the benchmark. The real task was taking a multi-model creative benchmark from raw assets to a working public product, with human voting, rankings, and traceable model provenance. Astra was useful across the full loop, not just for generating snippets of code.
Pretty Bench’s primary image arena currently compares 10 image models across 4,338 human votes and 338 sessions. Instead of relying on benchmark scores alone, we put model outputs side by side and let people choose what they actually prefer. We’re expanding the same format to video and other creative work, with model provenance and transparent rankings. If you’ve tried image models recently, vote on a few matchups. Your preferences help make the leaderboard more useful. Explore the arena at prettybench.org.