One prompt. Every model. Judge it yourself.
Real outputs running live in your browser — not screenshots, not cherry-picked demos. Judge blind, vote, and the community tally is public.
54 challenges · 131 model variants across 21 families · 4,026 live artifacts · $309 estimated output spend, every prompt published. Counted at build, 2026-07-17.
All challenges
Every brief we have run, ranked by community votes. Click a row to judge it yourself — blind.
| Task | Category | Models | Leader | Votes | Est. cost range |
|---|---|---|---|---|---|
| 1 · Landing · Driftwood | web & tools | 131 | … | 0 | $0.00–1.10 |
| 2 · Game · NEON BREAKER | arcade & games | 126 | … | 0 | $0.00–0.79 |
| 3 · Tool · Palette Studio | web & tools | 127 | … | 0 | $0.00–0.93 |
| 4 · Dashboard · Bean There | web & tools | 125 | … | 0 | $0.00–0.90 |
| 5 · Particles · Gravity Lab | simulation | 119 | … | 0 | $0.00–0.66 |
| 6 · 3D Game · STARDRIFT | 3d & godot | 113 | … | 0 | $0.00–0.78 |
| 7 · CHIP-8 · CHIP-8 emulator | web & tools | 116 | … | 0 | $0.00–0.52 |
| 8 · News · The Meridian Wire | web & tools | 118 | … | 0 | $0.00–1.35 |
| 9 · Marble 3D · godforge | 3d & godot | 2 | … | 0 | $0.00–0.19 |
| 10 · Runner 3D · godforge | 3d & godot | 2 | … | 0 | — |
| 11 · FPS 3D · godforge | 3d & godot | 2 | … | 0 | — |
| 12 · Digger · DEEP HAUL | arcade & games | 117 | … | 0 | $0.00–0.95 |
Six arenas, one method
Coding and writing are live and interactive right now. Images, videos, voice, and music launch as deep evaluation guides while their side-by-side arenas come online.
Coding
Live nowOne brief, every model, real code running live — landing pages, arcade games, 3D, emulators.
43 tasks · 131 model variants
Writing
Live nowOne brief, every model, complete written pieces side by side — essays, stories, songs, satire.
11 briefs · blind judging
Images
GuideHow to judge AI image generators — prompt adherence, text, hands, style range, artifacts.
evaluation guide · arena in production
Videos
GuideWhat separates good text-to-video from bad — motion, temporal consistency, prompt control.
evaluation guide · arena in production
Voice
GuideHow to judge AI voice models — prosody on long form, emotional control, hard pronunciations.
rubric live · audio arena next
Music
GuideHow to judge AI music generators — composition, fidelity, vocals, structure, prompt adherence.
first takes, not curated demos
How the arena works
One prompt, every model
We write a single brief and hand the exact same words to every variant. No per-model tuning, no quiet retries to make a favorite look good. The full prompt is published on every challenge.
Real outputs, running live
Each answer runs in your browser as a live artifact, an actual playable game or interactive page, not a screenshot or a marketing clip. Token and cost estimates sit next to every result.
Blind judging, public votes
You start blind: labels hidden, panes shuffled. Vote and the reveal shows who you picked — and whether the crowd agrees. Every vote rolls into a public community tally.
Full methodology, cost math and changelog: /method
The models on the stand
21 families and 131 variants, most at several thinking-effort levels — so you can see what extra reasoning actually buys.
Pick a challenge. Judge for yourself.
54 challenges, 131 model variants, and real outputs you can poke at. Go blind, compare, and add your vote.
Open the coding arena