THE AI MODEL ARENA

One prompt. Every model. Judge it yourself.

Real outputs running live in your browser — not screenshots, not cherry-picked demos. Judge blind, vote, and the community tally is public.

54 challenges · 139 model variants across 23 families · 4,414 live artifacts · $436 estimated output spend, every prompt published. Counted at build, 2026-07-25.

All challenges

Every brief we have run, ranked by community votes. Click a row to judge it yourself — blind.

TaskCategoryModelsLeaderVotesEst. cost range
1 · Landingweb & tools1390$0.00–1.10
2 · Gamearcade & games1330$0.00–0.79
3 · Toolweb & tools1350$0.00–1.40
4 · Dashboardweb & tools1320$0.00–0.90
5 · Particlessimulation1270$0.00–0.66
6 · 3D Game3d & godot1200$0.00–1.77
7 · CHIP-8web & tools1230$0.00–0.52
8 · Newsweb & tools1240$0.00–1.35
9 · Marble 3D3d & godot20$0.00–0.19
10 · Runner 3D3d & godot20
11 · FPS 3D3d & godot20
12 · Diggerarcade & games1230$0.00–0.95

Six arenas, one method

Coding and writing are live and interactive right now. Images, videos, voice, and music launch as deep evaluation guides while their side-by-side arenas come online.

How the arena works

01

One prompt, every model

We write a single brief and hand the exact same words to every variant. No per-model tuning, no quiet retries to make a favorite look good. The full prompt is published on every challenge.

02

Real outputs, running live

Each answer runs in your browser as a live artifact, an actual playable game or interactive page, not a screenshot or a marketing clip. Token and cost estimates sit next to every result.

03

Blind judging, public votes

You start blind: labels hidden, panes shuffled. Vote and the reveal shows who you picked — and whether the crowd agrees. Every vote rolls into a public community tally.

Full methodology, cost math and changelog: /method

The models on the stand

23 families and 139 variants, most at several thinking-effort levels — so you can see what extra reasoning actually buys.

Fable 5Opus 4.8Opus 5Sonnet 4.6Sonnet 5Haiku 4.5GLM-5.2GPT-5.5Gemini 3 FlashKimi K2.7 CodeQwen3.7 PlusDeepSeek V4 FlashMiniMax M3Mistral Large 2512Grok 4.3Seed 2.1 ProLongCat 2.0KAT-Coder-Pro V2MiMo-V2.5-ProHunyuan Hy3Nemotron 3 Ultragpt-oss-120bGemma 4 26B-A4BLlama 4 MaverickStep 3.7 FlashMuse Spark 1.1

Pick a challenge. Judge for yourself.

54 challenges, 139 model variants, and real outputs you can poke at. Go blind, compare, and add your vote.

Open the coding arena