THE AI MODEL ARENA

One prompt. Every model. Judge it yourself.

Real outputs running live in your browser — not screenshots, not cherry-picked demos. Judge blind, vote, and the community tally is public.

54 challenges · 131 model variants across 21 families · 4,026 live artifacts · $309 estimated output spend, every prompt published. Counted at build, 2026-07-17.

All challenges

Every brief we have run, ranked by community votes. Click a row to judge it yourself — blind.

TaskCategoryModelsLeaderVotesEst. cost range
1 · Landingweb & tools1310$0.00–1.10
2 · Gamearcade & games1260$0.00–0.79
3 · Toolweb & tools1270$0.00–0.93
4 · Dashboardweb & tools1250$0.00–0.90
5 · Particlessimulation1190$0.00–0.66
6 · 3D Game3d & godot1130$0.00–0.78
7 · CHIP-8web & tools1160$0.00–0.52
8 · Newsweb & tools1180$0.00–1.35
9 · Marble 3D3d & godot20$0.00–0.19
10 · Runner 3D3d & godot20
11 · FPS 3D3d & godot20
12 · Diggerarcade & games1170$0.00–0.95

Six arenas, one method

Coding and writing are live and interactive right now. Images, videos, voice, and music launch as deep evaluation guides while their side-by-side arenas come online.

How the arena works

01

One prompt, every model

We write a single brief and hand the exact same words to every variant. No per-model tuning, no quiet retries to make a favorite look good. The full prompt is published on every challenge.

02

Real outputs, running live

Each answer runs in your browser as a live artifact, an actual playable game or interactive page, not a screenshot or a marketing clip. Token and cost estimates sit next to every result.

03

Blind judging, public votes

You start blind: labels hidden, panes shuffled. Vote and the reveal shows who you picked — and whether the crowd agrees. Every vote rolls into a public community tally.

Full methodology, cost math and changelog: /method

The models on the stand

21 families and 131 variants, most at several thinking-effort levels — so you can see what extra reasoning actually buys.

Fable 5Opus 4.8Sonnet 4.6Sonnet 5Haiku 4.5GLM-5.2GPT-5.5Gemini 3 FlashKimi K2.7 CodeQwen3.7 PlusDeepSeek V4 FlashMiniMax M3Mistral Large 2512Grok 4.3Seed 2.1 ProLongCat 2.0KAT-Coder-Pro V2MiMo-V2.5-ProHunyuan Hy3Nemotron 3 Ultragpt-oss-120bGemma 4 26B-A4BLlama 4 MaverickStep 3.7 FlashMuse Spark 1.1

Pick a challenge. Judge for yourself.

54 challenges, 131 model variants, and real outputs you can poke at. Go blind, compare, and add your vote.

Open the coding arena