Elo · Cost · Speed

Model leaderboard

Every generated entry, aggregated per model variant and ranked by Elo rating from blind pairwise votes, alongside community votes, output volume, estimated cost per task, and measured generation time. Pick an arena, click a column to sort.

Wondering what these numbers say about the best AI for coding? Read the full breakdown →

Model
Muse Spark 1.1 · Thinking1500?0-0-005~10,413$0.04
Fable 5 · low1500?0-0-0034~11,240$0.5620m 18s
Fable 5 · medium1500?0-0-0020~13,622$0.6826m 06s
Fable 5 · high1500?0-0-0032~13,137$0.6632m 10s
Fable 5 · xhigh1500?0-0-0032~12,646$0.6333m 36s
Fable 5 · max (post-ban)1500?0-0-0019~14,612$0.7334m 34s
Fable 5 (pre-ban)1500?0-0-0018~15,339$0.77
Opus 4.8 · low1500?0-0-0040~13,851$0.35
Opus 4.8 · medium1500?0-0-0040~14,769$0.37
Opus 4.8 · high1500?0-0-0040~16,804$0.42
Opus 4.8 · xhigh1500?0-0-0040~16,503$0.41
Opus 4.8 · max1500?0-0-0040~16,891$0.42
Opus 5 · low1500?0-0-0040~15,777$0.39
Opus 5 · medium1500?0-0-0040~16,001$0.40
Opus 5 · high1500?0-0-0040~15,541$0.39
Opus 5 · xhigh1500?0-0-0040~16,030$0.40
Opus 5 · max1500?0-0-0040~16,374$0.41
Opus 5 · ultracode1500?0-0-0016~29,278$0.73
Sonnet 4.6 · low1500?0-0-0040~9,603$0.14
Sonnet 4.6 · medium1500?0-0-0040~10,308$0.15
Sonnet 4.6 · high1500?0-0-0040~11,062$0.17
Sonnet 4.6 · max1500?0-0-0040~11,290$0.17
Sonnet 5 · low1500?0-0-0011~10,154$0.15
Sonnet 5 · medium1500?0-0-007~13,008$0.20
Sonnet 5 · high1500?0-0-0010~12,128$0.18
Sonnet 5 · xhigh1500?0-0-008~12,583$0.19
Sonnet 5 · max1500?0-0-008~13,260$0.20
Haiku 4.51500?0-0-0040~7,235$0.04
GLM-5.21500?0-0-0019~11,407$0.05
GLM-5.2 · flash1500?0-0-0040~12,940$0.06
GLM-5.11500?0-0-0010~12,764$0.06
GLM-51500?0-0-0010~15,523$0.07
GLM-5 Turbo1500?0-0-0033~8,632$0.04
GLM-4.71500?0-0-0040~9,471$0.04
GLM-4.61500?0-0-0040~8,976$0.04
GLM-4.51500?0-0-0040~10,441$0.05
GLM-4.5 Air1500?0-0-0040~8,936$0.04
GPT-5.51500?0-0-0040~6,002$0.06
Gemini 3 Flash1500?0-0-0040~7,137$0.02
Kimi K2.7 Code1500?0-0-0040~7,868$0.03
Qwen3.7 Plus1500?0-0-0040~7,373$0.01
DeepSeek V4 Flash1500?0-0-0040~14,960$0.00
MiniMax M31500?0-0-0038~6,447$0.01
Kimi K2.61500?0-0-0040~11,664$0.05
Kimi K31500?0-0-0036~8,522$0.03
Qwen3.7 Max1500?0-0-0040~8,156$0.01
Qwen3.8 Max1500?0-0-0035~12,513$0.02
DeepSeek V4 Pro1500?0-0-0040~16,903$0.00
MiniMax M2.71500?0-0-0040~10,658$0.01
GPT-5.5 · low1500?0-0-0040~4,567$0.05
GPT-5.5 · medium1500?0-0-0040~5,031$0.05
GPT-5.5 · high1500?0-0-0040~6,266$0.06
GPT-5.5 · xhigh1500?0-0-0040~8,413$0.08
GPT-5.6 Luna · low1500?0-0-0040~3,569$0.04
GPT-5.6 Luna · max1500?0-0-0040~12,188$0.12
GPT-5.6 Terra · low1500?0-0-0040~3,904$0.04
GPT-5.6 Terra · max1500?0-0-0040~9,746$0.10
GPT-5.6 Sol · low1500?0-0-0040~5,329$0.05
GPT-5.6 Sol · max1500?0-0-0040~12,092$0.12
GPT-5.6 Luna · ultra1500?0-0-0040~16,692$0.17
GPT-5.6 Terra · ultra1500?0-0-0040~18,160$0.18
GPT-5.6 Sol · ultra1500?0-0-0040~20,575$0.21
GPT-5.6 Luna Pro1500?0-0-001~6,246$0.06
GPT-5.6 Luna · medium1500?0-0-0040~4,650$0.05
GPT-5.6 Luna · high1500?0-0-0040~8,501$0.09
GPT-5.6 Terra · medium1500?0-0-0040~4,105$0.04
GPT-5.6 Terra · high1500?0-0-0040~4,950$0.05
GPT-5.6 Sol · medium1500?0-0-0040~7,536$0.08
GPT-5.6 Sol · high1500?0-0-0040~10,704$0.11
Mistral Large 25121500?0-0-0040~11,170$0.02
Mistral Medium 3.51500?0-0-0040~9,347$0.02
Mistral Medium 3 (2508)1500?0-0-0039~17,512$0.04
Mistral Medium (2505)1500?0-0-0039~8,534$0.02
Mistral Small 26031500?0-0-0040~6,386$0.01
Mistral Small 25061500?0-0-0039~5,542$0.01
Codestral 25081500?0-0-0040~7,946$0.02
Devstral 25121500?0-0-0039~10,327$0.02
Magistral Medium · reasoning1500?0-0-0010~6,892$0.01
Magistral Medium · no-reason1500?0-0-0039~8,877$0.02
Magistral Small · reasoning1500?0-0-0040~6,119$0.01
Magistral Small · no-reason1500?0-0-0040~5,536$0.01
Ministral 14B1500?0-0-0039~14,725$0.03
Ministral 8B1500?0-0-0037~12,817$0.03
Ministral 3B1500?0-0-0038~11,777$0.02
Mistral Nemo1500?0-0-0040~1,707$0.00
Grok 4.31500?0-0-0040~3,338$0.01
Grok 4.20 · non-reasoning1500?0-0-0040~11,924$0.03
Grok 4.20 · reasoning1500?0-0-0039~11,550$0.03
Grok 4.5 · high1500?0-0-0039~11,422$0.03
Grok 4.5 · low1500?0-0-0038~8,915$0.02
GLM-5.2 · OR fp81500?0-0-0018~9,913$0.04
GLM-5.1 · OR fp81500?0-0-0018~10,210$0.04
GLM-5 · OR fp81500?0-0-001~12,716$0.06
GLM-4.7 · OR fp81500?0-0-0020~9,770$0.04
GLM-4.7 Flash · OR fp81500?0-0-0018~8,716$0.04
GLM-4.5 Air · OR fp81500?0-0-0019~11,441$0.05
DeepSeek V4 Flash · OR fp81500?0-0-0020~16,445$0.00
DeepSeek V4 Pro · OR fp81500?0-0-0020~15,988$0.00
DeepSeek V3.1 · OR fp81500?0-0-0020~7,119$0.00
DeepSeek R1-0528 · OR fp81500?0-0-0020~11,440$0.00
Kimi K2.5 · OR fp81500?0-0-0020~8,958$0.04
Kimi K2.6 · OR fp81500?0-0-0019~11,931$0.05
Kimi K2.7 Code · OR fp81500?0-0-0018~8,965$0.04
Kimi K2-0905 · OR fp81500?0-0-004~10,882$0.04
MiniMax M2.7 · OR fp81500?0-0-0020~10,883$0.01
MiniMax M2.5 · OR fp81500?0-0-004~13,211$0.02
MiniMax M3 · OR fp81500?0-0-007~13,385$0.02
MiniMax M2.1 · OR fp81500?0-0-004~11,294$0.01
MiniMax M2 · OR fp81500?0-0-004~11,771$0.01
Qwen3.6 27B · OR fp81500?0-0-0019~9,642$0.02
Qwen3.6 35B-A3B · OR fp81500?0-0-0020~10,059$0.02
Qwen3.5 122B-A10B · OR fp81500?0-0-0020~7,611$0.01
Qwen3.5 397B-A17B · OR fp81500?0-0-0020~8,165$0.01
Qwen3.5 27B · OR fp81500?0-0-0019~7,643$0.01
Qwen3.5 9B · OR fp81500?0-0-0019~6,802$0.01
Qwen3.5 35B-A3B · OR fp81500?0-0-0020~8,688$0.01
Qwen3 Coder Next · OR fp81500?0-0-0020~9,954$0.02
Qwen3 Coder 480B · OR fp81500?0-0-0020~7,619$0.01
Qwen3 Coder 30B · OR fp81500?0-0-0017~7,835$0.01
Qwen3 235B-2507 · OR fp81500?0-0-0020~11,653$0.02
Qwen3 32B · OR fp81500?0-0-0018~5,788$0.01
Qwen3 30B-A3B · OR fp81500?0-0-0020~11,637$0.02
gpt-oss-120b · OR fp81500?0-0-0020~3,767$0.00
gpt-oss-20b · OR fp81500?0-0-0020~2,798$0.00
Gemma 4 26B-A4B · OR fp81500?0-0-0020~5,836$0.00
Gemma 4 31B · OR fp81500?0-0-0018~5,571$0.00
Gemma 3 27B · OR fp81500?0-0-0020~2,954$0.00
Llama 4 Maverick · OR fp81500?0-0-0020~3,194$0.00
Llama 4 Scout · OR fp81500?0-0-0020~2,376$0.00
Llama 3.3 70B · OR fp81500?0-0-0020~2,922$0.00
Llama 3.1 8B · OR fp81500?0-0-0015~5,576$0.00
MiMo V2.5 · OR fp81500?0-0-0014~9,344$0.01
MiMo V2.5 Pro · OR fp81500?0-0-005~11,484$0.01
Step 3.7 Flash · OR fp81500?0-0-0018~10,897$0.02
Nemotron 3 Super 120B · OR fp81500?0-0-0018~4,905$0.01
Nemotron 3 Nano 30B · OR fp81500?0-0-0017~3,643$0.01
Mistral Small 3.2 · OR fp81500?0-0-0019~5,856$0.01
Ministral 14B · OR fp81500?0-0-0016~13,755$0.03
Hunyuan A13B · OR fp81500?0-0-0015~3,326$0.00

Elo ratings replay the public pairwise match log: every blind vote records the models that were on screen, the winner beats each shown opponent (ties count 0.5), starting at 1500 with K 32 for a variant's first 30 matches and 16 after. One vote carries the same total weight no matter how many panes were compared. Variants with no rated matches yet display the 1500 baseline. Ratings are per arena and started fresh on 2026-07-10; earlier votes lacked opponent data and stay in the raw vote counts only. Token counts are estimated from artifact file size (chars ÷ 4). Cost = estimated output tokens × the model's published per-1M output price; input and reasoning tokens aren't counted, so true cost is higher, especially at high thinking effort. Generation time is wall-clock and only recorded for runs generated after 2026-07-02; it includes queue and throttling waits. This page compares like-for-like one-shot generations, not lab benchmarks.