Claude Fable 5.1 in the Arena: Seven Ways to Run It, and What Each Costs
Claude Fable 5.1 runs in this arena as seven variants, from low effort to a multi-agent ultracode mode. Here is what each one costs, how long it takes, and how much code it delivers on the same briefs.
Claude Fable 5.1 is the newest model in Anthropic's Claude 5 family, and it sits in the Mythos class - the capability tier the company positions above Opus. As with the previous generation, there are two names on one model: Fable 5.1 is the generally available version, shipping with additional safety measures around dual-use capabilities, while Mythos 5.1 is the same underlying model without those measures, offered to approved organizations only. If the naming is new to you, the Fable versus Mythos explainer covers it.
What matters here is that Fable 5.1 entered the arena roster not as one entry but as a ladder of seven: fable51-low, fable51-medium, fable51-high, fable51-xhigh, fable51-max, fable51-ultrathink and fable51-ultracode. Same brief, same one-shot contract every other model gets, seven different amounts of thinking. Every one of those cells has a measured output-token count, wall-clock time and cost.
What the seven rungs are
- low, medium, high, xhigh, max - the standard thinking-effort dial, five settings on the same model. More effort means more reasoning tokens before the model starts writing.
- ultrathink - max effort plus the literal word ultrathink at the top of the prompt. It is not a separate setting; it is what a user actually gets when they type that word at max.
- ultracode - not an effort level at all. A multi-agent mode where the model is allowed tools, fans out across the brief, builds, then adversarially verifies its own result against every line of the requirements before finishing.
Twenty-three Fable 5.1 artifacts are live so far, covering four briefs: a landing page, a brick-breaker game, a designer colour tool and an analytics dashboard. Low, medium, high and xhigh ran all four. Max and ultrathink ran three. Ultracode ran one. The full matrix is 43 briefs across 7 rungs, so 23 of 301 runs are done so far - every figure below is drawn from four briefs, which is a sample, not a verdict.
The ladder, measured
Here are the four rungs that ran the complete four-brief set. Cost and wall clock are summed across the four cells; generation tokens are the model's own output count including thinking; delivered size is the arena's file-size proxy for the finished HTML. Cost throughout is the run's tokens priced at Anthropic's published rates.
| Rung | Cost, 4 briefs | Wall clock | Generation tokens | Delivered (~tok) |
|---|---|---|---|---|
| low | $6.86 | 17.5 min | 99,014 | 42,168 |
| medium | $9.34 | 27.6 min | 151,294 | 41,881 |
| high | $12.30 | 38.7 min | 210,510 | 47,591 |
| xhigh | $22.48 | 75.6 min | 413,524 | 65,839 |
Low to xhigh is 3.3x the cost, 4.3x the wall clock and 4.2x the generation tokens, for 1.56x the delivered file. The medium rung is the one worth staring at: it cost 36% more than low and delivered 0.7% less code across the same four briefs. That is not one bad cell dragging an average. On the brick-breaker brief specifically, low produced ~8,903 delivered tokens for $1.37 and medium produced ~6,840 for $2.60 - nearly double the money for roughly a quarter less output.
Where the curve breaks
On the landing-page brief, the bottom three rungs are close to indistinguishable. Low: 27,403 generation tokens, 293 seconds, $2.17. Medium: 26,521 tokens, 295 seconds, $2.10. High: 29,690 tokens, 326 seconds, $2.25. Three effort levels inside 12% of each other on tokens, 11% on time and 7% on cost. Then xhigh: 95,653 tokens, 1,080 seconds, $5.58. The dial does close to nothing for three clicks and then more than doubles the bill on the fourth.
The designer-tool brief is the counter-example, and it is the cleanest ladder in the set: $1.26, $1.42, $2.90, $5.31 for low, medium, high and xhigh, with delivered size climbing the whole way, ~9,732 to ~18,055. Same model, same day, same one-shot prompt. Whether effort buys anything depends on the brief, not on the setting - which is the empirical version of a point I made before about thinking-effort levels.
The top of the ladder: max and ultrathink
Three briefs ran at xhigh, max and ultrathink, so those three rungs compare directly.
| Rung | Cost, 3 briefs | Wall clock | Generation tokens | Delivered (~tok) |
|---|---|---|---|---|
| xhigh | $17.17 | 58.0 min | 312,701 | 47,784 |
| max | $39.69 | 106.5 min | 540,745 | 53,535 |
| ultrathink | $34.36 | 98.6 min | 496,254 | 53,065 |
Max costs 2.3x what xhigh costs and takes 1.8x as long, for 12% more delivered code. Ultrathink and max land within 0.9% of each other on delivered size while ultrathink runs 13% cheaper, which is about what you would expect from two labels on the same effort level plus ordinary run-to-run variance. In practice they are the same rung.
The most expensive one-shot cell in the grid is the brick-breaker at max: 198,799 generation tokens, 2,287 seconds, $16.39. That same cell is also the weakest value per delivered token, at about $1.02 per thousand. The dashboard at ultrathink is close behind: $12.12 for a ~12,085-token file, which is 1.95x the cost of the same brief at xhigh for 9% less output. Across the whole grid, cost per thousand delivered tokens ranges from about $0.13 at the cheap end to $1.02 at the expensive end - a near eightfold spread on the same model.
What ultracode costs
The multi-agent mode is not an effort setting, and it is not priced like one. The single ultracode cell, on the landing-page brief, took 3,728 seconds of wall clock and cost $82.94. It has no generation-token figure of its own, because the work happens across sub-agents rather than in one reply.
The file it delivered is ~19,467 tokens, which is 16% smaller than what the same brief produced at xhigh for $5.58. Per thousand delivered tokens that works out to roughly $4.26, against about $0.13 for the cheapest one-shot cell in the grid. Whether the extra verification pass makes the file better is exactly the kind of question a blind vote answers and a spend curve does not.
How Fable 5.1 sizes up against the rest of the roster
Cost figures are not available for every Claude family on the same file, but delivered size is, and it is recorded for every variant. Below is the largest one-shot artifact each family produced on each brief, with the rung that produced it. Multi-agent runs and the earlier pre-ban Fable 5 entry are excluded, because they are not the same kind of run.
| Brief | Fable 5.1 | Fable 5 | Opus 5 | Opus 4.8 |
|---|---|---|---|---|
| Landing | 23,618 (ultrathink) | 21,072 (high) | 20,292 (high) | 18,884 (high) |
| Game | 17,362 (ultrathink) | 15,724 (xhigh) | 16,379 (low) | 13,771 (medium) |
| Tool | 18,055 (xhigh) | 18,650 (max) | 21,009 (medium) | 19,435 (xhigh) |
| Dashboard | 16,496 (max) | 17,981 (xhigh) | 15,876 (low) | 13,637 (xhigh) |
Fable 5.1 ships the largest artifact on two of the four briefs and loses on the other two. Look at which rung won for each family and the earlier point repeats itself: Opus 5's biggest game file came from low effort, and its biggest tool file from medium. Size does not track effort, and size was never quality anyway - a bigger file can mean more implemented features or more dead markup, and the only way to tell is to open both and look. Fable 5.1 is listed here at the same $50 per million output tokens as Fable 5, twice the Opus 5 rate.
Fable 5.1 has only just started collecting blind votes. Its seven rungs sit on the leaderboard at the 1500 starting rating with a 0-0-0 record, so there are no standings to report yet. Anyone quoting a Fable 5.1 arena ranking today is quoting an empty table.
What I would run
On this sample of four briefs: low is the default, medium is hard to justify on any of them, and high only separates from low on the tool brief. Xhigh is the one genuine step up and it costs about 3x low. Max and ultrathink are the same rung wearing different labels, and they buy roughly 12% more delivered code for 2.3x the price of xhigh. Ultracode is a mode to reach for when a brief needs the verification pass, not a default, at $82.94 and an hour of wall clock for a file smaller than the $5.58 one.
None of that is a quality claim. It is a spend curve. Check the quality yourself: open the coding arena, put Fable 5.1 low against ultrathink on the landing brief, vote blind, and see whether you can pick the $11.64 run out of a lineup. The leaderboard will carry the tally once there is one worth showing. If you want the wider context on this family, start with the Fable 5 report.
Frequently asked questions
What is Claude Fable 5.1?
Claude Fable 5.1 is the newest model in Anthropic's Claude 5 family, in the Mythos class, the capability tier Anthropic positions above Opus. Fable 5.1 and Mythos 5.1 are the same underlying model. Fable 5.1 is the generally available version and ships with additional safety measures for dual-use capabilities; Mythos 5.1 is the version without those measures and is offered to approved organizations only.
What effort levels does Claude Fable 5.1 have in the arena?
Seven variants: fable51-low, fable51-medium, fable51-high, fable51-xhigh, fable51-max, fable51-ultrathink and fable51-ultracode. Ultrathink is max effort with the literal word ultrathink added to the prompt, not a separate setting. Ultracode is a multi-agent mode where the model gets tools and verifies its own work against the brief, instead of answering in a single one-shot message.
Does higher thinking effort make Claude Fable 5.1 produce more code?
Not reliably. Across the four briefs measured, going from low to xhigh multiplied cost by 3.3x and wall-clock time by 4.3x while the delivered file grew only 1.56x. The medium rung cost 36% more than low and delivered slightly less. On one brief, the designer tool, delivered size did climb steadily with effort. It depends on the brief.
How much does one Claude Fable 5.1 run cost?
On the four briefs measured, single one-shot cells ranged from $1.26 at low effort to $16.39 at max, with the run's tokens priced at Anthropic's published rates. The multi-agent ultracode cell cost $82.94 and took 3,728 seconds. Fable 5.1 is listed at $50 per million output tokens, the same as Fable 5.
Is Claude Fable 5.1 better than Fable 5?
There is no answer from this data yet. Fable 5.1 has only just started collecting blind votes and still sits at the 1500 starting rating with no decided matchups, so there are no vote results to cite. On delivered file size across four shared briefs, Fable 5.1 ships the larger artifact on two and the smaller on two, and file size is not a quality measure.
Don’t take the post’s word for it
The arena runs every model’s real output live. Pick a challenge, go blind, and cast a vote that counts in the public tally.
Open the arena