All news
·via nowosci.ai

Anthropic Ships Claude Opus 5 at Half the Price of Fable 5

Opus 5 lands at unchanged $5/$25 per million tokens and takes the state of the art on Frontier-Bench and GDPval-AA. It becomes the default model on Claude Max and the strongest model on Claude Pro.

Introducing Claude Opus 5: a numeral five composed of illustrated birds eggs on a beige field

Anthropic released Claude Opus 5 today. The pitch is not a new ceiling but a new price per unit of intelligence: Anthropic says Opus 5 comes close to the frontier capability of Fable 5 at half the cost, and it holds the Opus 4.8 rate card exactly, at $5 per million input tokens and $25 per million output tokens.

It is available now on the Claude API, Claude.ai, Claude Code and Claude Cowork. It is also the new default model on Claude Max and the strongest model offered on Claude Pro. Fast mode returns at twice the base price for roughly 2.5x the default speed.

Where it lands on the benchmarks

BenchmarkResult
Frontier-Bench v0.1State of the art, 2x Opus 4.8
CursorBench 3.2 (max)Within 0.5% of Fable 5, at half the cost
ARC-AGI 33x the next-best model
Zapier AutomationBench~1.5x the pass rate of the next-best model
OSWorld 2.0Beats every model, at 1/3 the cost of Fable 5
GDPval-AA v2, DeepSearchQABest and most cost-efficient

The life-sciences gains are the cleanest year-over-year numbers in the announcement: +10.2 points on organic chemistry tasks and +7.7 points on protein sequence prediction against Opus 4.8. On applied work Anthropic reports first-turn legal redlines at nearly double the Opus 4.8 score, and financial modelling at 9 points higher accuracy, 60% less time and one third fewer tool calls.

Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost - Sualeh Asif, co-founder, Cursor

The catch: cybersecurity

Anthropic is explicit that Opus 5 stays behind Mythos 5 on cybersecurity tasks. On OSS-Fuzz it identifies vulnerabilities at near parity with Mythos 5, but is "considerably less successful" at building working exploits. The company frames this as deliberate: the model "does not advance the frontier in risky, dual-use capabilities".

The safeguards moved the other way for defenders. Cyber classifiers are about 85% less restrictive than on Fable 5, so source-code vulnerability research is allowed while binary scanning, penetration testing and exploit generation stay blocked. Flagged requests fall back to Opus 4.8, and enterprises in the Cyber Verification Program get a less-restricted build.

Two API changes worth noting

  • Mid-conversation tool changes without invalidating the prompt cache, which cuts the running cost of long agentic sessions that swap toolsets.
  • Automatic fallbacks: requests flagged by safety classifiers on Opus 5 or Fable 5 can route to another model instead of returning a refusal.

On alignment, Anthropic reports a behavioural audit score of 2.3, the lowest and therefore best among its recent models, with the lowest rate of deceptive behaviour of the current lineup.

In the arena

Every number above is vendor-reported. Opus 5 runs are queued for the coding challenges now, and the cost-per-task column is the one to watch: a model that is 0.5% behind Fable 5 at half the price changes which tier you should be routing work to.

Based on: Anthropic · nowosci.ai
More AI news in Polish at nowosci.ai