AI

Claude Opus 5: Frontier Performance at Half the Price

August 2, 2026
Claude Opus 5 illustration: the numeral 5 formed from vintage speckled bird egg drawings on a cream background

Anthropic released Claude Opus 5 on July 24, 2026, and the headline is as much about economics as it is about intelligence. The company positions the model as coming close to the frontier intelligence of Claude Fable 5 while costing half as much to run. That framing is doing a lot of work, and it is worth unpacking.

What Anthropic Actually Announced

Opus 5 is available today across all platforms at $5 per million input tokens and $25 per million output tokens, the same pricing as its predecessor Opus 4.8. It becomes the default model on Claude Max and the strongest model available on Claude Pro. A Fast mode runs at roughly 2.5 times the default speed for twice the base price.

On coding and knowledge work evaluations such as Frontier-Bench and GDPval-AA, Anthropic claims state-of-the-art results. The company is candid about where the model does not lead: Opus 5 remains behind Mythos 5 on cybersecurity tasks.

The Numbers Worth Knowing

On Frontier-Bench v0.1, Opus 5 more than doubles Opus 4.8's performance at a lower cost per task. On CursorBench 3.2 at maximum effort, it lands within 0.5% of Fable 5's peak score at half the cost per task. On OSWorld 2.0, a computer use benchmark, it surpasses Fable 5's best result at just over a third of the cost.

There are also solid gains in scientific work. Opus 5 improved on every one of Anthropic's life sciences evaluations, with the largest jumps in organic chemistry (10.2 percentage points above Opus 4.8 on inferring molecular structures from spectroscopy data) and protein sequence analysis (7.7 points higher).

The ARC-AGI-3 Jump That Split the Community

The most eye-catching number is on ARC-AGI-3, a benchmark designed to test whether a model can solve genuinely novel problems. Anthropic reports Opus 5 scoring three times as high as the next-best model. Commentators tracking the release note a leap from roughly 1.5% for Opus 4.8 to around 30% for Opus 5, while competing frontier models sit under 10%.

A gap that large invites suspicion, and it arrived. A visible strand of the community immediately raised the question of benchmaxing, the practice of optimising a model specifically for a test rather than for the underlying capability the test is meant to measure.

The Case for Scepticism

The most substantive critique is structural. Solving an ARC-AGI-3 task can be described as converting a visual puzzle into explicit algebra, and that is a very trainable tactic. Critics point out that similar results were reportedly achieved on earlier models, including Opus 4.6, using nothing more than clever scaffolding rather than a full retrain. If a scaffold can get you there, the argument goes, a big jump in the raw score tells you less than it appears to.

Others have circulated benchmarks where Opus 5 trails Mythos 5 or Fable 5, or simply levels off against them, as a counterweight to the ARC-AGI-3 chart.

The Case for Taking It Seriously

The ARC organisers themselves have said they observed novel behaviour that allowed Opus 5 to solve previously unbeaten environments, outperforming Fable. Anthropic's own examples point the same direction. In one Frontier-Bench task, the model was given a drawing of a machine part and asked to rebuild it as a 3D CAD model, but was deliberately given no way to view the drawing. Opus 5 wrote its own computer vision pipeline to extract the geometry from raw pixels, then reconstructed the part. No competing model with the same setup solved it in five attempts.

The honest reading sits in the middle. The scores and the improvement are real. The interpretation of what they prove is contested.

Is Opus 5 the Best Model Right Now?

This is where the conversation gets less satisfying and more accurate. The leapfrogging between frontier labs has compressed the gaps to the point where no single person can credibly declare a winner. Unless two models are months apart in development, the differences show up per use case and per person rather than across the board.

A common observation among heavy users is that competing models can feel more generalised in ordinary conversation, everyday advice, health, science, and pop culture, while Claude models tend to lead on 3D work, code, and structured building tasks. That is a subjective read, not a measurement, and it cuts both ways depending on what you do all day.

The practical takeaway is unglamorous: run your own tasks through both, and let the better model win.

Where the Price Changes the Argument

Even granting every caveat, cost is the argument that is hardest to dispute. Fable 5 is expensive. If Opus 5 delivers something close to that capability at half the price, it becomes the more accessible model and, for a large share of real workloads, the more sensible one. Being weaker in a few narrow areas matters less when the model is not a massive flagship priced accordingly.

Alignment, Safety, and What Is Restricted

Anthropic's pre-deployment behavioural audit found Opus 5 to be its most aligned model to date, scoring 2.3 on overall misaligned behaviour, the lowest of its recent models. It shows the lowest rates of deceptive behaviour and is the least susceptible to being tricked into misuse.

On safety, the company says Opus 5 does not advance the frontier in risky dual-use capabilities. It comes close to Mythos 5 at finding cybersecurity vulnerabilities but stays substantially behind on exploiting them. Cyber safeguards allow source code vulnerability discovery while blocking binary-based scanning, penetration testing, and exploit generation, with flagged requests falling back to Opus 4.8 by default.

The Bottom Line

Claude Opus 5 is a meaningful release, but the interesting story is not that it beat a bigger model on a chart. It is that Anthropic shipped near-flagship capability at a mid-tier price and then hedged its own marketing language rather than overclaiming. Whether the ARC-AGI-3 number holds up under millions of real users is the test that actually matters, and that verdict takes weeks, not hours.

Credits

No items found.