AI / GenAI·6 min·26 July 2026·Claude Opus 5 — Part 3 of 4

The Opus 5 benchmarks, and the four it loses

I usually do not read model releases for the benchmarks. This time I did, for one reason: the table Anthropic published contains rows where their own new model loses. That is rare in an announcement, and it makes the rest of the table more credible.

Benchmark comparison of Claude Opus 5, Fable 5, Opus 4.8 and GPT-5.6 Sol across fourteen evaluations

The same table as text, so you can search and quote it. Bold is the highest score in each row.

EvaluationOpus 5Fable 5Opus 4.8GPT-5.6 Sol
Agentic terminal coding (Frontier-Bench v0.1)43.3%33.7%21.1%34.4%
Knowledge work (GDPval-AA v2)1861174715931736
Novel problem-solving (ARC-AGI-3)30.2%not scored1.5%7.8%
Agentic search (BrowseComp)90.8%87.4%84.3%90.4%
Multidisciplinary reasoning, no tools (Humanity’s Last Exam)56.3%56.5%49.8%not scored
Multidisciplinary reasoning, with tools (Humanity’s Last Exam)64.7%63.9%57.9%not scored
Computer use (OSWorld 2.0)70.6%66.1%55.7%62.6%
Agentic coding (DeepSWE v1.1)68.8%69.7%59.0%72.7%
Agentic coding (FrontierCode v1.1, Main)53.4%53.5%46.5%47.5%
Business workflows (AutomationBench)26.0%17.4%17.0%18.1%
Legal (Legal Agent Benchmark, held-out)11.7%13.3%10.4%2.5%
Health (HealthBench Professional)59.8%66.0% (Mythos 5)57.4%60.5%
Biology, hard (BioMysteryBench)49.4%46.5%42.4%not scored
Biology, human solved (BioMysteryBench)90.1%89.0% (Mythos 5)88.5%not scored

Source: Anthropic, Introducing Claude Opus 5.

What genuinely impresses

Three numbers stand out.

On Frontier-Bench v0.1, agentic work in the terminal, Opus 5 goes from 21.1 percent (Opus 4.8) to 43.3 percent. More than double in one generation, and well clear of GPT-5.6 Sol at 34.4. This is the kind of work I do most of, and the jump is big enough to notice outside a test harness.

On ARC-AGI-3, a test built from problems the model has not seen before, it scores 30.2 percent against 1.5 for Opus 4.8 and 7.8 for GPT-5.6 Sol. That is a factor of twenty over the previous version.

On AutomationBench, business tasks carried start to finish, Opus 5 reaches 26.0 percent where the rest of the field sits between 17 and 18. That is the row closest to ordinary office work.

The four benchmarks it loses

And then the more interesting side.

On DeepSWE v1.1, GPT-5.6 Sol wins with 72.7 percent against 68.8 for Opus 5. Not marginally, just better. On FrontierCode v1.1, Fable 5 wins 53.5 to 53.4, which is inside the noise but still means Opus 5 is not the top there. On the Legal Agent Benchmark, Fable 5 wins 13.3 to 11.7. And on HealthBench Professional, Opus 5 sits at 59.8, below both GPT-5.6 Sol (60.5) and Mythos 5 (66.0).

Four of twelve benchmarks. On Humanity’s Last Exam Opus 5 also loses the no-tools measurement, 56.3 against 56.5 for Fable 5, while winning the with-tools one, so that benchmark does not count as a loss here. For a model positioned as the new default that is honestly reported, and it is a useful signal too: if your work sits heavily in coding agents or in healthcare, “the newest Anthropic model” is not automatically the right answer.

The number that puts it all in perspective

Look at that legal row again. The best model in the table scores 13.3 percent. Opus 5 scores 11.7.

That is not “nearly there”. That is: on this test, the best model available fails in almost nine cases out of ten. If you read the announcement and conclude legal work can now be automated, the evidence that it cannot is in the same table.

The same applies more gently to ARC-AGI-3. Twenty times better than the previous version sounds like a breakthrough, and in relative terms it is. In absolute terms, seventy percent of novel problems remain unsolved.

I do not call that a criticism of the model. I call it the context that falls out of press releases the moment they get retold.

What the table does not tell you

Two things worth knowing.

The benchmarks were replaced, not just filled in. There is no SWE-bench Verified in this table, for years the number everyone watched. There is Frontier-Bench, FrontierCode and DeepSWE. That is partly fair, because old tests saturate once everyone scores ninety percent. But it makes cross-generation comparison hard, and it hands the vendor influence over the yardstick. When the test changes at the same time as the model, you do not know exactly what you are measuring.

It is a vendor measuring its own product. The setup looks careful and the losing rows help, but these are not independent measurements. GDPval-AA and the AA Coding Agent Index come from Artificial Analysis, a third party; most of the rest does not. Wait for independent replication before you put a number in a deck.

What I see myself

I have only had this model on real jobs since this morning, so my impression is fresh and no more than that. What stands out is not in the table: it finishes tasks outright more often, and it announces less about what it is going to do and simply does it. That saves turns, and turns are time and money.

Something else stands out too: it stretches its brief. A few times I got back more than I asked for, cleanly done but not requested. That is exactly the behaviour the next part is about, because if you build on the API that is not charm, it is a bug in your prompt.

This piece also appeared in Dutch: De benchmarks van Opus 5, en de vier waar het verliest.