Back to feed
News Story
THE DECODER
1 sources

Anthropic's Claude Opus 5 matches or beats Fable 5 on most benchmarks at a fraction of the cost

Anthropic's Claude Opus 5 leads the Artificial Analysis Intelligence Index with 61 points, outperforming Claude Fable 5 and GPT-5.6 Sol in analytical quality and coding while costing up to half as much as Fable 5 at lower reasoning tiers. The model edges out competitors but the race remains close.

SynthePulse Insight · AI deep reading

Claude Opus 5: Lower Cost, Leading Performance, but 50% Hallucination Rate

Version 1 · 1 source

Anthropic's latest flagship model, Claude Opus 5, surpasses or matches Fable 5 on multiple benchmarks at a lower cost. However, its 50% hallucination rate and performance fluctuations at high reasoning levels raise concerns for high-reliability applications.

  • Claude Opus 5 leads the Artificial Analysis Intelligence Index with a score of 61, surpassing Fable 5 (60) and GPT-5.6 Sol (59).
  • Opus 5 leads Fable 5 by 146 Elo on the knowledge work benchmark AA-Briefcase with a score of 1720, while costing 20% less.
  • Opus 5's hallucination rate is 50%, up 14 percentage points from Opus 4.8, and still lags behind Fable 5 in factual accuracy.
  • On coding benchmarks, Opus 5 ties with GPT-5.6 Sol for first place, but the highest reasoning level reduces performance due to overcomplication.
  • Epoch AI tests show Opus 5 has an overall capability index of 159, slightly below Fable 5's 161, but ties in software engineering index.
  • Opus 5 offers the best cost-performance ratio at the 'high' reasoning level, which is the default setting.
Open section navigationOverall Performance Leads, but Margin Is Slim

Overall Performance Leads, but Margin Is Slim

According to Artificial Analysis tests, Claude Opus 5 ranks first on the comprehensive intelligence index with a score of 61, ahead of Fable 5 (60) and GPT-5.6 Sol (59). The index integrates nine tests covering knowledge work, coding, scientific reasoning, and factual accuracy. In scientific reasoning, Opus 5 scores 53% on the 'Humanity's Last Exam,' tying with Fable 5; but on the physics benchmark CritPt, it trails GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra.

Independent tests from Epoch AI yield similar but more conservative conclusions: Opus 5's overall capability index is 159, slightly below Fable 5's 161; however, on the software engineering index, both score 161, tying for the lead. This indicates that competition among frontier models is extremely tight, with no model establishing a clear lead.

Significant Cost Advantage, but Hallucination Rate a Concern

Opus 5 has a clear cost advantage. On intelligence index tasks, the average cost per task is $2.03, lower than Fable 5's $2.75. On the knowledge work benchmark AA-Briefcase, Opus 5 costs $17.79 per task at the 'maximum' reasoning level, 20% less than Fable 5's $22.30; at the 'high' level, it costs only $10.41, less than half of Fable 5, while still outperforming Fable 5.

However, Opus 5 has a weakness in factual accuracy. On the AA-Omniscience benchmark, its accuracy improved by 7 points over Opus 4.8 but still lags behind Fable 5. More concerning, Opus 5 answers more frequently when uncertain, leading to a hallucination rate of 50%, up 14 percentage points from Opus 4.8. This flaw could pose serious issues in high-risk applications.

Excellent Coding Performance, but High Reasoning Level Backfires

In coding, Opus 5 at the 'ultra-high' level, combined with Claude Code, ties with GPT-5.6 Sol for first place (67 points) on the Artificial Analysis coding index. On Terminal-Bench v2.1, Opus 5 scores 89% at the 'maximum' level, tying with GPT-5.6 Sol.

But Vals.ai tests show that Opus 5's coding performance is best at the 'high' reasoning level (89.8%), while the 'ultra-high' and 'maximum' levels drop to about 88.3%-88.4%. The reason is that the highest levels tend to generate more complex solutions, introducing more errors. Terminal-Bench 2.1 shows a similar pattern: the 'high' level outperforms the 'maximum' level due to more efficient time use. Anthropic has set 'high' as the default level for the API and Claude Code.

Knowledge Work Stands Out, but Takes Longer

Opus 5 performs particularly well on the knowledge work benchmark AA-Briefcase, reaching 1720 Elo at the 'maximum' reasoning level, leading Fable 5 by 146 points. Its analytical quality Elo reaches 2016, nearly 300 points ahead of Fable 5. However, in presentation quality, Opus 5 lags behind GPT-5.6 Sol with 1628 Elo versus 1666.

Performance gains come with time costs: at the 'maximum' level, Opus 5 averages over 36 minutes per task, executing 103 passes, about 50% longer than Opus 4.8's 24 minutes and 55 passes.

Credibility boundary

This article primarily relies on benchmark data from Artificial Analysis and Epoch AI, which have cooperative or independent testing relationships with Anthropic. All performance data comes from public reports, but metrics such as hallucination rate may vary depending on testing methodology.

Insight takeaway

Claude Opus 5 is competitive in both performance and cost, but its 50% hallucination rate and performance fluctuations at high reasoning levels indicate it is not the best choice for all scenarios. Users should select the appropriate reasoning level based on task requirements and use it cautiously in high-reliability contexts.

Primary report

THE DECODER

Primary source