According to Anthropic's internal benchmarks, Opus 5 sets new highs on several evaluations. On Frontier-Bench v0.1, Opus 5's end-to-end coding score (43.3%) is more than double Opus 4.8 (21.1%) and surpasses Fable 5 (33.7%) and GPT-5.6 Sol (34.4%). On the knowledge work benchmark GDPval-AA v2, Opus 5's Elo rating (1,861) leads Fable 5 (1,747) and GPT-5.6 Sol (1,736).
The most striking result is on ARC-AGI-3, which tests a model's ability to solve novel problems. Opus 5 scores 30.2%, nearly four times that of second-place GPT-5.6 Sol (7.8%). However, this test lacks a comparison with Fable 5, and some observers believe this huge lead may not be reproducible in real-world use.
However, Opus 5 does not lead on all tests. On the agentic coding benchmark DeepSWE v1.1, GPT-5.6 Sol (72.7%) leads Fable 5 (69.7%) and Opus 5 (68.8%). On health and legal tasks, Fable 5 and Mythos 5 still perform better. Additionally, on cybersecurity tasks, Opus 5 approaches Mythos 5 in vulnerability discovery but lags significantly in exploit development.