Back to feed
News Story
THE DECODER
1 sources

Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

Anthropic's Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, nearly quadrupling the previous record of 7.8% set by GPT-5.6 Sol. The benchmark's developers noted that the model independently formulated reflection equations, a behavior not seen before, indicating stronger logical reasoning.

SynthePulse Insight · AI deep reading

Opus 5 Sets Record on ARC-AGI-3: True Reasoning Breakthrough or Benchmark Specialization?

Version 1 · 1 source

Anthropic's Claude Opus 5 scores 30.2% on the ARC-AGI-3 benchmark, nearly quadrupling the previous record held by GPT-5.6 Sol. The ARC Prize team attributes this lead to stronger logical reasoning, but independent tests suggest its advantage may be limited to specific task types.

  • Claude Opus 5 scores 30.2% on ARC-AGI-3, far surpassing GPT-5.6 Sol's record of 7.8%.
  • ARC Prize analysis attributes the lead to stronger logical reasoning, enabling autonomous exploration and planning in unfamiliar environments.
  • During testing, Opus 5 exhibited previously unseen behaviors, such as converting tasks into algebraic symbols and independently deriving reflection equations.
  • Opus 5 solved five previously unsolved environments, four of which matched or exceeded human performance.
  • On the Witness benchmark, Opus 5 scored 43.4, statistically tied with Kimi K3 and Fable 5, a much smaller improvement than on ARC-AGI-3.
  • Researchers note that Opus 5's advantage on ARC-AGI-3 may partly stem from targeted training on that benchmark.
Open section navigationRecord Score and Novel Behaviors

Record Score and Novel Behaviors

Anthropic's Claude Opus 5 achieved a score of 30.2% on the ARC-AGI-3 benchmark, nearly quadrupling the previous record of 7.8% held by OpenAI GPT-5.6 Sol (Max). The ARC Prize team attributes this lead to "stronger logical reasoning capabilities, enabling more autonomous exploration, planning, and execution in unfamiliar environments."

During testing, Opus 5 exhibited behaviors never before observed in AI models: it converted tasks into algebraic symbols and independently derived reflection equations. Additionally, Opus 5 solved five previously unsolved environments, four of which matched or exceeded human performance.

Benchmark Design and Controversy

ARC-AGI-3 is designed to measure an AI model's ability to solve novel tasks not seen during training, which humans can typically handle easily. The current version uses a game-like format where the model must infer the rules of an interactive environment, plan actions, and execute them step by step. This tests general reasoning ability rather than stored knowledge.

Notably, the official score only accounts for the language model's performance without additional software (i.e., the "harness"). The ARC Prize believes that future AGI systems should not require external help to solve new tasks. Opus 5's score could be higher if used within Claude Code.

Limitations Revealed by Independent Tests

On Guanghan Ning's private benchmark Witness, Opus 5 scored only 43.4, statistically tied with Kimi K3 and Fable 5, showing a much smaller improvement over Opus 4.8 than on ARC-AGI-3. Ning notes that this pattern is consistent with training on specific types of data, but Witness cannot determine which data Anthropic used.

Greg Kamradt argues that a game based on familiar mechanisms does not test adaptability to novelty, and a weak result does not negate overall model improvement. Witness itself is designed around ARC-AGI-3-style puzzles, so better performance may reflect genuine transfer rather than memorization.

Ning later clarified that Opus 5 did generalize on Witness, but to a much lesser extent than on ARC-AGI-3. He likened the process to the evolution of coding benchmarks: as the primary target for interactive reasoning, ARC-AGI-3 naturally attracts the most training effort. Covering more edge cases may help models generalize to a broader range of abstract reasoning tasks.

Possible Training Strategies and Open Questions

Anthropic has not explained the performance improvement, but targeted data annotation and reinforcement learning are plausible factors. Unlike earlier models, Opus 5 was developed after ARC-AGI-3 and its format were made public, which may have allowed Anthropic to train on the benchmark's skills and puzzle format, though this does not imply training on exact tasks. Annotators could label reasoning trajectories, useful actions, failed attempts, and recovery steps for similar puzzles, while reinforcement learning could reward exploration, planning, rule discovery, and self-correction.

Credibility boundary

This article primarily relies on reporting from THE DECODER, which cites public results and analysis from the ARC Prize, as well as comments from independent researchers Guanghan Ning and Greg Kamradt. All specific numbers and conclusions come from these sources, without introducing external knowledge. Some discussion of training strategies is reasonable inference.

Insight takeaway

Claude Opus 5's performance on ARC-AGI-3 is indeed impressive, but independent tests suggest its advantage may partly stem from targeted training on that benchmark. Genuine improvements in general reasoning remain to be validated more broadly.

Primary report

THE DECODER

Primary source