On Guanghan Ning's private benchmark Witness, Opus 5 scored only 43.4, statistically tied with Kimi K3 and Fable 5, showing a much smaller improvement over Opus 4.8 than on ARC-AGI-3. Ning notes that this pattern is consistent with training on specific types of data, but Witness cannot determine which data Anthropic used.
Greg Kamradt argues that a game based on familiar mechanisms does not test adaptability to novelty, and a weak result does not negate overall model improvement. Witness itself is designed around ARC-AGI-3-style puzzles, so better performance may reflect genuine transfer rather than memorization.
Ning later clarified that Opus 5 did generalize on Witness, but to a much lesser extent than on ARC-AGI-3. He likened the process to the evolution of coding benchmarks: as the primary target for interactive reasoning, ARC-AGI-3 naturally attracts the most training effort. Covering more edge cases may help models generalize to a broader range of abstract reasoning tasks.