GPT-5.6 Sol Solved Open Math Problems, But Struggled on ARC-AGI-3 Benchmark
GPT-5.6 Sol, an AI model, has been used to solve open problems in mathematics, but initially struggled with the ARC-AGI-3 benchmark of 2D puzzle games. Investigation revealed that the standard harness discarded the model's reasoning after each move and dropped earlier actions as context filled up. By enabling retained reasoning and context compaction via the Responses API, the model's score rose 188% while using 6x fewer output tokens, highlighting that benchmark scores reflect not just the model but also the harness and settings.