OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness
OpenAI says it can keep up on ARC-AGI-3. After Anthropic’s Claude Opus 5 quadrupled the record score on the logic benchmark, OpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5’s 30.2 percent. GPT-5.6 Sol’s ARC-AGI-3 scores jump dramatically when using OpenAI’s custom harness with retained reasoning and compaction, compared to the…