Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Key Points:
- Anthropic's Claude Opus 5 set a new record on the ARC-AGI-3 benchmark with a 30.2% score, significantly outperforming the previous leader, OpenAI's GPT-5.6 Sol (Max), which scored 7.8%.
- Opus 5 demonstrated advanced logical reasoning, including translating tasks into algebraic notation and formulating reflection equations, enabling it to solve five previously unsolved environments at or above human level.
- ARC-AGI-3 evaluates AI models' ability to solve novel tasks without prior training, focusing on general reasoning and autonomous exploration rather than relying on stored knowledge or external software assistance.
- While Opus 5's gains on ARC-AGI-3 are substantial, tests on other benchmarks like Witness show narrower improvements, suggesting the model's advancements may be more specialized to certain puzzle types or data.
- Researchers emphasize that further testing on a wider range of unfamiliar tasks is needed to determine whether Opus 5's reasoning improvements are broadly generalizable or more domain-specific.