OpenAI’s new GPT-6 Astra scored 62.7% under the benchmark’s standard, provider-neutral testing harness on ARC-AGI-3, an interactive benchmark that requires AI agents to explore unfamiliar games, infer their rules and goals, and plan effective actions without instructions. The same model scored 99.9% with an OpenAI-specific context-management adapter that preserves the model’s hidden reasoning state between calls.
When the ARC Prize Foundation launched ARC-AGI-3, the third generation of its Abstraction and Reasoning Corpus benchmark for artificial general intelligence, the state-of-the-art AI models it tested fared abysmally on the games, every frontier system tested scored below 1%. Opus 4.6, max lead the ranking with a score of 0.50%, while Google’s Gemini 3.1 Pro Preview, OpenAI’s GPT-5.4, high and xAI’s Grok-4.20 Beta had near zero scores. Human testing showed that every environment was solvable. In a max-effort run using OpenAI’s context-management adapter, Astra used fewer actions than the median human baseline on 96% of the levels it completed and averaged 51.7% fewer moves per level.
The advance came at substantial computational expense. The standard run cost $26,098 and the adapter-assisted run $18,817. Still, ARC Prize called the result a step-function change in frontier capability. It also noted that saturating the benchmark is not proof of AGI, which stands for artificial general. intelligence, since its environments are bounded and deterministic rather than open-ended.
In terms of open-ended tasks, our recent coverage of a preprint testing AI agents as autonomous researchers. During six-day runs, the agents with $3,000 budgets capably reviewed literature, wrote code, ran experiments and produced papers. They repeatedly pursued weak hypotheses, struggled to incorporate reviewer feedback and submitted papers that earned reject and strong-reject scores.
Astra fared less well in terms of the abstract “Intelligence” dimension on the independent Artificial Analysis ranking with a score of 61, putting it in third place and five points behind Fable 5.1 from Anthropic.




Tell Us What You Think!
You must be logged in to post a comment.