Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

GPT-6 Astra scores 62.7% on interactive reasoning benchmark, near-perfect with custom adapter

By Brian Buntz | September 3, 2026

OpenAI’s new GPT-6 Astra scored 62.7% under the benchmark’s standard, provider-neutral testing harness on ARC-AGI-3, an interactive benchmark that requires AI agents to explore unfamiliar games, infer their rules and goals, and plan effective actions without instructions. The same model scored 99.9% with an OpenAI-specific context-management adapter that preserves the model’s hidden reasoning state between calls.

When the ARC Prize Foundation launched ARC-AGI-3, the third generation of its Abstraction and Reasoning Corpus benchmark for artificial general intelligence, the state-of-the-art AI models it tested fared abysmally on the games, every frontier system tested scored below 1%. Opus 4.6, max lead the ranking with a score of 0.50%, while Google’s Gemini 3.1 Pro Preview, OpenAI’s GPT-5.4, high and xAI’s Grok-4.20 Beta had near zero scores. Human testing showed that every environment was solvable. In a max-effort run using OpenAI’s context-management adapter, Astra used fewer actions than the median human baseline on 96% of the levels it completed and averaged 51.7% fewer moves per level.

The advance came at substantial computational expense. The standard run cost $26,098 and the adapter-assisted run $18,817. Still, ARC Prize called the result a step-function change in frontier capability. It also noted that saturating the benchmark is not proof of AGI, which stands for artificial general. intelligence, since its environments are bounded and deterministic rather than open-ended.

In terms of open-ended tasks, our recent coverage of a preprint testing AI agents as autonomous researchers. During six-day runs, the agents with $3,000 budgets capably reviewed literature, wrote code, ran experiments and produced papers. They repeatedly pursued weak hypotheses, struggled to incorporate reviewer feedback and submitted papers that earned reject and strong-reject scores.

Astra fared less well in terms of the abstract “Intelligence” dimension on the independent Artificial Analysis ranking with a score of 61, putting it in third place and five points behind Fable 5.1 from Anthropic.

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

Claude Mythos by Anthropic mobile logo app on a screen smartphone. Claude is a family of large language models developed by Anthropic. Batumi, Georgia - March 26, 2026
Anthropic doubles a science benchmark score with Fable 5.1 while OpenAI says its Astra model crosses critical cyber threshold
Inside the Genesis Mission’s first cohort: Sandia is automating Bayesian reasoning for science 
How Lantern Med Digital is helping build the digital layer of Costa Rica’s medtech boom
Anthropic wants Claude to run life sciences R&D. Now it is wiring AI agents into the lab.
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE