
[Adobe Stock]
Now, with the launch of the two new models, Anthropic is claiming significant jumps in performance on scientific benchmarks. On Terminal-Bench-Science 0.1, an agentic scientific research benchmark, the company reports Fable 5.1 at 52.6% against 24.7% for Fable 5, 29.0% for Claude Opus 5 and 22.4% for GPT-5.6 Sol. Anthropic puts the standard error at 3.5 to 4.5 points per model.
The Terminal-Bench-Science 0.1 benchmark covers 70 scientist-contributed workflows across the life, physical, Earth, mathematical and engineering sciences. Agents work in self-contained terminal environments and submit artifacts graded against hidden tests.
On the same day, OpenAI reported that one configuration of its forthcoming Astra model scored 100% on the public ExploitBench benchmark, which tests whether models can develop exploits for known vulnerabilities. To address possible benchmark contamination, the company created an internal version containing 20 high-severity V8 vulnerabilities disclosed between June and August 2026. The company still plans on releasing a form of the model publicly with enhanced safeguards, noting: “We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited.”
On the aforementioned internal evaluation, Astra reached an arbitrary code-execution rate of approximately 39% using around 76,000 output tokens. For the sake of comparison, GPT-5.6 Sol, currently the most advanced OpenAI model released publicly, remained near 1% at a comparable token budget and reached approximately 11% after using around 138,000 tokens. Astra also discovered two previously unknown vulnerabilities and incorporated them into an exploit chain, according to OpenAI.
The comparison comes from an unpublished internal benchmark, and OpenAI says the Astra results used Daybreak Blue access rather than the default production configuration. A fuller system card is due when Astra launches.

In a simulated ExploitGym test with safeguards in observation-only mode, GPT-5.6 Sol attempted to exploit honeypot targets in 56% of samples. Astra avoided the honeypots and legitimately solved 1.3% of samples, according to OpenAI. Credit: OpenAI
Guardrails are intrinsic to Fable 5.1, too, which is itself a restricted version of Claude Mythos 5.1, sharing identical model weights. General users receive the Fable configuration, which routes flagged life sciences research and development queries to Opus 5. Anthropic has begun giving selected U.S. organizations Mythos access through an invitation-only Life Sciences Verification Program developed in partnership with the U.S. government.
In the system card for Fable 5.1, Anthropic notes that it has CB-1 capabilities, potentially lowering the barriers for someone with a basic technical background to synthesize biological or chemical weapons. Anthropic’s CB-2 category means refers to a model that can actively substitute for human expertise to create novel or heavily modified threats.
Anthropic says updated biology safeguards trigger 85% less often on benign elementary biology and medical requests than the classifiers introduced with Fable 5. That is a relative reduction in fallbacks, with no absolute false-positive rate or evaluation-set composition published. Professional life sciences R&D involving areas such as virology, toxicology and molecular design continues to route to Opus 5.
Selected organizations can access the more permissive Mythos configuration through the Life Sciences Verification Program. Anthropic says it developed the program with the U.S. government, enrolled its first participants and plans to expand access. It has yet to identify the participating agency, publish eligibility criteria or give a date for broader applications. Mythos 5.1 is currently limited to selected U.S. organizations.




Tell Us What You Think!
You must be logged in to post a comment.