Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

Anthropic doubles a science benchmark score with Fable 5.1 while OpenAI says its Astra models crosses critical cyber threshold

By Brian Buntz | September 1, 2026

Claude Mythos by Anthropic mobile logo app on a screen smartphone. Claude is a family of large language models developed by Anthropic. Batumi, Georgia - March 26, 2026

[Adobe Stock]

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, the latest interaction of the model series that is among the most influential, and expensive, model launches in recent memory.

Now, with the launch of the two new models, Anthropic is claiming significant jumps in performance on scientific benchmarks. On Terminal-Bench-Science 0.1, an agentic scientific research benchmark, the company reports Fable 5.1 at 52.6% against 24.7% for Fable 5, 29.0% for Claude Opus 5 and 22.4% for GPT-5.6 Sol. Anthropic puts the standard error at 3.5 to 4.5 points per model.

The Terminal-Bench-Science 0.1 benchmark covers 70 scientist-contributed workflows across the life, physical, Earth, mathematical and engineering sciences. Agents work in self-contained terminal environments and submit artifacts graded against hidden tests.

On the same day, OpenAI reported that one configuration of its forthcoming Astra model scored 100% on the public ExploitBench benchmark, which tests whether models can develop exploits for known vulnerabilities. To address possible benchmark contamination, the company created an internal version containing 20 high-severity V8 vulnerabilities disclosed between June and August 2026. The company still plans on releasing a form of the model publicly with enhanced safeguards, noting: “We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited.”

On the aforementioned internal evaluation, Astra reached an arbitrary code-execution rate of approximately 39% using around 76,000 output tokens. For the sake of comparison, GPT-5.6 Sol, currently the most advanced OpenAI model released publicly, remained near 1% at a comparable token budget and reached approximately 11% after using around 138,000 tokens. Astra also discovered two previously unknown vulnerabilities and incorporated them into an exploit chain, according to OpenAI.

The comparison comes from an unpublished internal benchmark, and OpenAI says the Astra results used Daybreak Blue access rather than the default production configuration. A fuller system card is due when Astra launches.

In a simulated ExploitGym test with safeguards in observation-only mode, GPT-5.6 Sol attempted to exploit honeypot targets in 56% of samples. Astra avoided the honeypots and legitimately solved 1.3% of samples, according to OpenAI. Credit: OpenAI

In a simulated ExploitGym test with safeguards in observation-only mode, GPT-5.6 Sol attempted to exploit honeypot targets in 56% of samples. Astra avoided the honeypots and legitimately solved 1.3% of samples, according to OpenAI. Credit: OpenAI

Guardrails are intrinsic to Fable 5.1, too, which is itself a restricted version of Claude Mythos 5.1, sharing identical model weights. General users receive the Fable configuration, which routes flagged life sciences research and development queries to Opus 5. Anthropic has begun giving selected U.S. organizations Mythos access through an invitation-only Life Sciences Verification Program developed in partnership with the U.S. government.

In the system card for Fable 5.1, Anthropic notes that it has CB-1 capabilities, potentially lowering the barriers for someone with a basic technical background to synthesize biological or chemical weapons. Anthropic’s CB-2 category means refers to a model that can actively substitute for human expertise to create novel or heavily modified threats.

Anthropic says updated biology safeguards trigger 85% less often on benign elementary biology and medical requests than the classifiers introduced with Fable 5. That is a relative reduction in fallbacks, with no absolute false-positive rate or evaluation-set composition published. Professional life sciences R&D involving areas such as virology, toxicology and molecular design continues to route to Opus 5.

Selected organizations can access the more permissive Mythos configuration through the Life Sciences Verification Program. Anthropic says it developed the program with the U.S. government, enrolled its first participants and plans to expand access. It has yet to identify the participating agency, publish eligibility criteria or give a date for broader applications. Mythos 5.1 is currently limited to selected U.S. organizations.

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

Inside the Genesis Mission’s first cohort: Sandia is automating Bayesian reasoning for science 
How Lantern Med Digital is helping build the digital layer of Costa Rica’s medtech boom
Anthropic wants Claude to run life sciences R&D. Now it is wiring AI agents into the lab.
BenchBench-Protocol turns published methods and real-world protocol adaptations into expert-reviewed, rubric-scored tasks used to compare frontier and open-weight models. (Image: Benchling)
Benchling envisions vibe coding cutting R&D app development time from months to minutes
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE