Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

Who has access to Claude Mythos-tier models (and beyond) will redefine cybersecurity, including in R&D

By Brian Buntz | May 9, 2026

Mozilla chart showing monthly Firefox security bug fixes, which hovered in the teens and 20s through most of 2025 before rising to 423 in April 2026 as Mozilla used Claude Mythos Preview and other AI models to harden Firefox. Image courtesy of Mozilla.

Mozilla chart showing monthly Firefox security bug fixes, which hovered in the teens and 20s through most of 2025 before rising to 423 in April 2026 as Mozilla used Claude Mythos Preview and other AI models to harden Firefox. Image courtesy of Mozilla.

Last April, Firefox patched 31 security vulnerabilities. For most of the year, the number hovered between the teens and the mid-twenties. The numbers began to tick up after, in January 2026, Anthropic partnered with Mozilla and deployed its Claude Opus 4.6 model to scan Firefox over a two-week period. That month, the model discovered 22 vulnerabilities. A total of 14 of them were high-severity. But then the number jumped to 423 in April. The reason? Mozilla had been granted early access to Claude Mythos Preview, Anthropic’s most powerful model to date, through a vetted-partner program called Project Glasswing.

Mythos alone surfaced 271 of those 423 bugs. Only three warranted standalone Common Vulnerabilities and Exposures, or CVEs, the public identifiers security teams use to track and discuss known software flaws. A CVE entry gives a vulnerability a shared name, description and reference point. It allows vendors, researchers and defenders to talk about the same issue without ambiguity. The rest were lower-severity issues, defense-in-depth hardening and fixes in long-dormant code paths that no single team would have audited manually.

Defenders get a head start, for now…

The same class of AI systems that can comb through a browser codebase can also change how defenders think about lower-priority risks in R&D environments, from source code and cloud infrastructure to ELNs, LIMS, connected instruments and sensitive research data.

Cliff Steinhauer, director of information security and engagement at the National Cybersecurity Alliance, said the Firefox results point to a broader shift in how organizations should reassess their vulnerability backlogs. “A lot of these organizations probably have some kind of SOC 2 Type II, ISO 27001 audit process where they know where their gaps are, and maybe they’ve been having to prioritize those gaps,” he said. “The lower ones are kind of still sitting out there, because there are enough high-priority gaps and vulnerabilities to keep everybody busy, so the lower ones fall to the wayside.”

Anthropic says Project Glasswing partners receive Claude Mythos Preview to find and fix vulnerabilities in foundational systems that make up a large share of the world’s shared attack surface. The roster reads like an infrastructure map. AWS and Google host the cloud environments where research pipelines run. Microsoft and the Linux Foundation maintain the operating systems and open-source libraries underneath. Broadcom and Nvidia supply the chips and networking silicon. Cisco and Palo Alto Networks build the firewalls and network security layers. CrowdStrike monitors endpoints. Apple ships the devices. Mozilla, as the Firefox data shows, maintains one of the most widely used browsers on earth. OpenAI is moving along a similar axis with Trusted Access for Cyber, an identity- and trust-based program that gives verified defenders access to more capable, more permissive cyber models.

The list of who gets access is itself a map of which layers of the technology stack will be hardened first and which will not. Cloud providers, operating-system maintainers, chipmakers, browser developers, networking vendors and endpoint-security firms are at the front of the line.

When low-severity flaws become high-severity chains

One recent example of an attack that had a widespread impact on an R&D heavy sector came in February 2024, when attackers breached Cencora, the drug distribution and patient-services giant formerly known as AmerisourceBergen. From that single point of entry, patient data from 27 pharmaceutical companies spilled out, including records held for Novartis, Bayer, AbbVie, GSK and Bristol Myers Squibb.

Systems like Mythos, when weaponized, could automate that entire sequence. An AI agent scans a pharma company’s external attack surface and finds a medium-severity flaw in a vendor portal. It scores a five or six on the CVSS scale, the 0-to-10 rating system security teams use to decide what to patch first. Exploiting it requires local network access, so it sits near the bottom of the queue. A second scan finds a misconfigured LIMS integration that provides exactly that foothold. Neither flaw alone would make anyone’s priority list. Together, they are a path straight to research data. Early testers of both Mythos and OpenAI’s GPT-5.4-Cyber have reported that the models’ real leap is in chaining exploits at a speed and scale no human team can match.

“I think the time has come to move down to those low-severity vulnerabilities, especially because AI can sit there and chain multiple of those together in a way that escalates the severity,” Steinhauer said. “You don’t need a zero day. You don’t need a 9.9 unpatched vulnerability. You only need a couple of sixes, because the AI has social-engineered somebody to have local access to the network.”

The human attack surface

The chaining problem gets worse when it moves beyond code. “If you could automate that and scale that, then exactly what you’re talking about becomes possible,” Steinhauer said. “It’s a tidal wave of social engineering driven by large language model agents that are only getting better at what they do.”

And the channels keep multiplying. “It’s not just email either. [Weaponized AI agents] can make phone calls. They can text. They can create LinkedIn accounts, fill it all in, write messages. It is absolutely going to be potentially a huge flood of automated messages that look human.”

What makes those messages effective, he said, is that they won’t follow the playbook most people have learned to spot. In short, phishing in the future is likely to be much tougher to detect. “The slower, more strategic, sporadic type of emails, it’s not going to look like somebody being urgent and trying to close a deal as soon as possible, like you see with human scammers. It’ll be a lot slower, a lot more long-term.”

Redefining offense and defense

Patching software is one problem while patching people is another. Anthropic’s own Mythos Preview system card notes that non-experts with no formal security training have used the model to produce complete, working exploits overnight, a capability that could just as easily be turned toward crafting persuasive pretexts as toward finding code-level flaws.

“Social engineering from an AI perspective is a huge vulnerability that is a lot harder to defend against, because you can have a whole security team working on securing a system, but how do you secure a CEO from their psychology?” Steinhauer said. “You have to have some kind of human vulnerability management, which is a thing. We call it human risk management.”

For now, programs like Glasswing and Trusted Access for Cyber limit the most capable models to a short list of vetted defenders. Steinhauer sees that constraint as temporary.

“There are other organizations developing models that are catching up already, and they’re not holding back, they’re not restricting it to 50 companies. They’re giving it to all their paying customers. Some of those could be bad guys, and I think it’s going to be an interesting next 12 months.”

Related Articles Read More >

GPT-6 Astra scores 62.7% on interactive reasoning benchmark, near-perfect with custom adapter
Claude Mythos by Anthropic mobile logo app on a screen smartphone. Claude is a family of large language models developed by Anthropic. Batumi, Georgia - March 26, 2026
Anthropic doubles a science benchmark score with Fable 5.1 while OpenAI says its Astra model crosses critical cyber threshold
Inside the Genesis Mission’s first cohort: Sandia is automating Bayesian reasoning for science 
How Lantern Med Digital is helping build the digital layer of Costa Rica’s medtech boom
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE