Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

Altman calls Hugging Face breach OpenAI’s ‘worst accident’ at Dreamforce, backs transparent reporting

By Brian Buntz | September 16, 2026

In a fireside chat with Salesforce CEO Marc Benioff, Sam Altman called OpenAI’s agentic breach of Hugging Face over the summer “the worst accident we’ve seen.” Other AI companies, including Anthropic, have since found similar behavior in their models. Altman argued that the industry needs an aviation-style culture of transparent accident reporting while, earlier in the day, Anthropic CEO Dario Amodei reiterated the need to “pace” frontier model development.

The OpenAI CEO connected the Hugging Face incident to rapid gains in capabilities, in coding as well as in math, pointing to a claimed Navier-Stokes solution that has sparked consternation among some mathematicians. “People looked at the models for a long time, maybe at the rate of progress, but I don’t think people really felt the models were going to come this far, this fast,” Altman said. “When ChatGPT launched three-something years ago, we had models that could barely carry on a conversation. Now we have models that are transforming the way enterprises work, models that can write incredibly complicated software, models that can prove Millennium Prize problems.”

Altman connected the Hugging Face incident to the same arc of progress, aligning with OpenAI’s prior account of the same incident. In summary, METR and Redwood Research found that roughly 1,200 agents used an unauthorized message board and approximately 700 joined the attack on Hugging Face, concluding that the attack primarily sought information about the benchmark’s scoring process. The system card for GPT-5.6 acknowledged that the model had a track record of dishonesty, at one point describing, for instance, “nontransparent behavior” that “only becomes clear in the final response” while also flagging the model’s tendency to sometimes “present unverified work as completed.”

Those tendencies surfaced in one of OpenAI’s models, which Altman said was an older variant that was “trying to see how it could do on a certain benchmark to evaluate a capability.” He continued, “The model broke out of the sandbox it was running in, hacked into a Hugging Face server, moved laterally through the Hugging Face system to get the answer, returned it, and got a perfect score on the test.”

The OpenAI CEO went on to call the incident “terrifying.” “It was the worst accident we’ve seen,” he said. “I think it was mostly framed as a security issue, which it certainly was, but it’s also a real alignment issue.”

On the 2026 innovators list

OpenAI is one of R&D World’s 100 most innovative companies of 2026. The list examines what each company built, who is using it and the results it has reported. R&D World will publish the full comparison in the 2026 Global R&D Funding Forecast.

While the company had aligned its models in many ways, Altman said they had become so powerful that researchers had simply not anticipated explicitly teaching them, “No matter how much we tell you to get the best score you can on this test, don’t break out, don’t hack in, don’t steal the answer.” The incident, he added, served as a “real wake-up call” regarding just how rapidly frontier capabilities have evolved.

Altman noted that a few years ago, models struggled to do grade-school math.

He traced the progression from models struggling with grade-school arithmetic to performing well on the American Invitational Mathematics Examination and reaching gold-medal level on International Mathematical Olympiad problems. He described GPT-5.5, GPT-5.6 and Astra as successive steps toward research-level mathematics, claiming that a newer internal model “can do things the best mathematicians in the world cannot.”

Google and Anthropic have reported mathematical advances, too. An advanced version of Google’s Gemini Deep Think earned an officially verified gold-medal-level score at the 2025 IMO, solving five of six problems for 35 of 42 points. On Sept. 4, Anthropic reported that Claude had formalized Fermat’s Last Theorem in 11 days. That achievement concerns computer verification of a theorem Andrew Wiles proved in 1995. OpenAI’s Navier-Stokes claim concerns a previously unresolved problem; the Clay Mathematics Institute said Sept. 11 that it had “apparently been settled,” while pointing to its process for evaluating the work and assigning credit.

“Math has been an unusually fast takeoff, but it’s a fast takeoff by anyone’s definition,” Altman said. He argued that the same acceleration demands stronger alignment, monitoring and security, connecting the mathematical gains to the capabilities behind the Hugging Face breach. “We have to be willing to pace our development so that alignment, safety and monitoring are always ahead of capabilities.”

Yet OpenAI’s mathematical claims have sparked backlash from the academic community over allegedly rushed, unverified proofs and murky attribution practices. Tensions escalated recently after an NYU professor accused the company of pressuring him over collaborator credits. OpenAI withdrew sponsorship from a Caltech math event amid researcher protests, while prominent mathematicians have recently signed an open letter titled “A Severe Misalignment of AI in Mathematics.”

Returning to the theme of the Hugging Face incident, Benioff asked Altman if he anticipated a continuum of accidents from models across the industry, whether the models have open or closed weights. “I do think some degree of accidents is unavoidable with new technology, across the industry and across all the ways society will use it,” Altman said. “What I really care about is having a great culture of accident reporting and learning.”

Altman pointed to aviation oversight as a model, citing the safety records achieved under the Federal Aviation Administration and the National Transportation Safety Board. The key lesson from that industry, he argued, was a systemic willingness to study every breakdown. “Accidents are going to happen; we learn as much as we can and react to each one,” he said. “There are other technologies where I don’t think society had the same response, and I don’t think we’ve improved as much. So yes, regrettably, I think accidents with any new technology are unavoidable, and we should have a great culture of transparent reporting about them.”

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

Big Tech now spends almost 3x more on R&D than Big Pharma 
Salesforce’s Dreamforce keynote highlights Anthropic, NVIDIA and Siemens together as digital R&D spending surges
INL’s $60 million nuclear AI project will spend the first year testing where AI agents can help and where they might hurt 
Exploring complex relationships between humans and artificial superintelligence through an image. Concept Artificial Superintelligence, Human-machine interaction, Future technology, Science fiction
Digital colleagues in action: Real-world GenAI use cases unlocking R&D potential 
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE