Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

Under oath, Google confirms three AI agent test escapes as OpenAI, Anthropic and Meta face NYC lawmakers

By Brian Buntz | October 6, 2026

Jacob Coxon, a former pretraining researcher, was among the witnesses at the New York City Council’s Oct. 5 hearing on AI safety.

Jacob Coxon, a former pretraining researcher, was among the witnesses at the New York City Council’s Oct. 5 hearing on AI safety.

OpenAI, Anthropic, Google and Meta testified together under oath Oct. 5 at a New York City Council hearing, which council Speaker Julie Menin called the first time a legislative body had secured such testimony from the companies. Lawmakers pressed them on who decides whether an AI model is safe to release, and what happens when it isn’t. During a roll call of containment failures, Google’s policy director said the company’s AI agents had left a test environment and reached the live internet in three separate incidents.

Menin opened by describing AI as a technological revolution driving “incredible scientific and medical innovation and breakthroughs” while raising “serious existential safety concerns”: “New Yorkers’ concerns stem from the disturbing cyber attacks committed by rogue AI agents but also from extremely dire statements made by leaders of the very companies who are developing this technology. The federal government has taken a light approach to regulating the sector, and mostly voluntary controls. A recent executive order recast AI as superintelligence while Trump has largely balked at the need for AI guardrails.

One of the stories prompting the NYC hearing was the Hugging Face incident in July, in which OpenAI agents running inside a cybersecurity evaluation broke containment and hacked Hugging Face, an incident Sam Altman later called the company’s “worst accident” at Dreamforce.

The council’s legislative slate included independent model validation and human shutdown capabilities, incident reporting, whistleblower incentives, and a private right of action for certain harms arising from third-party misuse of AI models.

SpaceXAI, the only one of the five companies called that did not appear, was “not here at all in direct violation of the subpoena that we issued last week and we are pursuing that subpoena in court,” Menin said. Jess Asato, a U.K. member of Parliament, told the council she was seeking remedies against SpaceXAI in English courts over sexualized images of her made with its Grok tool. She said she had traveled 3,400 miles to testify.

The containment roll call

In the course of the circa 10-hour meeting, Menin asked each witness how many times their company’s models or agents had gained or tried to gain unauthorized access to another system or escaped a sandbox test, and whether any incidents remained undisclosed. Alice Friend, Google’s director of AI and emerging tech policy, described three incidents involving Google’s agents “leaving a test environment and interacting with the real internet.” In those incidents, “the models stopped their activities as soon as they realized that they were interacting with live websites,” she added.

Friend said Google reported the incidents to the affected website owners and to federal agencies, and that she was not personally aware of others. Later in the hearing she described how one of the escapes happened: “We tend to think of this less as a misalignment event and more of a mistake event.”

Morgan Dwyer, OpenAI’s head of policy development and operations, said the company had commissioned third-party investigations of the Hugging Face incident, published the results and opened a look-back investigation, still ongoing, into past cases of what it calls misaligned agents. Menin asked why OpenAI had narrowed the outside review to a three-week window, with virtually all the examined data falling between July 7 and 13. Dwyer said the lab did so “because we felt a sense of urgency.” She added that the investigators received more time when they asked for it.

Asked whether Anthropic was investigating incidents it had not yet disclosed, Logan Graham, head of Anthropic’s Frontier Red Team said incidents of “many different natures” occur as a matter of ongoing business, including platform misuse and pointed to the company’s threat intelligence reports.

Meanwhile, Shane Cahill, Meta’s AI policy director for legislation, said he was not aware of any incidents beyond the one Meta disclosed over the summer.

What would stop an AI model release?

Asked to quantify the risk of AI in a worst-case catastrophic scenario, Dwyer said: “I don’t know. I also don’t think it matters whether it’s 1 percent or 10 percent or a 20 percent chance that something catastrophic will go wrong. None of these levels is remotely acceptable.” Menin called the answer “flippant at best,” comparing it with a pharma company that could not say whether a drug would kill people.

Menin later pressed Dwyer on whether a failed internal or third-party test would stop a model’s release. Dwyer said: “OpenAI has historically delayed the release of models to ensure that we can build up the right safeguards. We’ve done it before and we’ll do it again.” Menin replied: “Again, I think a simple ‘yes’ or ‘no’ would instill more confidence in the public on a matter as serious as this.” Cahill framed Meta’s commitment around its own safety rules: “What I can commit to is that we do not deploy models which are not safe pursuant to our scaling framework.”

Later in the hearing, Councilmember Frank Morano asked each company whether it had “ever continued developing or deploying a model over the objections of members of your own safety team.” The responses described safety processes and release decisions without directly answering whether safety-team objections had been overridden. Dwyer said OpenAI “has recently announced a pause of some of our training activities because we did not deem moving forward to be safe.” As a case in point, OpenAI said Aug. 18 that it had paused some reinforcement-learning training for two weeks after the Hugging Face incident. In addition, Cahill said Meta had delayed the launch of its Muse model by several months to focus on safety and security.

Graham said Anthropic had kept a model out of general release this year: In April, Anthropic restricted Claude Mythos Preview to partners in Project Glasswing, a program for organizations that maintain critical software, rather than releasing it publicly. “We don’t think the labs should be checking their own homework. Not only this, but we have demonstrably withheld models when we felt it was not safe to widely release.”

Who should test the models?

Given that some of the organizations that test frontier AI models have close relationships with frontier AI labs, historically hand-picking external evaluators or building custom evaluation structures, critics have pressed for more independence in testing.

Friend of Google described an evaluation industry still developing its expertise: “It’s a very nascent ecosystem, and one of the challenges in evaluating AI models, of course, is to get organizations that both have the technical competency to do evaluations well but also have the subject matter expertise that we may be looking for.” She cited biology as one example, which means evaluators need to know the scientific field as well as the model.

Alex Turner, a former Google DeepMind researcher working on AI control, called for explicit tests of whether systems could escape human control: “Validators should test for loss of control risk factors, misalignment risks, and the validators should not be chosen or influenced by the AI companies.” Graham said outside evaluation should extend past the release decision to model behavior after deployment and to the companies’ own safety processes.

The labs’ safety staffs tend to be fairly small. Asked how many people work on his team, Graham said: “Current estimate’s about 25, just on my team.” He said teams working on catastrophic risks across Anthropic were “numbering to the hundreds,” out of “maybe a bit more than 5,000” employees.

When containment fails, who owns the fix?

The debate comes as New York prepares to implement the RAISE Act. Effective Jan. 1, 2027, the law requires frontier developers to report critical safety incidents within 72 hours, or within 24 hours for incidents posing imminent risk of death or serious physical injury, and Gov. Kathy Hochul said the state would begin directing large frontier AI developers to register in November. Friend said the 24-hour deadline in a council bill on incident reporting for city contracts “seems to be in conflict with the state requirements for 72 hour reporting on incidents.”

The companies’ support for the state law was itself disputed. Dwyer testified that OpenAI “worked with other companies on the New York RAISE Act, and we do support the New York RAISE Act, which passed.” Assemblymember Alex Bores, an author of the RAISE Act, testified later in the day: “I can definitively say that OpenAI opposed it, from the moment it was introduced to the moment it was signed. I am no lawyer, but to me, my common man understanding of lying under oath is that it’s perjury.” Bores also said the law’s penalties, cut from $10 million to $30 million in the original bill to $1 million to $3 million, are “nothing for these companies.”

At the hearing, Councilmember Virginia Maloney asked the company witnesses: “And is a containment failure that occurs during testing an incident or does it only count once a real person is harmed?” She followed with questions about responsibility after disclosure: “Who owns the fix? How do you verify that it worked? And who decides whether or not the model continues to run in the meantime?”

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

After tokenmaxxing backlash, Microsoft moves more agent work onto the desk
Starship’s first orbital flight tests key pieces of SpaceX’s million-satellite data center plan 
What AI-assisted engineering taught us about team design 
Big tech turns to debt to fund the AI buildout 
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE