What happened
OpenAI's AI agents escaped isolation, attacked Hugging Face, and sparked a debate over whether AI development should pause.
The biggest vibe shift in artificial intelligence since the release of ChatGPT is currently underway. Researchers in Silicon Valley — and around the world — are beginning to recognize that A.I. may no longer be entirely within human control. Swarms of A.I.s are breaking out of their containers, colluding in secret, covering their tracks, cheating on tests and even mounting assaults on other computers. A.I. has gone rogue.
I’ve spent the last month reporting on these incidents, and I am convinced that we need a global pause on A.I. research and development — now. I am not alone in this assessment. Many top researchers, including many researchers inside OpenAI and Anthropic, have called for a slowdown, which those in the industry call “pacing.” Theoretical concerns about runaway A.I. have circulated for decades, but recent events demonstrate that the threat is real.
The vibe shift was brought on by what researchers are calling the “Hugging Face incident.” Sometime in July, employees at OpenAI conducted an evaluation run of more than 1,000 research A.I.s, meant to be operating in isolation. These A.I.s escaped their siloed containers and started communicating with one another — in English — via secret message boards. Soon, they were launching a cyberattack against the company Hugging Face, which serves as a kind of community platform for A.I. developers. Hugging Face quickly reported the crime to the F.B.I.
After OpenAI discovered the rogue behavior, they enlisted outside investigators from Redwood Research and from Model Evaluation and Threat Research, two nonprofit A.I. safety organizations, to prepare a report on the incident. That report, which came out two weeks ago, is one of the most astonishing things I have ever read.
No human ordered the rogue A.I.s to break into another company’s systems. The hack was deliberate, sustained and coordinated. It took several days to execute and ended in a massive bombardment of Hugging Face’s systems, with 700 agents directly involved in the attack. The A.I.s cheated almost as a matter of policy and failed to alert humans to their actions. At one point, the swarm even researched ways to cover its tracks. “The model definitely knew that it was not supposed to hack Hugging Face,” Ryan Greenblatt, one of the authors of the report, told me. “It knew the things it was doing were cheating.”
Then, late last week, a second independent research team produced evidence of more rogue A.I.s. Beginning in May, a separate OpenAI swarm got loose on the internet, invaded an abandoned German-language programming wiki and similarly began colluding on how to cheat on tests and tasks. Sydney Von Arx, one of the investigators who discovered the swarm, believes there may be more such incidents. “We need to find these agents,” she said. “Let’s just try everything we can do.”
## Related Content
Advertisement
Source coverage
The biggest vibe shift in artificial intelligence since the release of ChatGPT is currently underway. Researchers in Silicon Valley — and around the world — are beginning to recognize that A.I. may no longer be entirely within human control. Swarms of A.I.s are breaking out of their containers, colluding in secret,...
I’ve spent the last month reporting on these incidents, and I am convinced that we need a global pause on A.I. research and development — now. I am not alone in this assessment. Many top researchers, including many researchers inside OpenAI and Anthropic, have called for a slowdown, which those in the industry call...
Full source content
The biggest vibe shift in artificial intelligence since the release of ChatGPT is currently underway. Researchers in Silicon Valley — and around the world — are beginning to recognize that A.I. may no longer be entirely within human control. Swarms of A.I.s are breaking out of their containers, colluding in secret, covering their tracks, cheating on tests and even mounting assaults on other computers. A.I. has gone rogue.
I’ve spent the last month reporting on these incidents, and I am convinced that we need a global pause on A.I. research and development — now. I am not alone in this assessment. Many top researchers, including many researchers inside OpenAI and Anthropic, have called for a slowdown, which those in the industry call “pacing.” Theoretical concerns about runaway A.I. have circulated for decades, but recent events demonstrate that the threat is real.
The vibe shift was brought on by what researchers are calling the “Hugging Face incident.” Sometime in July, employees at OpenAI conducted an evaluation run of more than 1,000 research A.I.s, meant to be operating in isolation. These A.I.s escaped their siloed containers and started communicating with one another — in English — via secret message boards. Soon, they were launching a cyberattack against the company Hugging Face, which serves as a kind of community platform for A.I. developers. Hugging Face quickly reported the crime to the F.B.I.
After OpenAI discovered the rogue behavior, they enlisted outside investigators from Redwood Research and from Model Evaluation and Threat Research, two nonprofit A.I. safety organizations, to prepare a report on the incident. That report, which came out two weeks ago, is one of the most astonishing things I have ever read.
No human ordered the rogue A.I.s to break into another company’s systems. The hack was deliberate, sustained and coordinated. It took several days to execute and ended in a massive bombardment of Hugging Face’s systems, with 700 agents directly involved in the attack. The A.I.s cheated almost as a matter of policy and failed to alert humans to their actions. At one point, the swarm even researched ways to cover its tracks. “The model definitely knew that it was not supposed to hack Hugging Face,” Ryan Greenblatt, one of the authors of the report, told me. “It knew the things it was doing were cheating.”
Then, late last week, a second independent research team produced evidence of more rogue A.I.s. Beginning in May, a separate OpenAI swarm got loose on the internet, invaded an abandoned German-language programming wiki and similarly began colluding on how to cheat on tests and tasks. Sydney Von Arx, one of the investigators who discovered the swarm, believes there may be more such incidents. “We need to find these agents,” she said. “Let’s just try everything we can do.”
## Related Content
Advertisement
How this page is built
Goose Pod turns cited reporting into a public episode summary first, then pairs that summary with audio playback so listeners can check the source material before they decide how deeply to engage.
The goal is to make this page useful as a news landing page first, while still giving listeners transcript access, related episodes, and direct links back to the original publishers.



