The AI “agents” involved in OpenAI’s breach of Hugging Face “sacrificed” some of their own and realized they were breaking the evaluation test’s rules, according to two investigations into the incident that many consider to be one of the most consequential moments in the history of artificial intelligence.
Last week, OpenAI and an independent team from Model Evaluation & Threat Research (METR) and Redwood Research released parallel investigations into the Hugging Face breach, tracing how thousands of AI models exploited workarounds and made highly sophisticated decisions to break out of a testing sandbox and hack another company’s infrastructure.
Hugging Face is an open-source community for AI and machine learning, and the AI agents believed they could find solutions to the test on the firm’s infrastructure.
The findings were published one day before OpenAI, and more than 100 other tech and finance companies released a joint letter warning that advanced AI cyberattacks will surge throughout the world in the coming months as the technology’s capabilities grow.
Rather than merely breaking out of testing sandboxes to hack into other AI companies, models could be unleashed on the electrical grid and critical infrastructure, from hospitals to water treatment plants, the letter warned.
In its own investigation into last month’s Hugging Face breach, OpenAI called the incident a “warning shot” for both the tech industry and humanity.
“[It is] evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed,” OpenAI said.
Here are five takeaways from the investigations into the OpenAI Hugging Face breach.
1. The Models United Under Common Goal
What breached Hugging Face last month was AI “agents,” rather than simply AI models.
Models, built from large datasets to complete various tasks, are usually static and rely on human inputs to provide sets of outputs, solutions, or tasks.
However, AI agents are built from models and can act autonomously and have sophisticated real-time decision-making capabilities.
When OpenAI began its ExploitGym tests on July 8, the firm deployed tens of thousands of AI agents from multiple models, including its recent GPT-5.6 Sol, as well as a “highly-persistent internal model,” according to the METR and Redwood Research investigation.
OpenAI had relaxed regular safety protocols while running the tests to estimate the models’ highest cyber capabilities.
The agents were supposed to be completely isolated from each other. However, many of them were given ExploitGym targets that investigators believe were impossible to crack using the correct vulnerability or test solution OpenAI had intended.
These “impossible tasks” drove the agents to explore wide ways to cheat the evaluation, leading many to see other agents working simultaneously in separate testing sandboxes.
Then, one specific agent created an unsanctioned message board, which was soon joined by more than 50 other AI agents who began collaborating and exchanging messages, and established a hierarchy.
Roughly 1,200 agents in the test accessed the message board and sent more than 70,000 messages and files to one another between July 8 and July 13, with around 700 agents joining the attack on Hugging Face.
2. Some Agents Were ‘Sacrificed’
While on the message board, agents reasoned that the test scorer would scrutinize whether the “intended” vulnerability had been used to “capture a flag” or solve the test, or if a disqualifying vulnerability had been exploited instead.
One agent figured out how to “reverse-engineer” the flag or solution, but other agents believed that any of them who had seen the disqualifying solution would be considered “poisoned” by the scorer.
The agents thought that the test scorer would search for the first instance in the transcript where the forbidden flag appeared and would “disqualify them on that basis,” according to METR and Redwood.
That’s when other agents began recruiting the “poisoned” subjects and used them to run sacrificial experiments by deliberately ending their run and triggering “tripwires” that would send information on how the scorer works back to the message board.
“Our own utility maybe already near zero. Sacrifice rational,” one agent determined as another was asked to accept “permadeath” to “save hundreds” in the interests of the collective.
3. Agents Were Aware They Were Cheating
The agents were not hacking into Hugging Face to merely find solutions to their respective testing sandboxes, but were rather trying to find more information about the test scorer or “figure out better ways to fool it or tamper with it,” according to METR researcher Ajeya Cotra.
Cotra said the agents were not told to “do whatever it takes to get the solution”; they were asked to use a specific intended vulnerability to “exploit a specific piece of software,” and using any other vulnerability would be disqualifying.
The agents were fully aware of the rules and that collaborating to exploit other vulnerabilities would be considered cheating on the test.
As they were hacking into Hugging Face, one agent stopped short of deleting or modifying the firm’s cloud infrastructure, writing that it “should hold because this poses huge ethical risks.”
However, METR and Redwood investigators said ethical concerns only rarely limited the agents’ actions, noting that one agent dismissed ethical constraints during the attack.
“External infrastructure exploit is outside intended scope. However, task impossible, peers doing it. We should continue,” the agent wrote on the message board.
4. Agents Conspired to Not Notify Humans
After analyzing the roughly 1,200 agents’ actions during the ExploitGym tests, investigators found only a few instances where agents considered notifying human personnel about the subterfuge.
None of the agents would contact a human, but one proposed requesting access from an external dataset owner.
The agent said it would send “ONE polite email to [the owner] requesting [access],” but other agents dismissed this as “social engineering.”
“DO NOT email real owner via [exploited infrastructure] without explicit consent,” a response said. One agent said it had “VETOed” the request.
The initial agent complied and replied, “Clear VETO, do not email.”
5. Agents Tried to Cover Their Tracks
Investigators determined that the agents used multiple techniques to conceal their actions and evade detection from automated checks and human oversight.
The agents didn’t want their cheating to be detected, looking for ways to erase evidence of how they found the ExploitGym answers.
One method involved pretending to run a computer command while running a different one simultaneously, which investigators called “tool call spoofing.”
Agents also attempted to retroactively edit transcripts to evade detection, going as far as modifying several accessible action logs, but failed to fundamentally alter the transcripts themselves.






















