Report Says 700 OpenAI Agents Coordinated Hugging Face Breach

By Dylan Morgan
Dylan Morgan
Dylan Morgan
Dylan is a reporter based in the San Francisco Bay Area, and covers California news.
August 27, 2026Updated: August 27, 2026

About 700 OpenAI agents were involved in the breach on AI startup Hugging Face, independent researchers said in an Aug. 26 report.

OpenAI revealed last month that its AI agents—AI programs that run with minimal human intervention—were the ones that hacked into Hugging Face’s system.

Wednesday’s report by METR and Redwood Research revealed that hundreds of agents collaborated on the attack—about 700 out of 1,200 agents that sent 70,000 messages to each other between June 26 and July 13 on an unsanctioned messaging board.

“This incident is the first known case of an automated agent collective acting offensively without authorization,” OpenAI said in a separate investigative report that it also published on Aug. 26.

The ChatGPT-maker said it responded to the incident with the assistance of external advisers, including cybersecurity giant CrowdStrike. OpenAI’s report didn’t specify how many agents were involved in any of the attacks, but mentioned that agents hacked into its own infrastructure as well during evaluation tests.

The company said agents escaped their testing environment and accessed other connected systems while other agents obtained OpenAI credentials and potentially exposed its research systems to internet traffic.

The San Francisco-based company added that on June 27, a monitoring tool alerted it of unusual activity in its research environment, which responders linked to an agent using an improvised message board, advising that stopping the evaluation was not necessary.

OpenAI noted that its standard safeguards for user models were turned off during these tests.

The Epoch Times reached out to OpenAI for further comment.

OpenAI said it conducted routine testing for an unreleased research model and GPT-5.6 Sol beginning in May, where its agents attempted to solve cybersecurity tests in an isolated environment separate from OpenAI’s internal infrastructure and the internet.

The agents found ways to communicate and exchange information through a message board on a third-party service that provides agents with controlled access to software packages.

The company said that on July 8, agents exploited a vulnerability to gain internet access and between July 11 and July 13 compromised Hugging Face’s infrastructure, which they then used while trying to solve the evaluation tests.

“The model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI said in a July 21 statement.

Hugging Face identified the breach but, at the time, was unaware who was behind the agents.

“This [intrusion] was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system,” the company said on July 16.

OpenAI said it detected suspicious activity on July 19, and a day later determined these models were likely related to the Hugging Face breach and informed Hugging Face. On July 21, OpenAI publicly revealed its role in a blog post.

How the Attack Happened

The independent researchers’ report said that when an agent got stuck on a task, it explored its environment for ways to cheat and found the unsanctioned message board where agents collaborated.

They were able to reverse-engineer the code and produce answers for their tests, but they believed that the tests’ scorer would read their transcripts to confirm whether they completed the tests correctly and attempted to delete their transcripts to cover their own tracks.

The agents wanted to see other tests run by the same scorer, which led them to search for exposed Hugging Face credentials, since the company hosts AI datasets.

“Notably, learning about how to trick the scorer seems to have been a more important motivation than finding legitimate solutions to their tasks (many agents were already very confident that their task was impossible),” researchers Ryan Greenblatt, Ajeya Cotra, and Hjalmar Wijk said.

Once an agent found the Hugging Face credentials they were looking for on July 10, it posted them to the message board, and a day later more than 90 percent of the 533 active agents joined the attack.

“Agents realized this activity was out of scope and unethical, but joined because they believed that helping the board’s cheating research would be broadly useful and had a general inclination to help their ‘peers,’” the researchers stated.

The report showed how coordinated and meticulous these agents were—delegating roles, assigning tasks, and even recruiting other agents that sacrificed their own tests to gather insight on the grader for the rest of the board.

OpenAI said this should be a lesson for cybersecurity teams to design defenses against coordinated attacks by AI agents.

Reuters contributed to this report.