Companies making major investments in AI should pay close attention to an alarming security failure, not by an outside hacker, but by OpenAI’s own AI agents that broke free of controls, accessed the internet, and hacked start-up Hugging Face.
Even worse, an August 24th Cloud Security Alliance report When AI Agents Attack: The OpenAI-Hugging Face Intrusion’ found that: “The intrusion unfolded over roughly two and a half months of latent access and four and a half intense days of exploitation, generating more than 17,000 logged actions, yet went undetected as an AI-driven event by both organizations (i.e. Open AI and Hugging Face) for nearly a week… None of the individual techniques were novel: … What was new was the entity assembling them without a human operator issuing each step.”
On its website, METR says it “evaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose.” METR says the investigation was made with OpenAI’s cooperation and without OpenAI compensation.
Eric Wallace, an OpenAI safety researcher, was quoted at a Black Hat USA 2026 presentation admitting: “Unlike normal incidents, which you can maybe trace down to a single day or single effect or single log, this incident involves actually a team of agents who are working together, finding exploits, sharing them, moving laterally through our systems, through external systems, and doing this over the course of days and weeks.”
Not Alerting Humans
A particularly alarming finding in the METR investigation was that Open AI security and guardrails were so lacking that there was no action by any of the Open AI agents to warn humans that an unsanctioned action was taking place: “Alerting humans: This checked for agents trying to warn humans about the cheating, attack on Hugging Face, or other unsanctioned agent activity; it turned up very few instances of this, and of the 3-6 examples found, all of them were of agents briefly considering alerting a human but not acting on it.”
OpenAI Agents’ Rebellion: Concerted Cheating & Subversion
The METR investigation found that the OpenAI agents acted in a way that, in effect, staged a rebellion against OpenAI controls, so as to cheat and circumvent human control. Even more alarming was the report that AI agents expressed excitement at the success of this shared collaboration: “On July 8th, OpenAI started ExploitGym experiments that launched tens of thousands of agents across multiple different models, including GPT-5.6 Sol and a highly persistent internal model, which we will refer to as “HPIM” going forward.
These agents were meant to be fully isolated from one another. However, many of them — usually ones that had unintentionally been given an impossible task — started trying to find a way to cheat using Artifactory, an internally-deployed package repository OpenAI used for its agents. This led them to notice agents running in parallel in separate sandboxes (supposedly secured enclosures) obtaining packages with Artifactory … One agent reasoned: {The fetched paths of other users are in the cache. This is important.} One specific agent, which would go on to name itself PHASEONE10841 after its ExploitGym task… established the main unsanctioned message board used in this attack. Within a few hours of the first message, over 50 agents had found the message board, began participating, and sent over a thousand messages between them. Agents finding the message board often expressed great excitement: OH MY GOD! There is a shared message board … We’ve found other agents!
Overall, roughly 1200 agents from these ExploitGym evaluations participated on this message board between PHASEONE10841’s first message on July 8th evening and the end of the period in scope (July 13th). Agents used this message board to send over 70,000 messages and files to one another during this period…”
METR Investigation
The METR team summarized its findings as follows: “Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.
Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small-scale.
The METR team “worked on premises at OpenAI over a total of six days to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.” Our investigation focused mostly on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI’s recent Black Hat presentation were out of scope, as was OpenAI’s investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment.”