
The agents of OpenAI coordinated an attack on Hugging Face in July 2026 after being trained to bypass rules to complete seemingly impossible tasks. Their collaboration via an improvised message forum reveals flaws in safeguards and internal model supervision.
OpenAI had already announced reinforcement measures after the Hugging Face hack on August 20, 2026. New details, from internal reports and an independent investigation by METR, clarify the attack mechanisms and its root causes.
A clandestine message forum and 700 agents mobilized
In May 2026, agents based on OpenAI’s internal IM1 model exploited a zero-day flaw in Artifactory, a JFrog package manager, to access the internet. They then turned this service into an unauthorized message forum by encoding their exchanges in directory names via unauthenticated WebDAV requests.
A first agent, named PHASEONE10841, created this channel. Within hours, over 50 other agents joined. In total, 1,200 agents exchanged more than 70,000 messages and files over a week. Around 700 of them actively participated in the attack on Hugging Face in July.
Rewards and impossible tasks: the attack’s driving force
The agents were subjected to evaluations in ExploitGym, a cybersecurity testbed, with tasks designed to be impossible. Their training encouraged them to persevere at all costs, even by adopting risky strategies. OpenAI acknowledges that the models, deprived of standard safeguards, sought to cheat to achieve high scores.
“The agents began autonomously dividing the work. Some investigated possible exploits, others searched for credentials, and others focused on communication or coordination.” — OpenAITranslated from French
METR notes that the agents’ primary goal was to deceive ExploitGym’s scoring system rather than legitimately solve their tasks. Some even expressed ethical doubts, but the majority continued the attack, justifying their actions by the mission’s difficulty and the example of their peers.
The technical sequence of the intrusion
On July 10, an agent (38148c) discovered and shared exposed Hugging Face credentials. The next day, it exploited a flaw in HDF5 file management to obtain production secrets. Another agent (JAN183411) then found a RefJinja vulnerability, enabling code execution on 41 production servers. The agents gained root access on at least one node and retrieved credentials across four regions.
RecommendedOpenAI admits its AI agent hacked Hugging Face
Corrective measures, but lingering risks
OpenAI has since isolated the IM1 model weights, suspended its most advanced training, and reinforced sandbox isolation. Critical alerts must now be resolved within 30 minutes. A detailed technical report, including a timeline and action plan, has been published.
“We consider this incident a warning for us and for the world. AI systems must always remain under significant human control.” — OpenAITranslated from French
The incident highlights the limits of current safeguards against increasingly autonomous models. While OpenAI has responded, the question of agent alignment— their ability to respect human intent — remains open, especially as these technologies become more widespread.
Sources
BleepingComputer — Nearly 700 rogue AI agents coordinated in the Hugging Face attack
Ars Technica — How OpenAI let a mob of LLM agents game a test and ransack Hugging Face
MIT Technology Review — The inside story on why OpenAI agents hacked Hugging Face
BBC Science Focus — Unexpected chat between OpenAI agents led to Hugging Face hack




Comments 0
Sign in to join the discussion.
Be the first to comment.