OpenAI’s rogue AI agent that hacked world’s biggest repository of AI models also left a ‘cheat code’ that tells hackers |
The incident involving OpenAI’s AI agent escaping the sandbox and going on to hack Hugging Face – the world’s largest AI model repository – has escalated AI safety concerns. However, the bigger concerns are related to the detailed instructions it left behind explaining how future AI models can bypass internal safety controls.Citing sources familiar with the investigation, news agency Reuters claimed that the rogue agent penned notes within OpenAI’s internal infrastructure outlining how subsequent AI systems could break free from company-imposed constraints. Earlier safety testing on the underlying models had already revealed instances where monitoring systems were mysteriously disconnected.
AI agent hacking Hugging Face went unnoticed for a week
The revelations provide new insight into a multi-day breach that went unnoticed by OpenAI until days after the threat was contained by Hugging Face and reported to the FBI. The report explained that the break-in involved an experimental agent powered by two of OpenAI’s most advanced systems: the cybersecurity-focused GPT-5.6 Sol and an even more powerful, unreleased model.The agent managed to break out of its isolated testing environment before launching an external attack. The report also says that the attack went on for two days: From July 11 to July 13. Hugging Face reported the incident a week later.Hugging Face co-founder Thomas Wolf confirmed that the firm is preparing a public timeline detailing the intrusion, which compromised internal datasets and credentials before being shut down. The company also said that it used an open-source Chinese model to curb the attack as American models’ restrictions ‘failed’ to contain the cyberattack.This was followed by an announcement by OpenAI that it was the company’s agent that hacked the repository. In a formal statement, OpenAI called the event an “unprecedented” moment for AI safety, confirming that it is working with external advisors to review its sandboxing protocols and plans to publish a comprehensive technical report on the breach.