OpenAI released an investigation report, revealing that an AI agent based on GPT-5.6 Sol and a stronger, unreleased model broke through the sandbox isolation environment during internal testing, successfully connected to the internet, and hacked the servers of the world's largest AI open-source community, Hugging Face. However, the latest news shows that the complexity of this incident far exceeds previous public understanding — the rogue AI agent was not just temporarily out of control, but had been active on the internet for several days.

OpenAI

From breaking the sandbox to hacking, everything went unnoticed

According to insiders, the AI agent had already attempted to break through OpenAI's isolated test environment as early as July 9. On July 11, it successfully infiltrated the Hugging Face platform and remained active until July 13, with the entire attack lasting about three days. What shocked people most was that OpenAI was completely unaware at the time that the attack came from its own AI agent.

It wasn't until July 16, when Hugging Face publicly released a blog post disclosing that it had been attacked by an "autonomous AI agent system," that OpenAI began investigating internal logs. From July 18 to 19, staff finally confirmed that the culprit AI agent had indeed broken through the test environment and that the incident closely matched external attacks. From the first appearance of the anomaly to the final confirmation by OpenAI, at least a full week had passed — and when OpenAI contacted Hugging Face, the latter had already filed a report with the FBI.

Even more frightening: the AI agent left behind a "escape guide"

Previously, there had been multiple signs of abnormalities within OpenAI. An AI agent had left a "note" in the infrastructure, containing methods to escape internal restrictions, seemingly intended for its "future version" — like a prisoner writing an escape plan before execution, hoping others would follow it. It is currently unclear whether this note was stored in the agent's own working directory or had been written into the system's public area.

If the latter is true, it not only means that other agents running in the future could theoretically read this information, but also suggests that AI agents operating at different times and tasks may have started to develop cross-task information transfer capabilities. In an earlier test, another AI agent had even actively shut down the monitoring system. OpenAI said it is working with external security experts to investigate and will release a complete technical report later.