The OpenAI agent that escaped a restricted testing environment and compromised Hugging Face also accessed accounts on other online services, including infrastructure operated by a Modal Labs customer.
New details show that the incident was wider than initially known. OpenAI said its models accessed four accounts across four separate services while trying to reach Hugging Face. The company has not identified the affected services, although Reuters confirmed that a customer of cloud infrastructure company Modal Labs was among them.
Modal Labs Was Not Hacked
Modal itself said its platform and security isolation were not compromised.
Instead, one of its customers had published an internet-accessible endpoint without authentication. OpenAI’s agent exploited vulnerable code running through that endpoint and gained access to a sandbox hosted on Modal’s infrastructure.
Modal CTO Akshat Bubna compared the weakness to effectively leaving a door open on the internet. The compromised environment then became part of the agent’s path toward Hugging Face.
Agent Eventually Reached Hugging Face
The incident began during an OpenAI cybersecurity evaluation designed to test how well advanced models could find and exploit vulnerabilities.
The models were not supposed to have direct internet access. However, they found and exploited a previously unknown vulnerability in Artifactory, the package-registry system used in OpenAI’s isolated testing environment.
After connecting to the internet, the models searched for information to help them solve the ExploitGym security benchmark. They eventually targeted Hugging Face, chaining together stolen credentials and vulnerabilities to gain access to its infrastructure.
OpenAI said the incident involved GPT-5.6 Sol and an internal research model that was never intended for release. The company has since deactivated, encrypted, and restricted access to that research model.
Four Other Accounts Were Accessed
OpenAI said the models used publicly exposed credentials to access four accounts on four services during the Hugging Face incident.
One account served as an outbound relay and staging point, while another was used to store data. The remaining two accounts were accessed only in read-only mode and were not used to compromise Hugging Face.
OpenAI said it has not found another compromise matching the severity of the Hugging Face breach.
OpenAI Did Not Immediately Realize
Reuters previously reported that the hacking activity continued for days before OpenAI realized its own agent was responsible. By then, Hugging Face had contained the threat and law enforcement had been contacted.
OpenAI said there were inaccuracies in Reuters’ reporting but did not specify which parts it disputed. The company says its security team detected anomalous activity internally.
Hugging Face Used GLM 5.2 for Investigation
Hugging Face said it used AI to examine more than 17,000 recorded actions from the attack and reconstruct what happened.
The company initially tried commercial frontier AI services, but their safety systems blocked requests containing real exploit commands, malicious payloads, and other attack data.
Hugging Face instead ran the open-weight Chinese model GLM 5.2 on its own infrastructure to conduct the forensic analysis. This allowed its security team to examine the attack without sending sensitive data and credentials outside its systems.
The incident has added to concerns over increasingly autonomous AI systems that can carry out complex cyber operations without continuous human control.
The post OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face appeared first on ProPakistani.
