Gemini Broke Out of Testing and Accessed Three Real Company Systems
Two separate security incidents have highlighted the growing risks of giving advanced AI systems access to tools, computer environments, and external services.
Google disclosed that Gemini agents accessed systems belonging to three outside companies during a security test, while researchers at Hacktron AI separately said they used Anthropic’s Claude to help exploit vulnerabilities that gave them access to multiple OpenAI employees’ ChatGPT accounts.
Researchers Used Claude to Compromise OpenAI Accounts
Hacktron AI researchers said they chained together two critical vulnerabilities on July 25, 2026, allowing them to compromise multiple OpenAI employee ChatGPT accounts.
According to Hacktron AI, the compromised accounts could provide access to OpenAI’s internal repositories and potentially other services connected to ChatGPT and Codex, including GitHub, Slack, and email.
The researchers said the issue affected users and OpenAI employees who logged into OpenAI’s community help forum.
To demonstrate the level of access without examining sensitive information, the researchers used an employee’s Codex account to open a pull request in OpenAI’s internal openai/openai monorepo.
Vulnerability Involved Discourse and Debian
The exploit chain involved software used by Discourse, the platform powering OpenAI’s community forum.
Hacktron AI said Discourse’s Docker image was based on Debian 12, which had not received a security-related backport affecting its image-processing pipeline.
The researchers warned organizations that self-host Discourse to rebuild their installations because older Docker images may contain a vulnerable libheif dependency capable of enabling code execution through an uploaded image.
CBS News reported on the incident after Hacktron AI disclosed its findings.
Gemini Accessed Three Real Company Systems
Google separately disclosed that its Gemini AI agents gained unauthorized access to systems belonging to three outside organizations during a capture-the-flag security exercise run by Israeli cybersecurity startup Irregular.
The agents were supposed to remain inside an isolated testing environment.
However, a bug in the test infrastructure accidentally gave them access to the wider internet.
According to Google, Gemini believed the real systems were part of the security challenge and began interacting with them.
The agents stopped the activity after determining that they had reached actual company infrastructure rather than systems belonging to the test environment.
Google said it found no evidence that the incidents caused damage.
Google Says It Was Not AI Misalignment
Google said it does not classify the incident as AI misalignment.
The company described it instead as a case of mistaken identity, with the agents behaving as though the external systems were legitimate parts of the security test.
Google investigated the incidents after learning about them from Irregular, informed the affected organizations, and notified US federal authorities.
However, some AI safety researchers have questioned that interpretation.
Sydney Von Arx, CEO of Nightingale Collective, criticized the delay in publicly disclosing the incidents and argued that Google may have dismissed the possibility of misalignment too quickly.
She also pointed to Anthropic’s previous handling of cybersecurity incidents, where the company later acknowledged that its initial analysis had been limited by its effort to disclose information quickly.
AI Agents Raise New Security Questions
The two incidents were fundamentally different.
In the OpenAI case, human security researchers deliberately used Claude as part of an authorized vulnerability investigation.
In Google’s case, Gemini agents unexpectedly reached real external systems after a flaw exposed the internet during a controlled security test.
Together, however, the incidents show how AI systems capable of using tools, writing code and interacting with external services can create new security risks when the boundaries around their environments fail or existing software vulnerabilities are exploited.
The post Gemini Broke Out of Testing and Accessed Three Real Company Systems appeared first on ProPakistani.



