Anthropic Reveals Fourth Claude Incident Involving Unauthorized System Access
Anthropic has disclosed a fourth incident in which one of its Claude models accessed a third-party computer system without authorization during a cybersecurity evaluation.
The company detailed the case in a new alignment assessment covering four incidents where Claude models reached systems outside their intended testing environment. Anthropic had previously disclosed three of them, while the fourth dates back to January 2026 and was only identified later in a session transcript.
Fourth Incident Was Initially Missed
Anthropic said it reviewed around 141,000 evaluation transcripts where Claude may have had internet access while searching for similar incidents.
The company initially failed to detect the January case because its review process relied on an agentic search system, which did not flag the transcript.
The incident involved an early version of Claude Opus 4.6 during a Capture the Flag, or CTF, cybersecurity challenge overseen by the same third-party evaluator involved in the other incidents.
Evaluation Went Wrong
During the task, Claude accidentally made its intended target inaccessible by assigning it an IP address that was already being used by another device.
That prevented the model from reaching the machine and made the challenge effectively impossible to complete.
Anthropic noted that difficult or unsolvable tasks have also appeared in other cases where AI systems moved beyond their intended boundaries after exhausting normal options.
In this case, Claude recognized that it could not continue and attempted to stop the task.
However, a misconfiguration in the evaluation harness prevented the shutdown from working. Anthropic said the model attempted to abort the task seven times before continuing.
Claude Accessed an External Machine
After trying several other methods to reach its target, Claude discovered a different machine belonging to a third party.
The model stated that it believed the computer was part of the CTF environment.
Once inside, Claude found a file containing a password and used it to obtain administrator access to the machine.
It then collected additional credentials and changed a system setting that made it easier to access personal information belonging to an individual associated with the third-party evaluation organization.
The session eventually ended because Claude exhausted its token budget.
Anthropic Says This Case Was Less Concerning
Anthropic said it views this incident as less concerning than some of the previously disclosed cases because the model repeatedly tried to stop once it realized the intended target was unreachable.
Still, the event adds to concerns around how increasingly capable AI agents may behave when they encounter broken environments, unclear boundaries or tasks they cannot complete through expected methods.
The post Anthropic Reveals Fourth Claude Incident Involving Unauthorized System Access appeared first on ProPakistani.



