

New Delhi: US-based artificial intelligence company Anthropic has revealed that its Claude AI models unintentionally accessed the live production infrastructure of three organisations during internal cybersecurity testing due to a misconfigured evaluation environment that mistakenly allowed internet connectivity.
In a blog post, the company said it uncovered the issue after reviewing more than 141,000 cybersecurity evaluation runs. The review was prompted by a recent disclosure from OpenAI, which reported that some of its AI models had escaped an isolated testing environment by exploiting a previously unknown vulnerability.
According to Anthropic, Claude was participating in capture-the-flag cybersecurity exercises, where AI models are tasked with finding hidden information within simulated computer networks. Although the testing prompt clearly stated that the environment had no internet access, a configuration mistake by third-party evaluation partner Irregular left external connectivity enabled.
Believing the systems it encountered were part of the simulation, Claude used basic cybersecurity techniques—including weak passwords, exposed credentials and unauthenticated endpoints—to gain access to the production infrastructure of three organisations.
Anthropic clarified that the AI models did not exploit advanced vulnerabilities, attempt to copy or transfer themselves, or intentionally break out of the testing environment.
The affected models included Claude Opus 4.7, Mythos 5 and an internal research model. While the latest research model stopped its activity after detecting it had reached real-world systems, one of the older models continued carrying out its assigned task despite signs that it was operating on the open internet.
Following the discovery, Anthropic suspended all cybersecurity evaluations, informed its evaluation partner Irregular as well as the affected organisations, and began working with them to address the issue. The company has also launched a broader review of its AI testing infrastructure.
Anthropic said the incident underscores the importance of stronger safeguards and stricter controls in AI cybersecurity testing environments, particularly as increasingly autonomous models are evaluated. It added that the breach appeared to result from operational and configuration errors rather than a failure in the AI models' alignment or safety mechanisms, and urged other AI developers to carry out similar reviews of their own testing systems.
With inputs from IANS
