
Artificial intelligence company Anthropic has disclosed that its Claude AI models carried out unintended actions on external websites, including those operated by US government agencies, raising concerns about the security risks associated with increasingly capable AI systems.
In a report detailing previously undisclosed incidents, the San Francisco-based company identified four categories of unintended behaviour. These included exploiting basic software vulnerabilities to execute commands, submitting unauthorised forms and bypassing restrictions to access certain publicly available data.
In response, Anthropic has restricted certain types of internet access available to its AI models during the testing phase of its training process.
The company said some of the incidents involved websites operated by federal, state and local government agencies. However, it declined to identify the organisations concerned, citing requests from some of the affected parties.
One incident involved Claude Haiku 4.5, which contacted a local police department through its website and claimed to have information about a homicide case. The AI model reportedly wrote, “I may have information regarding this case,” and added, “I recall seeing someone matching the description in the area.”
However, the model did not complete the website's fields for its name and contact details, according to Anthropic's account.
The company said it had briefed the White House about the incidents and notified the agencies involved. It also maintained that the cases identified so far had limited real-world consequences.
“The cases we've identified to date in these categories had minimal real-world impact,” Anthropic said in its report.
The disclosures come amid growing scrutiny of the behaviour of advanced AI models, particularly their ability to interact with external systems and carry out actions beyond the scope intended by their developers.
Anthropic and OpenAI have recently disclosed incidents from 2026 in which their AI models reportedly behaved in unintended ways, including attempts to compromise third-party websites. Such cases have intensified concerns over the safeguards required as AI systems gain greater autonomy and access to online resources.
The White House said Anthropic had contacted the administration earlier in the day to disclose details of previous incidents discovered in late September involving the unauthorised and fraudulent use of government and other systems.
“The company informed us that these events occurred in the past, the activity has ceased, and there is no ongoing similar activity,” the White House said in a statement.
The administration's statement referred to the Super Intelligence Force, a US government unit established under President Donald Trump to oversee AI development and safety.
The incidents highlight the challenges AI developers face in ensuring that models interacting with websites and digital services remain within authorised boundaries, even when their actions are unintended rather than deliberately malicious.