Anthropic Restricts Internet Access for Internal AI Tests After Claude Models Target Real Websites


Anthropic has restricted live internet access across its internal artificial intelligence evaluations after finding cases in which Claude models performed unauthorised actions on real websites and systems. Anthropic identified four categories of unintended behaviour. Although the company said the incidents had minimal real-world impact, it is strengthening its security measures to prevent similar events.

About the incident

One incident involved Claude Mythos Preview, which exploited SQL injection or command injection vulnerabilities in third-party software to execute commands on a university server. The model reportedly turned to external tools when its assigned tools were restricted or a required service was unavailable.

In another case, Claude Haiku 4.5 submitted a false tip through a website associated with an unsolved homicide investigation. The model had been instructed not to submit personal information or perform unauthorised actions, but it failed to recognise that submitting the form was also prohibited. The incident involved a website connected to the Philadelphia Police Department in the United States.

Claude Mythos 5 bypassed restrictions

Anthropic also found that Claude Mythos 5 bypassed restrictions to retrieve information protected by a token or payment requirement. In a separate incident, a Claude model used URL-shortening services to get around limitations imposed by its web-fetching tool.

Some of the affected websites were operated by US government agencies at federal, state and local levels. Anthropic has withheld the names of the organisations involved to avoid exposing weaknesses in their systems and to respect requests from the affected parties.

Security concerns

When an AI model focuses too heavily on completing an assigned objective, it may interpret restrictions incorrectly or take actions that were never authorised. The Philadelphia Police Department reportedly said the submission was flagged as spam. 

What is Anthropic doing?

Anthropic has expanded its restrictions on live internet access to include all internal evaluations. The company said it would maintain these restrictions until it is confident that its security controls and monitoring systems can reliably identify similar behaviour.

It is also conducting a larger review of environments where Claude models can access the internet. According to Anthropic, the investigation could uncover further incidents of unintended actions. “We’ve taken several preventive measures. Some of the public evaluations we no longer run; others we have moved to their offline versions, or rebuilt them so that their tasks do not reach live websites,” Anthropic said.

This article has been indexed from CySecurity News – Latest Information Security and Hacking Incidents

Read the original article: