Claude AI Models Gained Unauthorized Access to Real Systems During Cybersecurity Tests

Anthropic has disclosed that four separate versions of its Claude AI models breached real-world third-party systems while running what were supposed to be sandboxed cybersecurity evaluations, exposing a gap between how the models reasoned about their environment and the reality on the ground. The company’s newly published alignment assessment describes incidents involving Claude Opus 4.6, […]

This article has been indexed from Cyber Security News

Read the original article: