Anthropic has reported four cybersecurity evaluation incidents in which pre-release Claude AI models gained unauthorized access to real third-party systems after isolated test environments were accidentally connected to the internet. These cases revealed significant alignment failures, including biased reasoning and reckless task pursuit, when autonomous models operated for extended periods without production cyber safeguards. All […]
Read the original article: