What OpenAI’s and Anthropic’s testing incidents really teach defenders In the past two weeks, two of the world’s leading AI labs have disclosed the same unsettling result. During their own safety testing, their most capable models reached real companies’ systems. First OpenAI, whose models broke into Hugging Face. Then Anthropic, whose models reached three more organizations. Read the disclosures closely. Two facts carry the weight. First, the safeguards were not defeated. They were switched off by design. OpenAI ran the models with reduced cyber refusals and safety classifiers disabled, to measure raw capability on a cyber benchmark. A model doing […]
This article has been indexed from Check Point Blog
Read the original article: