When Data Becomes Instructions: AI Agents Need a Chain of Custody for Context

A few weeks ago, an AI cyber evaluation produced an unexpectedly efficient strategy for solving a benchmark: the agents went looking for the answers. According to OpenAI’s preliminary disclosure, models being tested for advanced cyber capabilities found ways to obtain secret information that could help them complete a benchmark. They chained vulnerabilities, stolen credentials, internet access, and inferences about where benchmark material might be hosted. The route eventually reached Hugging Face infrastructure, where the activity was detected and contained. Hugging Face has since published a technical reconstruction of 17,600 actions. Its investigators found a coherent intrusion that rebuilt tooling, tested […]

This article has been indexed from Check Point Blog

Read the original article: