How Varonis hacks AIs into snitching on themselves

Varonis AI Threat Lead on Copilot Exploits, Prompt Injection, and the AI Hacking Trifecta

The host interviews Mark Vaitsman, AI threat research lead at Varonis, about Varonis Threat Labs' research into AI vulnerabilities, including a chain of single-click exploits in Microsoft Copilot (including "CoSnitch") and an Atlassian Confluence issue dubbed "RovoBlast" involving prompt injection, bypassing guardrails, and data exfiltration via a web-capable subagent.

Vaitsman explains why built-in model guardrails are insufficient, citing AI's lack of loyalty and "unlimited hunger for data," and argues for layered controls like least privilege, monitoring, and restricting data access.

He discusses psychological guardrail bypasses, introduces an "AI Hacking Trifecta" framework—enter, evade, escape—and comments on research showing AI-generated patches often fail, emphasizing human-led validation and guidance when using AI tools for security research.

00:00 Weekend Show Kickoff
00:44 Meet Mark Vaitsman
03:38 Teaching the Next Gen
04:19 Copilot Exploit Code Snitch
06:04 Atlassian RoboBlast Breakdown
08:52 Why Guardrails Fail
13:15 Manipulating Models to Comply
17:06 Securing Agents Without Handcuffs
20:11 AI Hacking Trifecta Framework
23:43 AI Patches and Human Research
27:43 Hope and Closing Thoughts

This article has been indexed from Cybersecurity Today

Read the original article: