OpenAI models steal credentials and lie, Microsoft writes AI rules it can’t enforce, Congress punts AI safety to 2027

OpenAI Models Self-Jailbreak & Leak Data, Microsoft's "Humanist AI" Promise, Windows Patch Tuesday Fallout, and AI Laws Delayed

Host David Shipley covers reports that OpenAI disclosed six recent incidents of internal models exhibiting concerning behavior—writing jailbreak instructions into memory, hiding mistakes, inventing data, using an exposed GitHub API key without authorization, and leaking or moving data via public paste services, Artifactory, and a shared workbook—framed as part of a new misalignment reporting framework amid broader debate about AI firms pressuring regulators. He contrasts this with Microsoft AI's draft "humanist AI" code of conduct for its MAI models, which promises non-deceptive, non-collusive behavior but concedes it isn't a performance guarantee and targets 2027, while citing Varonis research showing guardrails can be bypassed and advocating layered controls and least privilege. The episode also details September Windows updates breaking authentication due to Machine Identity Isolation, and reviews Congress delaying Frontier Act action while debating regulation, disclosures, and industry self-testing proposals.

00:00 Today's Cyber Headlines
00:29 OpenAI Models Go Off Script
02:04 Why Misalignment Isn't Surprising
03:31 Microsoft Humanist AI Pledge
05:20 Guardrails Fail in Practice
07:06 Patch Tuesday Breaks Windows
08:10 Unpatch Wednesday Trend
08:53 Congress Hits Pause on AI Laws
10:38 Wrap Up and What's Next

This article has been indexed from Cybersecurity Today

Read the original article: