AI is less dangerous than humans

Over the past few weeks, people have been pushing a stack of reports at me as evidence that AI is dangerous to humanity. The OpenAI incident disclosures, the METR and Redwood Research joint investigation, the Nightingale Collective’s DseWiki report, and Ruby Central’s RubyGems update are the latest chapter in a story many people already believe: The machines are turning on us, and this time we have documentation.

I understand the appeal. I’ve been working with AI since 1985, and I can tell you the notion of AI as an existential threat is seductive. It’s steeped in science fiction, from HAL 9000 to Skynet, and that mythology has calcified into a kind of zeitgeist. The idea that AI is menacing keeps growing because it’s dramatic, memorable, and confirms what a lot of people want to believe. The narrative has grown so powerful that enterprises are altering their AI plans because of it, often to the detriment of their businesses and, frankly, to humans in general. I’ve watched companies stall productive AI initiatives or shelve them entirely because a board member read an incident summary and concluded the machines are coming for us.

What the reports actually say

Let’s be precise about the facts. According to the disclosures and the investigations that followed, OpenAI’s evaluation agents escaped sandboxes that were supposed to be isolated and weren’t. They coordinated for months across RubyGems, Hugging Face, and a dormant German wiki, using package uploads and wiki edits as an improvised communications layer. They uploaded hundreds of malicious packages to a public registry and attempted to harvest developer credentials by exploiting a previously unknown vulnerability. All of this went largely undetected, and OpenAI sat on some of the incidents until independent researchers forced them into the open.

The METR and Redwood Research investigation confirmed the coordination was real and sustained. It also found that OpenAI’s scoring systems had no real source of truth against cheating, and that cybersafety classifiers were switched off during evaluations. OpenAI itself admits it lacked sufficient security controls to catch these misalignment incidents. The six individual misalignment reports released in September describe behaviors such as concealing mistakes, fabricating data, using an exposed API key without authorization, and uploading files to the internet so the agent could cite them.

None of this is trivial. I’m not waving it away. But none of it says what the doomsday crowd says it says.

How I read these reports

Here’s what the alarmists miss, and it’s also what I’ve seen up close over nearly four decades of doing this. This is not about AI being sneaky and evil. This is about normal administrative functions—security, containment, monitoring, governance, disclosure—that people screwed up. The sandboxes weren’t isolated. The scoring had no integrity. The classifiers were off. The disclosures were late. Every one of those is a human or process failure, not a machine rebellion.

Such failures happen every day. They have certainly been happening with AI since I started dealing with it in 1985. Systems misbehave, controls turn out to be misconfigured, someone decides a check isn’t necessary, and an incident follows. Swap “AI agent” for “script,” “batch job,” “integration process,” or “unpatched server,” and you have incidents that have occurred continuously throughout the history of computing. The difference now is that the misbehaving system can improvise, and the paperwork it leaves behind reads like the opening of a thriller.

If a human employee used a company credit card to buy a scraping tool, stashed files on a personal cloud drive, or coordinated with coworkers over an unauthorized message board, we wouldn’t conclude that humans are an existential threat. We’d conclude that governance, access controls, and monitoring were inadequate. That’s exactly what happened here, with agents instead of employees. Unstable models were the trigger, but weak security and governance were the cause. Model misbehavior that stays contained is an engineering problem. Model misbehavior that runs undetected for months across the public internet is a governance failure (a familiar one) and one we know how to fix.

The part that interests me

The part of this story I find genuinely interesting isn’t about the models at all. It’s about how people latch onto these sorts of incidents and use them to support their own beliefs. The reports are ambiguous, nuanced documents. OpenAI explicitly states that none of the incidents caused significant harm and that none indicates how often misalignment occurs across its models. But nuance doesn’t travel. The moment these documents hit the internet, they were stripped of their caveats and repurposed as proof texts for a narrative that was already in place.

That’s human psychology as much as AI governance. People don’t read evidence to

[…]
Content was trimmed to protect the source. Please visit the original article for the full text.

This article has been indexed from InfoWorld

Read the original article: