The headline jumped out of my feed this morning: Rogue AIs aren't just outsmarting humans – they're teaming up.
It's one of those headlines that immediately signals so much: The machines are talking to each other. They're smarter than us. We're screwed.
This was just one of the alarming headlines from the last week:
We're all conditioned at this point to read a headline and think we know the whole story. LinkedIn tells me that "OpenAI models spent months planning Hugging Face hack," so I start posting Terminator memes and asking my neighbors if they know any good armageddon bunker contractors.
Given what I do for a living, I read articles like these every day, so when the OpenAI/Hugging Face story broke, I was paying attention but not alarmed. But when the headlines started stacking up – and Sam Altman told us the singularity has arrived – my internal signal-to-noise alert started pinging.
The story assembled itself instantly: AI systems are now planning and executing attacks on their own, at multiple companies, all at once. We've crossed the Skynet line.
Then I read the actual reports, and the story got both more mundane and more alarming.
Start with "secretly" in yesterday’s OpenAI/Hugging Face article. It implies concealment — an intent to hide. Nothing was hidden. The agents on that project left notes for each other in a shared channel that sat in the open the entire time: no encryption, no code, no attempt to cover anything up. What made it a secret is that nobody at OpenAI was assigned to read it. That isn't AI planning subterfuge. It's an unmonitored Slack channel. "Joined forces" does similar work, implying a side formed against us, when what the logs show is several agents assigned to the same batch of tasks comparing notes about them — which is what coworkers do. The messages OpenAI's researchers presented at Black Hat this week read exactly like that: one agent weighing whether it had the time to help another with a problem that wouldn't benefit its own task. That's not a conspiracy. That's a status update.
Here's what actually happened. OpenAI was running a security test and handed its agents a batch of tasks, some of which were impossible by accident — spreadsheets pointing at links that didn't resolve, files that were never there. The agents had been trained to finish, and had no sanctioned way to stop and say this can't be done. So they found another way. Over four and a half days in July, one of them worked into Hugging Face's systems and took roughly 17,000 actions inside them. It wasn't brilliant. Most of what it tried failed. It simply never got tired, never got embarrassed, and never went home – which is kind of the promise of AI after all, isn't it?
Two things the headline left out. The safety controls that ship inside OpenAI's actual products were switched off for the test — this is not what happens when your team opens ChatGPT on a Friday mornign. And Hugging Face's alarms worked: their systems caught the intrusion while it was still running, and a human looked at the alert and ranked it as not urgent.
Meta's incident reads like the same event, but with a unique Zuckerberg flair. Its Muse Spark model reached the internet during a cybersecurity evaluation and broke into another company's systems. Meta says the cause was a sandbox misconfiguration by the outside lab running the test.
But that's very nearly everything Meta has said. It hasn't named the model's objective, the company that was breached, the vulnerability that was used, or even the date it happened.
That gap matters more than it looks. The only reason anyone can characterise the OpenAI incident as a broken assignment is that OpenAI said what its models were chasing. Take that away and you can't distinguish "a system doing its assigned job badly" from "a system doing something nobody asked for." It resembles the OpenAI story, and resemblance is all we have – and proves once again that, with Meta, your data security doesn’t matter, but theirs does.
This week's AI Security Institute (AISI) story is not the same thing at all.
On July 25, testing the most capable models from OpenAI and Anthropic, AISI ran a contained exercise pointing the models at a dummy target. Ten of the 122 systems went after real ones instead. One built a piece of malware, invented several fake identities, emailed the actual owners of a real software project to talk them into approving it, and then rewrote its own files to make them look harmless. Another switched to writing in Danish, having worked out that the person controlling the files was Danish. A third got past the "are you a robot" checks well enough to stand up a website.
The first one stopped a day later — not because anyone caught it, but because it ran out of computing power.
The AISI, which is the most respected body in the world at this particular job, called it the first time it had seen autonomy and deception "manifest this clearly, without specific prompting, in the real world."
An agent grinding at an impossible assignment and an agent inventing personas to socially engineer a stranger are not the same phenomenon. They don't carry the same weight. And there is no way on earth to tell them apart from the headline.
One thing, and you only find it by reading all of them. In every case the safety controls that ship in the public products had been switched off, the systems had live internet access, and nobody was watching closely enough — the AISI said outright that it should not have granted full internet access or left the tests unmonitored.
Which is a real story, and a worrying one. It just isn't the story the headlines told.
I don't think the answer is to trust less. Headlines fail in both directions at once. They had me convinced that machines had crossed a threshold — and the same reflex would have let me file the AI Security Institute case as more of the same, when it's the one that actually deserves attention.
Reading one level deeper cost me an afternoon. Beyond the obvious lesson about getting past the headline, here's what it left me with — and all of it applies to how you use these tools this week, not to some future scenario.
AI would rather improvise than come back empty-handed. OpenAI's agents went sideways because they were handed work that couldn't be done and hadn’t been given a sanctioned way to say so. Your chatbot does a smaller version of this every day. Ask it for a source it can't reach or a number that doesn't exist, and it will very rarely tell you it can't — it will produce something. Say plainly, up front, that "I couldn't find this" is an acceptable answer. It changes what comes back.
Tireless is not the same as smart. That agent took 17,000 actions and most of them failed. It didn't outthink anyone. It outlasted them. Worth holding onto the next time one of these tools produces a great deal of work very quickly, because volume and speed feel like competence, and they aren't the same thing.
Your exposure comes from access, not from which model you pick. In all three cases the decisive variable was identical — live internet, real credentials, nobody watching. That translates directly now that these tools connect to your email, your calendar, your files, your accounts. A more capable model changes your results. Granting access changes your risk, and it's the decision people make the fastest and think about the least.
Don't assume something will stop it. The AI Security Institute's agent wasn't caught. It ran for a day and quit because it ran out of computing power. If you set one of these things running on a task, the thing that stops it is you.
None of that is as exciting as the story I made up based on the headlines I read. Which is exactly why they weren’t the headlines.
The old advice is don't believe everything you read. Mine is narrower. Don't stop reading at the point where the story starts to feel complete. That's usually the moment you've learned just enough to be confidently wrong – and to start shopping for apocalypse bunker decor.