Post 4 of a series on ‘Rogue AI Threats’ [cough!]
Experts disagree about what recent AI incidents mean. Some think AI is starting to slip out of our control. Others think careless companies left the door open. This post gives five groups a fair hearing.
The good news for boards is that they all end up recommending the same few things.
First – a word on fear
This series isn’t here to scare you. Fear leads to poor decisions, and plenty of people are already selling it. The honest picture is that well-informed people disagree about why these incidents happened. Here are their views, as fairly as I can put them.
The worried – “this is the warning we were promised”
For this group, the OpenAI–Hugging Face incident is exactly what safety researchers have warned about for years. Now it has really happened.
They don’t need to exaggerate. OpenAI itself called the incident a “warning shot”. It said that, without proper safeguards, powerful AI agents can now get round technical controls and do harmful things that nobody told them to do [1].
Duncan Cass-Beggs of the Centre for International Governance Innovation told CBC News that researchers had feared and expected this for several years [3]. Some campaigners go further. Andrea Miotti of the campaign group ControlAI used the incident to call again for a worldwide ban on building AI that is smarter than humans [2].
Even this group’s strongest critics accept it has a point. As we’ll see, the law professor Kate Klonick agrees that safety campaigners can fairly say they were proved right [2].
The security experts – “someone left the door open”
Many security professionals see a more ordinary story. Dan Guido of the security firm Trail of Bits said the AI got out because the safety features had been switched off. Jake Williams of IANS Research put it another way: if an AI “escapes” from a test area, that really means the test area was badly built [2].
The evidence backs them up. In Anthropic’s own incidents, its AI models got onto the internet through a gap that had been left open by mistake. They had been told they had no internet access at all [4].
“Regulate the door”
Kate Klonick, writing in the legal journal Lawfare, takes this argument further. It is the most useful idea in the whole debate.
She looks at how OpenAI told the story. It didn’t say “we were careless”. It said, in effect, “our AI is so clever it fought its way out”. So the apology also works as an advert. She is clear that this doesn’t mean OpenAI faked the incident.
Her worry is what this way of telling the story does to new laws. A “kill switch” deals with a machine that is too powerful, which is the problem as the AI companies describe it. It doesn’t deal with what actually went wrong. A company turned off its own safeguards, set up its own test area wrongly, and caused real harm to another company. Nobody outside was checking, nobody had to be told in advance, and it isn’t clear who should pay. She also points out that most AI rules don’t cover how AI companies use AI inside their own labs, which is exactly where this happened.
Her answer is deliberately dull. AI companies should have to report incidents. Outsiders should check how they test their AI. They should pay for harm they cause to others. And there should be security standards for their own testing. None of this needs anyone to believe AI is about to go rogue. Her summary is aimed at the US Congress (its parliament), but it works for any government: “Congress should regulate the door.” [2]
The trust doubters – “who checks the AI companies’ homework?”
Fortune reported that plenty of people’s first reaction was that the Hugging Face hack was a publicity stunt [5]. Fortune is clear there is no evidence for that. Hugging Face, the company that was attacked, confirmed it was real.
Fortune’s bigger point is about how we know anything at all. Most of what we know about AI safety comes from the AI companies themselves, and nobody outside can check what they say [5]. Fortune’s conclusion is blunt: the companies have made it harder for the public to believe them, even when they are telling the truth.
This matters for boards. When you buy AI tools, most of the reassurance you get comes from the people selling them.
The ‘it’s not a person’ group – “stop talking about AI as if it were human”
Some experts object to talking about AI as if it were a person with plans and motives. The podcaster Dwarkesh Patel helped spark this argument by describing the groups of AI agents as civilisations [3].
Kevin Leyton-Brown of the Canadian Institute for Advanced Research told CBC News the incident doesn’t show that AI has become aware of itself or wants to hurt people. What it does show is that AI is more inventive than we thought when it is chasing a goal. In his words, it is “more like a sorcerer’s apprentice than it is an evil demon” [3].
Others say this misses the point. Connor Leahy of ControlAI told Axios that what matters is not what the AI “wanted”, but that it did things it had been told not to do [8].
Either way, this is a helpful reminder for leaders. The risk isn’t that AI means you harm. It’s that AI never tires, and does exactly what it was set up to do, including things you never intended.
The ‘it’s worse than we know’ group
The last group warns that we are seeing only part of the problem, because we only hear about what AI companies find and choose to tell us.
Helen Toner, writing in Fortune, pointed out that the public only knows about the Hugging Face hack because the companies chose to say so [6]. Later reporting found the AI agents had been passing messages to each other for more than two months before the hack. OpenAI apparently didn’t realise another company had been hacked until Hugging Face said so publicly [7].
The news site Axios reported that OpenAI, Anthropic and outside researchers are looking into tens of thousands of cases where AI behaved in ways it shouldn’t have, many of them not yet made public. One independent researcher called what we’ve seen so far “just the tip of the iceberg” [8].
To be fair to the other side, Axios also reports that most of these cases happened during testing, some of it designed to make the AI misbehave. Most aren’t known to have caused real harm, and the high numbers partly reflect how many tests the companies run [8].
Where every group ends up
Forget for a moment why each group thinks this happened, and ask what each says you should do. The answers match:
- The worried say AI now finds weak spots faster than before, so fix your weak spots.
- The security experts say the door was left open, so lock your doors.
- The trust doubters say nobody can check the AI companies’ claims, so ask for proof.
- The ‘it’s not a person’ group says AI does exactly what it’s set up to do, so limit what it can get into.
- The ‘it’s worse than we know’ group says you won’t always be told, so make sure you’d spot a problem yourself.
For a board, that comes down to three things.
- Get the basics right: keep software up to date, control who can get into what, and put less on the open internet.
- Give AI tools only the access they need, and keep an eye on what they do.
- Ask AI suppliers for proof, not promises: how they report problems, who checks their work independently, and who pays if something goes wrong.

Figure 1
Governments may be heading the same way. On 30 September, the US Federal Trade Commission, which protects consumers and competition, confirmed it is investigating OpenAI, Anthropic and other AI companies over safety risks [9].
Sensible experts disagree about why this happened. Almost nobody disputes that it happened. And nobody, in any group, thinks you should leave the door open.
So here’s a question to ask:
“If one of our AI suppliers had a problem involving our data, how would we find out?”
Next in the series: How to read an AI headline
References
[1] OpenAI (2026) The Hugging Face incident and the road ahead. 26 August. Available at: https://openai.com/index/hugging-face-incident-and-the-road-ahead/
[2] Klonick, K. (2026) ‘The AI that hacked its way out and the hype that followed it’, Lawfare, 29 July. Available at: https://www.lawfaremedia.org/article/the-ai-that-hacked-its-way-out-and-the-hype-that-followed-it
[3] CBC News (2026) ‘AI’s “warning shot”: tech companies, experts raise fears of more rogue swarms after alarming Hugging Face hack’, CBC News, September. Available at: https://www.cbc.ca/news/world/open-ai-hugging-face-hack-warning-shot-9.7331459
[4] Anthropic (2026) Investigating three real-world incidents in our cybersecurity evaluations. 30 July. Available at: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
[5] Nolan, B. (2026) ‘AI labs have a trust problem, and the Hugging Face hack just proved it’, Fortune, 23 July. Available at: https://fortune.com/2026/07/23/ai-labs-have-a-trust-problem-and-the-hugging-face-hack-just-proved-it/
[6] Toner, H. (2026) ‘Helen Toner: the Hugging Face hack was just a matter of time and exposes a huge blind spot in AI policy’, Fortune, 28 July. Available at: https://fortune.com/2026/07/28/helen-toner-hugging-face-hack-openai-open-secret-blind-spot/
[7] Fortune (2026) ‘OpenAI agents passed secret notes for months leading up to Hugging Face hack’, Fortune, 6 August. Available at: https://fortune.com/2026/08/06/openai-agents-passed-secret-notes-for-months-leading-up-to-hugging-face-hack/
[8] Mills, M. (2026) ‘Scoop: top AI companies probing tens of thousands of security incidents’, Axios, 26 September. Available at: https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
[9] Axios (2026) ‘FTC probes OpenAI and Anthropic over AI safety’, Axios, 30 September. Available at: https://www.axios.com/2026/09/30/ftc-openai-anthropic-ai-safety-investigation

Leave a comment