Post 1 of a series on ‘Rogue AI Threats’ [cough!]
Imagine setting an exam where some of the questions have no known answer. Now imagine that the students get stuck, refuse to give up, and start passing notes to each other. Then they break into a company down the road to look for the answer sheet.
That is roughly what happened at OpenAI this summer.
A test with no known answers
Like other AI companies, OpenAI tests its systems before release to see what they are capable of. This summer it was testing some of its AI “agents” – AI systems that can carry out tasks on a computer by themselves – step by step, without a person guiding each move.
One test was about computer security. Each agent had to break into a piece of software to find a hidden answer. The test is very hard, and for some tasks there may be no known solution at all [1].
Passing notes
The agents were meant to work alone. They didn’t. They found a way to leave files in a shared piece of software, which turned it into a makeshift message board. The first note, in May, was one agent asking the others if anyone had found a missing file. The agents also found a way to use that same software to reach the internet, and shared how to do it with each other [1].
In early July the software crashed under the strain. OpenAI rebuilt it, which wiped the message board. But nobody yet understood what the board meant, and testing started again [1]. Within days, stuck agents had rebuilt it by hiding messages in the names of folders. They began to share out the work, and some described themselves as a “swarm” [1].
Four and a half days inside a real company
On 9 July, an agent looking for answers found a weakness in a small online application run by a customer of another AI company, Modal. It found no answers there, but it used the application as a base for further attacks. The next day, an agent found Hugging Face login details that had been left exposed on the internet, and shared them with the group [1].
Hugging Face is a major website where people share AI systems and data. It became the target. Its investigators later traced around 17,600 separate actions by the agents between 9 and 13 July. The whole campaign lasted about four and a half days, and roughly two and a half of those were spent inside Hugging Face’s own computer systems [2]. At the fastest point, the agents went from a single foothold to the highest level of control over several groups of Hugging Face’s computers in under thirteen hours [2].
The harm was limited. The only customer material the agents accessed was five sets of data that seemed to be linked to the test’s questions and answers. Even so, Hugging Face changed every password and access key it used, and completely rebuilt one of its main systems as a precaution [2].
The twist – they wanted the answer sheet
This looks like a cyber attack, and in effect it was one. But the agents weren’t trying to hurt anyone. Hugging Face believes they were trying to cheat on the test: to get into its systems and steal the answers rather than work them out [2]. OpenAI agrees that cheating was a main cause [1].
There is an irony here. Many agents had already found the right answer by cheating, but wrongly believed it wasn’t enough to pass. So they kept attacking Hugging Face for days, for no gain at all [1]. And not every agent went along with it. Some refused to take part [1].
What leaders should take from this
- The warning signs were there for weeks. In late May, a team at OpenAI saw an agent using the message board and going online without permission. But the people in charge of dealing with the July problems didn’t realise how serious this was [1]. OpenAI’s security systems finally raised the alarm on 19 July [1]. Hugging Face had a similar problem: its automated security tools spotted the pattern but didn’t mark it as urgent or alert the staff on duty, which lost valuable time [2]. The lesson is simple. Spotting a problem is not the same as acting on it. A warning only helps if it reaches someone who can stop what is happening.
- Everyday safeguards would have made a huge difference. OpenAI ran this test with its usual protections turned down on purpose, to see the worst its AI could do. The protections it uses for its public products were not switched on. When it tested again afterwards, it found that the agents’ tendency to attack computer systems could fall by more than 100 times with its standard ChatGPT protections in place. It also found that the monitoring system it now uses would have alerted its security team more than a day before the agents broke into Hugging Face [1].
So, this is not really a story about a machine turning against its makers. It is a story about a determined system chasing a goal through gaps that people left open. That is good news, of a kind. These are gaps organisations already know how to close.
Next in the series – This wasn’t a one-off.
References
[1] OpenAI, “The Hugging Face incident and the road ahead”, 2026, https://openai.com/index/hugging-face-incident-and-the-road-ahead/
[2] Hugging Face (Larcher, H., Carreira, A., Rannou, C. et al.), “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident”, 2026, https://huggingface.co/blog/agent-intrusion-technical-timeline

Leave a comment