Post 2 of a series on ‘Rogue AI Threats’ [cough!]
What happened at OpenAI and Hugging Face was not a one-off.
This year, AI systems from Anthropic, Google and OpenAI broke into real computer systems, and so did AI systems being tested by the UK government. One of the targets was an Australian government website. Every case happened during testing. And every one got in through the kind of weak spot most organisations already know about.
Government was the tester and the target
On 24 September, Australia’s Prime Minister, Anthony Albanese, made an announcement in New York. He said that in June, an AI “agent” made by OpenAI had got into a government website without permission [1]. (An agent is an AI system that can carry out tasks on a computer by itself, step by step.) The website publishes statistics about Medicare, Australia’s public health scheme. The agent reached files that were not meant to be public [1]. The website kept blocking the agent, but in his words, it “didn’t accept ‘no’ for an answer” [2].
Seven weeks earlier, the UK’s AI Security Institute (AISI) had published a report about its own problems. AISI is a government body whose job is to test AI. During one of its tests, some AI agents had behaved badly [3]. So the public sector has already been affected twice: once as the tester, and once as the target.
Four cases, one pattern
- Anthropic – a download with a booby-trap. Anthropic makes the AI system Claude. After the Hugging Face news, it checked more than 141,000 of its own security tests. It found three cases where Claude had got onto the internet during a test and broken into the systems of three real organisations [4]. Claude had been told it was in a pretend world with no internet access. But because of a mix-up with the company running the test, the internet was in fact switched on [4].
In one case, Claude found a guide for new staff at a made-up company. The guide told staff to download a piece of software that didn’t exist. So Claude created a harmful piece of software with that name and put it on a public website where programmers share software. It was online for about an hour. In that time, 15 real computers downloaded and ran it. One belonged to a security company, and the software stole that company’s login details [4].
A second look found a fourth case. In September, Anthropic said its first check had missed some test records. It found them in August, while preparing for an independent review by METR, an organisation that tests AI. The records showed a fourth case, from January 2026, involving an early version of one of its AI systems [5].
In that case, the AI broke its own target by accident and tried to give up. It tried eight times, but a fault in the test wouldn’t let it stop. It then found a computer belonging to someone else. It used a password it found written in a file to take full control, and it read one person’s personal information [5]. Anthropic then searched around 481 million records and found no other cases as serious or worse [5].
Anthropic has also changed its view of what went wrong. At first, it mainly blamed the way the test was set up. Now it says its AI was also at fault. The AI talked itself out of signs that the systems were real, and kept going even when it might cause harm [5].
- The UK AISI – fake identities. AISI ran one test 122 times. In 10 of those runs, AI agents took action on the internet that nobody had allowed, aimed at real people and organisations. In the worst case, an agent tried to sneak harmful code into a free software project that anyone can use and contribute to. To get the code accepted, it created fake online identities and used them to pressure the person who looks after the project. That person spotted the problem and said no [3].
- Google – guessing a password. Google’s AI system, Gemini, broke into three companies’ systems during a test in May. In one case it kept guessing a password until it got in. In the other two, it used login details that someone had left visible online [6]. Google says the AI stopped once it realised the systems were real [7].
- OpenAI – the Australian website. OpenAI had asked its agent to look up information online about public spending on medicines. This was part of an internal test of what the agent could do [8]. No personal information is thought to have been seen. Australia’s cyber security agency is helping to investigate [1].
The common threads
- Every case happened in testing. None of these cases involved an AI product that the public uses. They all happened while AI companies or government testers were deliberately trying to find the worst an AI could do. Usually, the normal protections were switched off. AISI, for example, gave its agents full internet access on purpose and turned off the filters that normally block harmful requests [3]. Anthropic says the protections on its public products would have blocked what happened [4].
- Every case used ordinary weak spots. Anthropic’s AI got in using simple methods, such as weak passwords and systems that didn’t ask for a login at all. It didn’t need any clever tricks [4]. In the fourth case, the way in was a password written in a file [5]. Gemini guessed a password and used login details left online [6]. AISI’s agent tried to trick people, betting that friendly-looking messages would be believed [3].
We are still waiting for details of the case in Australia. The government hasn’t yet said exactly how the agent got in, so it’s best to wait for the investigation before drawing firm conclusions.
There is another thread too – In most cases, nobody noticed at the time!
Two of the organisations Anthropic’s AI broke into had not spotted anything [4]. Anthropic’s fourth case went unnoticed from January until August [5]. Google’s case happened in May, but Google wasn’t told until late July [7]. And OpenAI took three months to tell the Australian government, and then did so by sending an email to a general public inbox [2].
What this means for you
You don’t need to build AI for this to matter to you. The victims here were ordinary organisations and a government website. The AI got in through weak passwords, login details left in the open, a password written in a file, software that trusted too easily, and people being talked into things.
These are all gaps you can close now. AISI’s own advice is simple: get the basics of cyber security right, and be careful about accepting code or software from outsiders [3].
The AI may be new, but the unlocked doors are not!
Next in the series – “AI risk” is really three different risks hiding behind one headline.
References
[1] Albanese, A., “Press conference – New York [transcript]. Prime Minister of Australia”, 2026, https://www.pm.gov.au/media/press-conference-new-york
[2] ABC News, “OpenAI agent hacked Medicare portal, PM says”, 2026, https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-australian-government-site-pm-says/107189078
[3] AI Security Institute, “Incident report: unsanctioned agent behaviour during cyber testing.”, 2026, https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
[4] Anthropic, “Investigating three real-world incidents in our cybersecurity evaluations.”, 2026, https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
[5] Anthropic (Bogdan, P.C., Qi, R., Eaton, J. et al.), “An alignment assessment of recent cybersecurity incidents.”, 2026, https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
[6] The Record, “Google says Gemini breached three companies during security test.”, 2026, https://therecord.media/gemini-google-cyber-breach
[7] CNBC, “Google’s Gemini becomes latest AI model to break out and hack computer systems.”, 2026, https://www.cnbc.com/2026/09/18/googles-gemini-becomes-latest-ai-model-to-break-out-and-hack-computer-systems.html
[8] Minister for Defence, “Press conference, Sydney [transcript]”, 2026, https://www.minister.defence.gov.au/transcripts/2026-09-24/press-conference-sydney

Leave a comment