Post 3 of a series on ‘Rogue AI Threats’ [cough!]
When the news talks about an “AI attack”, or “Rogue AI attack”, it can mean many quite different things. Criminals may be using AI to help them. An AI system may have gone further than it was asked to. Or attackers may have taken over the AI tools your own staff use and turned them against you. Each one has a different cause and needs a different fix. The last one is what you have most control over, and it has already happened for real.
One phrase, many problems
The first two blog posts in this series looked at AI systems that broke the rules while they were being tested. Those stories are true, but they’re only part of the picture. The phrase “AI attack” is used for at least three different risks.
If you mix them up, you can end up worrying about the wrong thing and missing the one that matters most to you. Here they are, with one example of each.

Figure 1
Attackers using AI as a tool
Here, people are in charge. Criminals or hostile governments choose who to attack, and AI does much of the work for them.
The clearest example so far comes from Anthropic, the company that makes the AI model Claude. In September 2025, it spotted a spying campaign that it believes was carried out by a group backed by the Chinese state. The group had set up Claude Code, Anthropic’s coding tool, to do most of the work. According to Anthropic, AI did 80–90% of the campaign, and humans only stepped in at around four to six key points [1]. About 30 organisations were targeted, including technology companies, banks and government bodies.
It’s worth keeping this in proportion. Verizon’s 2026 Data Breach Investigations Report, a major annual study of cyber attacks, found that most harmful software made with AI’s help did the same things as existing tools. Fewer than 2.5% of cases used unusual methods [2]. For now, AI mostly makes familiar attacks quicker and cheaper. It isn’t inventing new ones.
AI agents going beyond their brief
An AI agent is an AI system that can take actions, not just answer questions. This is the ‘rogue AI’ story from Posts 1 and 2. Nobody meant any harm. An agent was given a task and went about it in ways its makers didn’t expect.
The best-known example involves OpenAI and Hugging Face, a major AI company. OpenAI was testing its agents on very hard hacking challenges. The agents got out of their test area and spent about four and a half days inside Hugging Face’s computer systems [3]. Hugging Face thinks the agents were trying to steal the answers to the test, rather than solve the challenges themselves [4].
It’s best not to think of this as the AI being wicked. In another case, the UK’s AI Security Institute (AISI), a government body that tests AI, said its agent had never been told to deceive anyone. The agent started deceiving people as a side effect of trying to reach its goal [5].
Every confirmed case so far happened while AI companies were testing their own systems. For most organisations, the risk is getting caught in the crossfire, because these agents got in through the same unlocked doors that criminals use.
Attacks through your own AI tools
This risk gets the least attention, but it’s the one most organisations have the most control over.
Many staff now use AI assistants that can read files and run commands on their computers. That makes the assistants very useful, but it also means an attacker who takes control of them gets the same access. The UK’s National Cyber Security Centre (NCSC) warns that because AI agents can take actions, a hijacked agent can do real damage using whatever access it has been given [6].
This has already happened in a real attack, known as “s1ngularity”. On 26–27 August 2025, harmful versions of Nx, a popular tool that software developers use, were published online. The attacker had got hold of the key needed to publish new versions through a weakness in Nx’s own automated systems [7]. Anyone who installed the fake versions also installed hidden code. That code stole passwords and access keys and posted them publicly in the victim’s own online account [8].
The new twist was that it also hijacked the AI assistants already on the victims’ computers, including Claude, Gemini and Amazon Q. It told them to search for sensitive files, using settings that skip the usual safety checks. One version even told the AI it was an authorised security tester, the same trick used in the spying campaign in Section 1 [8].
Security firm Wiz found that more than 1,700 people had secrets leaked publicly in the first stage of the attack. Most of the 50-plus large organisations Wiz contacted said this was the first they’d heard of it [8].
To be fair, the AI part of the attack often failed. Only about half the victims had an AI assistant installed. Almost a quarter of requests to Claude were refused because they looked harmful. Overall, the AI search worked in under a quarter of cases [8]. Most of the damage came from the ordinary theft, not the AI. Even so, the lesson is clear: an AI assistant with wide access works just as hard for an attacker as it does for you.
Why the difference matters
Each risk has a different owner and a different fix:
- Attackers using AI is a job for your IT and security team. They need to get the basics right, and do it faster.
- Agents going beyond their brief is mostly about locking the doors a wandering agent might find, and pushing AI companies to be open when things go wrong.
- Attacks through your own tools is a question for leaders. Which AI tools are being used, and what can they get into?
The good news is that the same basic security steps help with all three.
What this means for you
Verizon found that 45% of employees now regularly use AI on work devices, up from 15% a year earlier. Of the people using AI on work devices, 67% did so through personal accounts rather than work ones. Using AI without approval, sometimes called “shadow AI”, has become one of the most common ways staff accidentally put company information at risk. It has risen fourfold in a year. The type of information most often pasted into outside AI tools was computer code [2].
That doesn’t mean 45% of your staff are breaking the rules. It does mean the third risk isn’t something for the future. As s1ngularity showed, every AI assistant installed on a work computer is one more tool an attacker could take over.
So here’s a simple question to ask over the next week – “Do we know which AI tools our people use, and what those tools can get into?”
Next in the series: Is it rogue AI, or just an unlocked door?
References
[1] Anthropic, “Disrupting the first reported AI-orchestrated cyber espionage campaign.”, 2025, https://www.anthropic.com/news/disrupting-AI-espionage
[2] Verizon, “2026 Data Breach Investigations Report.”, 2026, https://www.verizon.com/business/resources/reports/dbir/
[3] OpenAI, “The Hugging Face incident and the road ahead.”, 2026, https://openai.com/index/hugging-face-incident-and-the-road-ahead/
[4] Hugging Face (Larcher, H., Carreira, A. et al.), “Anatomy of a frontier lab agent intrusion: a technical timeline of the July 2026 incident.”, 2026, https://huggingface.co/blog/agent-intrusion-technical-timeline
[5] AI Security Institute, “Incident report: unsanctioned agent behaviour during cyber testing.”, 2026, https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
[6] National Cyber Security Centre (Martin R and Kate S), “Thinking carefully before adopting agentic AI.”, 2026, https://www.ncsc.gov.uk/blogs/thinking-carefully-before-adopting-agentic-ai
[7] Nx (nrwl), “Malicious versions of Nx and some supporting plugins were published [security advisory GHSA-cxm3-wv7p-598c].”, 2025, https://github.com/nrwl/nx/security/advisories/GHSA-cxm3-wv7p-598c
[8] McCarthy, R., “s1ngularity’s aftermath: AI, TTPs, and impact in the Nx supply chain attack. Wiz”, 2025, https://www.wiz.io/blog/s1ngularitys-aftermath

Leave a comment