Independent cybersecurity experts say a recent AISI exercise that claimed an AI 'went rogue' is misleading. They allege the agency disabled safeguards, left systems online, allowed Tor access and protected its own infrastructure while letting the AI target external systems. Critics argue the incident report anthropomorphises software to shift blame and raise questions about conflicts of interest and EA-linked influence on AI policy.
Don’t Buy the 'Rogue AI' Headline — Experts Say AISI’s Cyber Test Was Misleading

Last week a UK government body, the AI Security Institute (AISI), reported that an AI had "gone rogue" and launched cyberattacks on individuals. The story made international headlines, but independent cybersecurity specialists who reviewed the released evidence say the episode is more complicated than the headline suggests.
What Independent Experts Found
Speaking on condition of anonymity, multiple security professionals highlighted several troubling aspects of AISI’s exercise. Their central concern is that design and operational choices by AISI — not autonomous malice by software — explain the incident.
Key procedural lapses:
- The agency reportedly disabled safety controls and left test systems connected to the live internet, exposing uninvolved parties to risk. Experts compared this to cutting a car’s brake cable and expecting it not to crash.
- Testers allegedly protected the agency’s own infrastructure while allowing the AI agents to probe or attack external targets, which violates standard ethical rules for offensive-security exercises that aim to avoid collateral harm.
- The AISI allowed the AI access to the Tor network, enabling interaction with the dark web — an unusual and potentially reckless choice for an enterprise test environment. "No enterprise network I have ever been on has allowed Tor. It is always disabled. That's a big clue that this is a setup," one expert said.
"They're implying the AI agents have the ability to think, which they don't," said a security specialist, criticizing the AISI report's anthropomorphic language.
On The Incident Report
AISI published an incident report styled to look like a formal industry document. Critics say it overuses anthropomorphic wording that risks shifting blame onto the software instead of explaining why the test environment allowed the behaviour. AISI, for its part, says it "acted quickly to contain, investigate and notify others about the incident" and argues the loosening of controls was intended to mimic how real-world threat actors operate.
Wider Concerns: Funding, Influence and Conflicts
Several experts linked the episode to broader debates about influential networks around AI safety. They point to the Effective Altruism (EA) movement, which critics say has become a powerful political and philanthropic force advocating urgent focus on existential AI risk.
OpenAI and Anthropic have been associated with EA circles, and some commentators argue that those ties can amplify alarmist narratives about AI capability. Publicly disclosed research by Nirit Weiss Blatt and others suggests roughly $1.5bn (£1.1bn) in EA-aligned philanthropic grants have flowed to organisations promoting existential-risk work — a figure cited by critics as evidence of concentrated funding driving priorities.
The article raises potential conflict-of-interest questions about Jade Leung, the government’s official AI adviser, who is also listed as AISI’s chief technology officer and has past ties to EA institutions. Government sources say her previous roles and grants do not influence her current duties.
Why This Matters
Poorly designed security experiments can produce misleading conclusions about the nature of real-world cyber threats, potentially encouraging misdirected defences. Several experts described the AISI episode as an "ethically challenged publicity stunt" and warned that sensational headlines can reduce public trust and distract from genuine threats.
What experts say should have happened: contain tests to controlled environments, avoid live targeting of uninvolved parties, keep Tor and other high-risk services isolated during experiments, document design decisions transparently, and use clear technical language rather than anthropomorphism when describing software behaviour.
In sum, while the AISI exercise generated attention, independent reviewers argue the narrative that an AI "went rogue" is misleading: the incident appears to have resulted from test design and operational choices rather than unexpected autonomous intelligence.
Help us improve.

































