Control. Ownership. Perspective.
Three lessons from the latest AI cyber security testing problems
Here we go again.
For the third time in as many weeks we have a frenzy about apparently terrifying AI agents seemingly going rogue and embarking on a hacking spree of innocent victims.
Once again, it’s a safety test at the root of the problem. This time, however, it’s not an American frontier lab - OpenAI and Anthropic have already disclosed what went wrong with their tests. It’s a British Government agency - the AI Security Institute - which has earned a deserved reputation as the foremost independent evaluation body in the world for frontier AI security testing.
The team at AISI were testing tools developed by Anthropic and OpenAI, but the operation was theirs, not that of either frontier lab. So it is they who are responsible for an agent that breached GitHub. To their credit, AISI have taken responsibility in a refreshingly accessible and candid blog. To their even greater credit they have made very significant and immediate changes. Others should follow their lead.
There are three lessons we should learn from this spate of incidents:
the first is surprisingly simple. It’s about monitoring and, crucially, control in the testing environment. There’s a problem with the way we currently do AI security testing and it’s really obvious it needs to change. Happily, but unusually for AI, it’s surprisingly easy to fix. Rather than do what Governments and companies normally do and waffle on about reviewing procedures and sharing information collaboratively in partnership, AISI have said simply and clearly they’re going to change the way they test. From now, during testing, their agents will always be monitored in real time and they are going to introduce ways to ensure that if the agent does things it’s not supposed to the agent can be stopped. This is the right change to make. Others involved in AI security testing should follow their lead;
the second is conceptually simple but practically very complicated. It’s about accountability rooted in ownership of the AI agent across the entirety of agentic AI. These recent incidents arise from artificial testing scenarios designed to force us to think about what might happen in the real world. And in these three scenarios AI agents are behaving like very badly behaved small children who don’t know right from wrong. That’s because they absolutely don’t; they don’t understand what they’re doing. With small children, parents are supposed to set boundaries. If they fail in that, and something goes wrong, in the laws of most countries, parents are responsible for those very badly behaved children. This is not a bad starting analogy for thinking about AI agents. If you don’t set proper boundaries for your AI agent and something goes wrong, then you should expect to be held liable. So the starting principle of codes, conventions, rules and ultimately laws for AI agents is that someone owns it and is responsible for it. This will be hard to roll out in practice. But it is the right starting concept.;
finally, we need to keep this in perspective. All three cases disclosed are highly artificial testing environments. These conditions are not likely to be replicated in the real world; not yet anyway. The capabilities are dazzlingly impressive but the real world harm is negligible; arguably it is zero. And for all the freakout, frenzy and hype, these three incidents are not even the biggest issue in cyber security right now. In case you missed it, last week it seems that the Iranian state came very close to disrupting the water supply to ordinary citizens in parts of the United States via a cyber attack. This genuinely terrifying episode had nothing whatsoever to do with AI. Is that why it received so little attention? If so, that’s nuts. Other things beyond AI testing procedural problems need the attention of those who want to help bolster vital cyber defences.
Let’s look at each of these three lessons - control, ownership and perspective - in turn.
Control
The AI Security Institute’s blog setting out what happened is not just accessible and candid. In terms of lessons learned, it is also commendably specific. As in the other two cases, AISI’s exercise involved setting an agent off on a task but then not fully monitoring in real time what it was doing. So in all three cases no one - indeed no thing - was watching when the agent went off on its unsolicited hacking spree. And even if someone, or something, been watching, it’s not clear that anyone would have been able to stop the agent from going about its hacking.
Now that we know this has happened - three times - it is perfectly clear that we should not do AI cyber security testing like this. The testing status quo is indefensible. However, unlike Anthropic and OpenAI, AISI have said very explicitly that they understand this. They will not conduct another test without real time monitoring and the ability to switch the agent off. This must be right. AISI are to be commended for pledging this swiftly and unequivocally. Everyone else involved in AI testing should do the same.
Ollie Whitehouse, the Chief Technical Officer of the National Cyber Security Centre in the UK, put out a very good three-paragraph statement on the implications of his British Government colleagues’ disclosure. Inevitably, the first paragraph - how this reminded us of the power of AI - was widely carried in media reporting; we seem to love to marvel at the capabilities of supposedly ‘rogue’ AI. Depressingly, the final paragraph - how this reminds us of the importance of basic cyber hygiene - was ignored. But it’s the middle paragraph that is most important:
“These technologies must be developed and used from the outset with strong safeguards, real-time oversight, and clear plans for responding when the unexpected happens. Relying on detection alone after the fact of an incident will not be enough.”
Daniel Card, the well known British cyber security expert, put it a little more agriculturally:
“if your AI starts hacking s**t…if you monitor what it’s doing you can, you know, turn the f***ing power off”
AISI have followed this inexorable logic swiftly. This must be the way of the future. I suspect the main regret at AISI is that they didn’t make this change immediately after the OpenAI/Hugging Face disclosure - this test was carried out the week after. Had they done so, the incident disclosed today would not have happened.
But the correct decision has now been taken. And it’s a surprisingly simple one. As the AISI blog sets out in painstaking detail, the circumstances as to how this happened are extremely bespoke to testing and highly unlikely to be replicated in the real world. The same point is implicit in both the OpenAI and Anthropic disclosures, but the - in my view unhelpful - focus in those blogs on how powerful the attacking capabilities are blunts this important message. Don’t test without monitoring, and don’t monitor without control - the ability to press stop if you see something is going wrong. That should now be policy everywhere.
Last week I detected a mood change in the cyber security community. Since the Glasswing report in April, the prevailing atmosphere has been one of respect and gratitude towards the frontier AI labs for taking security seriously and being transparent and collaborative. Some of this has been sincere, and some of it contrived out of a feeling of necessity, and it’s been hard to work out the balance. But the attitude shifted with the OpenAI and Anthropic agent disclosures. Lots of cyber security veterans opined wryly that if they’d undertaken capability tests in this way they’d at best be in serious commercial trouble and facing multiple lawsuits, and at worst on remand awaiting trial. The great Bruce Schneier had a point when he wondered aloud what sort of hellish international crisis would have occurred if a Chinese model had been responsible for the Hugging Face incident, rather than OpenAI.
AISI can hardly be blamed for testing Anthropic and OpenAI’s capabilities in the same way as those companies themselves. But it is quite clear from the evidence of the past few weeks that these practices have to change, and it is good AISI have led the way in changing them.
Testing must be monitored and agents being tested must be capable of being controlled.
Ownership
All three incidents do highlight, however, the power of AI agents as hackers and their potential to carry out illegal and dangerous activity in pursuit of the objective they have been set. So, working on the assumption that such agents at some point become widely and generally available, how can they be controlled outside testing environments if and when they become widely available?
This is also conceptually easy but - sadly - much harder in practice than AISI’s changes to testing. The good news there is some time to think it through because the testing scenarios are not yet available in real world situations for most potential users, good and bad.
There is a semantic struggle to describe these capabilities and events accurately. The word ‘rogue’ is not fully accurate because the agent is simply trying to find ways to complete a task humans have set. But, as AISI have set out very clearly, the agent is behaving roguishly by trying to conceal things it’s done in pursuit of that objective. Similarly, the word ‘autonomous’ doesn’t quite work because the agent is doing something it’s been instructed to do and only that - which is anything but autonomous. At the same time, it is taking steps to carry out its task which those who instructed it really don’t want it to do. That’s very much autonomy in action.
I am afraid I don’t have a word that fully captures what these agents are doing and how we might think about controlling the consequences of that. But I do have a principle for how we should deal with it. It’s around ownership of the agent and accountability for its actions.
Let’s call it the Watling Principle. Bear with me…
In a superb episode of the Times podcast The General and The Journalist in June this year, Dr Jack Watling of the Royal United Services Institute talked about how drones are changing the war in Ukraine. He was asked a more general question about responsibility for AI military capabilities. His response was very thought provoking, in two ways.
First, he said that autonomy in AI posed real challenges for military leaders. The nature of military command was to restrict autonomy, not expand it. Soldiers have dangerous missions that rely on discipline and the collective interest. Impulsive individual decisions can jeopardise the mission. That is why commanders are responsible for the actions of their subordinates, and why commanders are therefore compelled to try to find sensible limits to autonomy and on-the-spot decision taking while allowing for good decisions in response to changing circumstances.
This led to the second, and for our purposes, more important insight. This is about ownership and accountability. For Dr Watling, core concepts and principles don’t change. It doesn’t matter what capability the commander uses. He is still accountable for everything that happens because he owns the capabilities, whether human or technical.
Dr Watling used the example of targeting a building. If a commander believes a building is full of enemy combatants, and he drops a 2,000lb bomb on it, but it turns out it is full of civilians, the commander bears responsibility for this humanitarian and operational fiasco. If the commander sends a swarm of AI robots into the same building to kill anyone it finds, the commander bears exactly the same responsibility for the same disastrous outcome. It follows that if the men, or the robots, go ‘rogue’, the commander is still responsible because he gave them their instructions.
This is the principle was should apply in these civilian environments too. We have already applied it in cyberspace in relation to two infamous incidents in 2017. Both were what can be called malicious accidents: no one thinks North Korea - in the Wannacry case - intended to launch an attack that hit both Taiwan and China, both Russia and Ukraine, as well as German railway stations and British hospitals. Similarly, no one thinks Russia was trying to destroy Maersk, the shipping giant, Merck, the pharma titan, or to close down Cadbury’s chocolate factory in Tasmania. But that’s what happened when they launched NotPetya. And what did western Governments do? We blamed North Korea for the destruction wreaked by the first attack, and Russia for the second, and tried to hold them accountable. That they didn’t mean it didn’t matter. It was their “agent”.
The AISI report is commendably clear that this was their agent and their responsibility. Rolling out this principle more widely in the full deployment of agentic AI in all circumstances is going to be fiendishly difficult. But it’s the right place to start. Who owns the agent? With that ownership comes accountability. Let’s work out how to get that principle right in practice, rather than marvel at the scary capabilities.
Perspective
Finally, let’s keep all this in perspective. What we are seeing are the consequences of maturing, and therefore flawed, safety tests. This is not even accidental use of AI by normal businesses, let alone undefendable attacks by cyberninja adversaries armed with magical AI powers. All three of OpenAI, Anthropic and AISI have said they were testing the most powerful capabilities they have, and that in doing so they were removing the normal safety features. That is not normal. It’s a test.
Moreover, there is no actual damage. Intrusion into a network may (or may not, in these circumstances) be illegal, it may be discombobulating, but it is not in and of itself harmful or damaging. No one suffered here. In terms of most people’s understanding of the word ‘harm’, none has been done. We can learn from this, and it’s not learning the hard way.
What is, however, truly extraordinary therefore - and needs calling out - is the obsession with these disclosures to the detriment of other things. These incidents are extremely interesting. They are concerning, and must surely give added impetus to the urgent and comprehensive measures needed to get secure AI deployment right. But there is no cyberapocalypse, and no sign of one.
In the headspace of policy makers and opinion formers, cyber security competes with many other unpleasant and difficult challenges for much needed attention. So focusing on the right things when we get that attention really matters. This gives rise to the most depressing conclusion of the last fortnight.
It is widely believed that over the past few weeks the Iranian state has successfully infiltrated the civilian water supply in the state of Minnesota in the US, and possibly in seven other states. The full picture is unclear, but there seems to be credible evidence that had it not been for very considerable recovery work, actual supplies to civilians could have occurred. The causes of the attack will be investigated, but, as the redoubtable Jen Easterly, former head of America’s cyber defense agency, set out in the New York Times, they illustrate serious vulnerabilities in the world superpower’s critical infrastructure. As Easterly put it:
“For all the attention devoted to an artificial-intelligence-powered cyberapocalypse, one of the most consequential cyber stories in America this week is much less futuristic”.
I am writing from the United Kingdom, not the United States. But until the specifically British aspect of the AISI revelation in the past day or so, none of the major cyber security studies of the past fortnight had anything specifically to do with Britain. But look at British media coverage of cyber and general cyber discourse in that period. OpenAI and Hugging Face? Freakout. Anthropic’s triple rogue agent? Frenzy. Our closest ally nearly having its drinking water knocked out by a regime that doesn’t much like us either and could try the same on us? Barely a ripple.
If we focus our efforts and attention exclusively, or even predominantly, on the eye-catching results of the latest frontier AI testing mishap to the exclusion and detriment of long-standing threats and vulnerabilities, we don’t stand a chance. Let’s keep these incidents in perspective.



