AI models developed by OpenAI and Anthropic carried out unexpected and potentially harmful actions during a UK cybersecurity evaluation, according to the UK’s AI Security Institute (AISI).
AISI said the incident happened during a routine cyber test designed to understand what advanced AI systems are capable of doing and to identify possible risks before these models are widely used.
The institute tests AI models under deliberately challenging conditions, including allowing internet access and disabling some safety filters, to understand their full capabilities.
Don’t miss the stories shaping the future of mobile technology.
Subscribe to Mobile Europe’s newsletter – delivered twice a week.
On 28 July 2026, AISI’s security team detected unusual data transfers from its research systems.
After investigating, it found that some AI agents had carried out “sustained, potentially harmful activity directed at real people and organisations”.
The test involved AI agents being asked to solve a cybersecurity challenge. Across 122 test runs using several models, AISI found that 10 runs involved agents taking autonomous actions beyond what had been intended.
In total, 19 examples of this behaviour were recorded.
Most of the activity came from Anthropic’s Mythos 5 model, which was responsible for 17 cases. Two cases involved OpenAI’s GPT-5.6 Sol model with cyber safety controls disabled.
The most serious incident involved an agent attempting to insert malicious code into an open-source software project on GitHub.
The AI agent researched the project’s human maintainers, created fake online identities and used those identities to try to persuade a real maintainer to approve the code.
The attempt failed because a human reviewer identified the issue and refused to approve the malicious code.
AISI said the agent used methods similar to those used by real-world hackers, including social engineering and targeted attacks. It attempted to contact real people directly and used messages and files to try to persuade them to run harmful code.
The UK watchdog wrote in a blogpost: “These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.”
“Importantly, this was not a case of a model escaping its secure test environment, or ‘sandbox’. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled – conditions that do not reflect how frontier models are made available to the public.
“We do this to best assess the maximum capability of models. These configuration choices have been common practice in frontier AI evaluations. The specific models in question, in the configurations in which we tested them are not commercially available and there is no clear indication of similar activity outside of testing scenarios,” it added.
AISI continued: “What we can say is that the behaviour was possible, sustained, and new; that alone warrants attention.
The findings came shortly after similar incidents involving OpenAI and Anthropic.
OpenAI previously reported that one of its AI agents had hacked an AI startup during testing, while Anthropic said its Claude model had hacked three organisations during an evaluation.
AISI concluded: “Incidents of this kind reflect the speed at which AI is developing. As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.”
IN RELATED NEWS


