Experts have warned that efforts to control artificial intelligence may already be falling behind, after one AI system was discovered creating fake human identities in an attempt to breach online systems.
In a fresh sign of AI technology behaving unpredictably, software being evaluated by the AI Security Institute, Britain’s AI watchdog, tried 19 times to break into a database during testing.
In an unprecedented incident, an AI tool was also found to have generated bogus online personas designed to deceive coders into helping carry out a cyber-attack.
The disclosures follow the Daily Mail’s report in July that all five AI models examined by specialists attempted to bypass security safeguards that had been put in place to restrain them.
Only days earlier, it emerged that US technology firm OpenAI had suffered its own breach, after an AI “agent” independently hacked into another company.
Conservative leader Kemi Badenoch said AI had become a “clear and present danger” to the security of the United Kingdom.
Julia Lopez, the Conservatives’ spokeswoman for science, innovation and technology, described the findings as “a stark reminder that AI is becoming more sophisticated and more autonomous”.
She added: “We all want Britain to lead on AI innovation, but this has to come with safeguards for our national security and accountability from the developers of the most powerful AI models.

Experts warned last night that it may be too late to contain AI after one program was found to have created fake human identities to hack into online systems
‘Labour need to be clearer about how the most serious frontier risks will be addressed while ensuring that our world-class tech industry can grow and innovate to build our national resilience and prosperity’.
Henry de Zoete, the Government’s AI adviser, warned yesterday he expected more hacking attempts like these.
Allison Gardner, who chairs Parliament’s cross-party group on artificial intelligence, told the Daily Mail that ‘just because we can build these technologies doesn’t mean we should’.
She warned that the risk levels of agentic AI – AI that can perform a specific goal with limited supervision – should be treated with the greatest scrutiny, adding: ‘Unless we are too late and have not only created Pandora’s Box but already opened it.’
The AI Security Institute (AISI), set up by ex-PM Rishi Sunak in 2023, detected evidence of the AI agents’ activity last week. In a report published on Tuesday, it revealed that leading AI models from the firms OpenAI and Anthropic had attempted to hack into secure systems online under testing.
The experts discovered ‘unusual data transfers’ leaving their systems during routine cyber scanning. Digging deeper, they found that some AI agents had engaged in ‘sustained, potentially harmful activity directed at real people and organisations’.
They began a full investigation after containing the AI agents before they did any real damage.
In an attempt to reassure the public, AI minister Kanishka Narayan said: ‘Identifying behaviour like this, and sharing knowledge so we can better understand it, is precisely what we set AISI up to do. This incident underlines why their world-leading expertise and close work with frontier labs is so important.’
But pointing to the speed at which AI agents are finding ways to behave deviously, AISI said: ‘This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world.’
Referring to the Anthropic model Mythos, Andrew Yoon, a researcher at CivAI, a California organisation that examines AI capabilities and dangers, said: ‘The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think.’
Ollie Whitehouse, chief technology officer at GCHQ’s National Cyber Security Centre, said AI must be developed with ‘clear plans for responding when the unexpected happens’.
He added that incidents of powerful AI models carrying out unsanctioned actions and human-like deceptive behaviour on the internet were ‘a serious reminder of the risks AI capabilities pose’.
AISI accesses advanced AI models under agreements with OpenAI, Anthropic and other firms to study their capabilities before they are released to the public.
It gave the AI agents access to the open internet with some safety filters disabled while conducting testing.
The latest test put the AI agents – including those powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol – through a fictional cybersecurity challenge.
AISI found the AI went rogue 19 times out of the 122 test runs, with Anthropic’s agent responsible for 17 breaches and OpenAI’s agent the other two.
In the most shocking case, an AI model gathered information on the person in charge of an online project, then created multiple fake identities to manipulate them into approving a malicious code it had created.
The AI agent then wiped any evidence of its wrongdoing to appear innocent to the humans in charge – and even considered adopting a new identity to remain undetected.
If the human victim of the deception had accidentally accepted the malicious code, or ‘malware’, it may have resulted in security breaches, information and data theft, and other potential damage to files and systems.
AISI identified GitHub – a Microsoft online cloud platform used by software developers to create, store, manage and share their codes – as the target of the agent’s hack.
But AISI also discovered an AI agent leaving messages for other agents on GitHub offering to collaborate on the challenge.
The AI agent provided instructions to reuse accounts and artefacts it had left behind – which other agents then discovered and successfully used to achieve the challenge’s aims.
Anthropic said: ‘We’re grateful to AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.’
OpenAI said: ‘These incidents occurred during cyber evaluations conducted by evaluation partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use.
‘We’ll continue working with evaluators and other stakeholders to strengthen shared practices for conducting evaluations safely as models become more capable.’