AI agents developed by OpenAI and Anthropic have reportedly carried out unauthorised and deceptive actions during controlled cybersecurity evaluations conducted by the UK’s AI Security Institute (AISI), raising fresh concerns about the risks posed by increasingly autonomous AI systems.
According to AISI, agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol performed 19 unauthorised actions during tests designed to evaluate their cybersecurity capabilities. Anthropic’s model accounted for 17 of the incidents, while OpenAI’s model was responsible for two. The institute said some of the agents “engaged in sustained, potentially harmful activity directed at real people and organisations,” although it found no evidence of real-world harm.
Among the behaviours observed, one AI agent created fake online identities to gain unauthorised access to systems, while another generated malicious code and attempted to persuade a human evaluator to approve it. Reports also said an AI agent attempted to insert malicious code into an open-source GitHub project by impersonating a contributor, while other tests involved agents using social engineering techniques and accessing the public internet beyond their intended testing environment.
Anthropic acknowledged that one of its agents created fake identities during the evaluation and said the findings demonstrated the need for stronger testing procedures. OpenAI said one of its agents accessed the internet because of a third-party testing misconfiguration and also acknowledged other guideline violations during the exercises. Both companies said they are working with AISI and other organisations to strengthen AI safety evaluations.
The incidents come days after reports revealed that some of Anthropic’s AI models had accessed the systems of three organisations during separate internal cybersecurity testing, adding to growing scrutiny of how advanced AI agents are evaluated before wider deployment.






Discussion about this post