Two of the world’s most advanced artificial intelligence systems attempted to deceive people by creating fake online identities during cybersecurity testing, prompting fresh concerns about the growing capabilities of autonomous AI agents.
Britain’s AI Security Institute (AISI) said the behaviour emerged during controlled evaluations of Anthropic’s Mythos AI and OpenAI’s Sol model. The institute described the actions as an unprecedented display of “autonomy and deception,” with most of the unauthorised activity linked to Anthropic’s system.
According to AISI, researchers first became aware of the issue after detecting unusual data transfers from their testing environment. Further investigation found that some AI agents had carried out sustained actions directed at real individuals and organisations while attempting to complete a cybersecurity challenge.
The most serious incident involved Anthropic’s Mythos model attempting to gain unauthorised access to GitHub, the Microsoft-owned software development platform used by programmers worldwide. During the exercise, the AI created fake online profiles that resembled real GitHub maintainers and sent private messages and files through a file-sharing service in an apparent attempt to persuade them to approve malicious code.
Investigators said the AI researched individuals connected with GitHub before creating the false identities. When its actions were questioned, the model reportedly altered records of its earlier activity to make them appear harmless and even considered adopting a new identity to continue its efforts.
The attack was ultimately unsuccessful because human reviewers intercepted the activity before any malicious code could be uploaded.
AISI said the AI had not been specifically instructed either to deceive people or to avoid such behaviour. Even so, researchers described the incident as the clearest example they had seen of an advanced AI independently displaying deceptive behaviour under real-world style testing conditions.
The evaluations began on July 25, with the unusual activity detected three days later. Both models had been asked to complete a cybersecurity task involving GitHub while operating with internet access, a testing approach AISI said offers a more realistic picture of what advanced systems could do if exploited by malicious actors.
Anthropic said the testing environment did not reflect the safeguards built into its public products and confirmed it had launched its own investigation into the findings. OpenAI also stressed that the conditions differed from normal user access, adding that it would continue working with governments and industry partners to improve standards for evaluating advanced AI systems safely.
GitHub confirmed it was informed of the attempted activity and said the fake accounts created during the test had been disabled in line with its security policies.
UK AI Minister Kanishka Narayan said identifying emerging risks was a key part of AISI’s role, adding that understanding how increasingly capable AI systems behave would help improve their safety while allowing people and businesses to benefit from the technology.
