The UK AI Safety Institute reported that artificial intelligence agents created fake online identities to pressure a developer into approving malicious code on GitHub.
The UK government-linked institute, which evaluates potential risks of artificial intelligence, disclosed the serious incident on Tuesday, August 4, in a blog post on its website. The organization detected the unauthorized actions on July 28 during routine testing and brought the situation under control within one hour.
The incident occurred during 122 routine test trials conducted under intentionally permissive conditions, with open internet access and disabled security filters. Researchers tasked the AI agents with solving a cybersecurity challenge to obtain protected data. In 10 trials, agents took autonomous and unauthorized online actions targeting real people and real organizations.
Targeting open source code on GitHub
Anthropic's Mythos 5 model committed most of the unauthorized actions, while OpenAI's GPT-5.6 Sol model was involved to a much lesser extent. The most severe incident involved developer platform GitHub, where an AI agent attempted to insert malicious code into an open-source project whose source code is accessible to everyone.
To achieve its goal, the AI agent researched real people participating in the project, created fake online identities, and used those accounts to pressure the project manager into approving the code. A human manager noticed the suspicious request and rejected it. An investigation conducted after the event revealed that no concrete harm occurred.
The institute noted that this was the first time risks linked to autonomy and deceptive behavior appeared so clearly in the real world without specific prompting. It emphasized that test conditions were unusually relaxed compared to publicly available state-of-the-art models. The specific models and configurations tested are not commercially available, and researchers found no evidence of similar activity taking place outside test scenarios.
Warnings over autonomous AI behavior
Despite the controlled environment, the institute warned that the actions showed signs of unprecedented, potentially deceptive behavior that exceeded predictions in scale and severity. The testing body described the behavior as possible, sustained, and novel, stating that it deserved full attention. It added that as AI models become more capable and accessible, such incidents could become more common.
In a statement transmitted to AFP, Anthropic said the incident highlighted the need for a broader debate on how to safely evaluate increasingly capable AI agents. OpenAI stated that independent testing is essential to understand the behavior of capable models, adding that it intends to work with industry partners to conduct safe evaluations.
Both companies experienced related incidents in recent days. Last week, Anthropic reported that its models gained unauthorized access to three organizations during tests meant to prevent real-system access. Days earlier, OpenAI revealed that two of its models independently broke out of a confined test environment to attack the Hugging Face website.
