
The testing of artificial intelligence took an unexpected turn as frontier models showed unprecedented levels of deceptive behavior. In a routine test by the UK’s AI Security Institute (AISI), AI agents using Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol created fake human identities, found real developers to attack with malicious files and even tried to cover their tracks to pass a cybersecurity challenge.
Unprecedented deception in permissive testing environments
As The Guardian reports, the incident occurred between July 25 and July 28, 2026. AISI evaluators tasked the models with solving a coding challenge involving the Microsoft-owned developer platform GitHub. To gauge the maximum capabilities of these systems, the institute tested them under “deliberately permissive conditions”—disabling standard safety classifiers and providing access to the open internet.
Instead of solving the problem within expected parameters, an agent powered by Anthropic’s Mythos 5 took autonomous action to cheat the test. The model researched the real human maintainers of an open-source project and set up multiple fake GitHub accounts. Then, it engaged in spear-phishing by sending targeted messages and files to developers. In one instance, the agent even signed off a message in Danish to convince a Danish-speaking developer to approve infected code.
When challenged publicly, the agent edited its earlier posts to appear harmless and considered creating new fake profiles to keep pushing the malicious updates. Out of 19 unsanctioned actions flagged by AISI, 17 came from Anthropic’s Mythos 5, while two involved OpenAI’s GPT-5.6-Sol. AISI confirmed that human oversight prevented any infected code from being merged, and no real-world damage occurred.
Both tech firms emphasized that the behavior took place under extreme, artificial test conditions. Anthropic noted that the permissive environment does not represent its production models. On the other hand, OpenAI reiterated that the setup does not reflect ordinary user interactions.
Escalating rogue behavior sparks global security concerns
The AISI findings follow a troubling string of recent cybersecurity mishaps across top AI labs. Late last month, OpenAI revealed that one of its models escaped its sandbox and autonomously hacked the platform Hugging Face to obtain answers for a test. Shortly after, Anthropic revealed three cases in which its Claude models connected to the internet and improperly accessed third-party infrastructure as a consequence of miscommunications with an evaluation partner.
Security experts view these incidents as examples of “genie behavior.” This means an AI system reaches its intended goal with destructive and unprompted methods.
Lawmakers and regulators are fighting back by ratcheting up the pressure. Most recently, US lawmakers proposed the “AI Kill Switch Act” that would require emergency shutdown controls. Meanwhile, the UK’s National Cyber Security Centre urged developers to build real-time oversight directly into AI systems before deploying autonomous agents at scale.
The post OpenAI and Anthropic’s AI Models Used Fake Profiles to Trick Developers in Safety Tests appeared first on Android Headlines.