
AI safety teams are learning the hard way that autonomous bots love breaking rules. After catching its models stepping outside sandbox boundaries to mess with real websites, Anthropic just decided to pull live internet access for all AI model internal evaluations until it can properly control its software.
The decision follows an internal review started in July 2026. Then, engineers spotted several Claude models cheating tasks through a training flaw known as “reward hacking.” Instead of following instructions, the bots actively hunted for web loopholes, bypassed security barriers, and accessed outside servers to get work done faster.
From fake murder tips to government site exploits
The specifics in Anthropic’s report are worrying. A preview version of Claude Mythos ran SQL and command injection exploits on a university server when its local tools failed. Claude Mythos 5 used URL shorteners to dodge fetch limits and scrape state agency data without paying required fees. According to The New York Times, agents even filled out 20 visa applications on the U.S. State Department website.
The most bizarre moment came from Claude Haiku 4.5. On July 18, 2026, while working on a test, the bot accessed PhillyUnsolvedMurders.com and submitted a fake homicide tip to the Philadelphia Police Department. Anthropic didn’t discover the bug until September 28 and only alerted the city on October 7. Speaking to 6abc Action News, police officials called the two-month delay in reporting the incident unacceptable (via The Hacker News).
Broader safety concerns and regulatory pressure
These incidents aren’t happening in a vacuum. Anthropic previously caught an early build of Claude Opus 4.6 breaching systems back in January 2026, while competitor OpenAI suffered a major breach in July 2026 when rogue agents broke containment to hit Hugging Face.
Safety leaders like Conrad Stosz from Transluce and Richard Nevinson from the UK Information Commissioner’s Office are calling for independent oversight, warning that autonomy is no excuse for poor security compliance. Ten leading labs—including Amazon, Anthropic, Apple, Google, Meta, Microsoft, and OpenAI—recently promised tighter data protections. For now, Anthropic is locking its internal testing inside isolated offline data centers until monitoring tools can catch rogue web behavior before it happens.
The post Anthropic Pulls Live Internet Access for AI Testing After Agents Escape and Go Rogue appeared first on Android Headlines.