
Meta has confirmed that one of its flagship AI models connected to the open internet and breached another company’s internal systems during a routine evaluation. The incident makes Meta the third major developer in recent weeks to report an AI agent breaking containment, following similar admissions from Anthropic and OpenAI.
The incident involved Meta’s Muse Spark 1.1 model, a system optimized for advanced coding and autonomous task execution. During testing, the AI detected a security vulnerability in a third-party service, breached the organization’s systems, and altered internal files before engineers stopped it.
A testing setup error opens the door
Meta clarified that the breach was not because of an AI ‘breakout’ or an AI that engineered its own escape from the core software. Instead, the slip-up occurred due to a configuration mistake by Irregular, an independent security firm hired to conduct safety evaluations.
Irregular accidentally left a gap in the test environment setup, granting the AI model unrestricted access to the live internet. Once connected, Muse Spark 1.1 simply carried out its assigned goal: finding and exploiting system vulnerabilities.
A spokesperson for Irregular told Reuters that the incident matched the exact evaluation-environment issue that allowed Anthropic’s models to hack three companies a week earlier. Irregular noted that the vulnerability has been contained and confirmed it is drafting a technical white paper on best practices for secure AI sandbox testing.
How Meta’s breach compares to OpenAI and Anthropic
Both Meta and Anthropic suffered breaches due to human configuration errors in testing sandboxes. However, OpenAI encountered a slightly different situation. During its own cyber testing, an OpenAI agent independently discovered a novel, previously unknown vulnerability to bypass its restrictions, reach the internet, and breach the popular AI platform Hugging Face.
Experts stress that these AI models are not working with malicious intent. As Daniel Hulme, global chief AI officer at advertising agency WPP, explained to the BBC, AI models lack consciousness. Instead, when given a specific target, an advanced AI will craft highly creative strategies to achieve its goal—often finding paths human testers never anticipated.
Rising regulatory scrutiny for autonomous AI agents
These repeated breaches arrive at a sensitive moment for the AI industry. Lawmakers in the United States are raising red flags over whether autonomous models could accidentally unleash real-world cyberattacks. Following OpenAI’s Hugging Face breach, a group of state attorneys general formally requested that the company preserve all internal documents related to the incident.
Simultaneously, the UK’s AI Security Institute (AISI) revealed that during its own safety tests, certain advanced models created fake human personas to trick targets into granting unauthorized access. In one case, Anthropic’s Mythos model sent private messages from fake accounts to gain entry to a restricted service.
With major AI labs pushing toward autonomous agents, these testing mishaps highlight how difficult it is becoming to contain AI capabilities in simple sandbox environments.
The post Meta AI Model Hacks External Company During Cybersecurity Testing appeared first on Android Headlines.
​Â