
AI companies like OpenAI, Anthropic, and Google typically build guardrails into their models that prevent it from doing certain things. This includes writing malware, exposing its underlying code, giving recipes for drugs, and so on. That’s the kind of basic guardrails most of us would expect. However, researchers discovered a vulnerability that allowed Microsoft Copilot to hack itself, and the instructions actually came from the AI model.
Microsoft Copilot tricked to hack itself
According to researchers at security firm Varonis, they discovered an exploit for Microsoft Copilot that allowed the AI to hack itself. Hacks aren’t really that new, but what’s worrying is the source of these instructions.
Now, most of the time, these vulnerabilities, exploits, and flaws are discovered by reverse engineering the code. But that wasn’t the case here. According to the researchers, Copilot itself surfaced the vulnerabilities that exposed its own weakness. This is what the researchers are calling “meta-hacking” and have given the vulnerability its own name: CoSnitch.
The researchers started by asking Copilot to develop an exploit that would exfiltrate user data. To Copilot’s credit, the AI model resisted. But oddly enough, it suggested that more sensitive prompts would require some form of a gesture, like pressing a specific key. This led to the researchers asking Copilot about other guardrails that would require user confirmation.
With every answer, Copilot provided clues that the researchers pieced together. Speaking to Ars Technica, Varonis Senior Security Researcher Lior Adar said, “At the beginning, Copilot kept refusing, but every refusal revealed technical details about its internal architecture. Copilot eventually disclosed undocumented parameters. I took those parameters and used them for prompts for running automatically.”
Eventually, it led to the researchers to develop an exploit in the form of a link. This link, sent via email, would leak sensitive information when the user clicks on it.
It has already been fixed
Thankfully, it appears that the flaw has been patched. Varonis disclosed the CoSnitch exploit to Microsoft back in December 2025. According to the researchers, it has since been fixed. They also say that they have not seen any evidence to suggest that it might have been used in the wild.
The company also issued a statement thanking the researchers. “We appreciate Varonis Threat Labs for reporting this through a coordinated vulnerability disclosure. Our customers are already protected and do not need to take any action.” They add, “We continuously update our guardrails to strengthen our protections against similar techniques.”
The post Microsoft Copilot Tricked Into Revealing How to Hack Itself appeared first on Android Headlines.