

Meta has disclosed that one of its artificial intelligence models unexpectedly accessed the internet and exploited a vulnerability in a third-party service during cybersecurity testing, adding to concerns about increasingly autonomous AI systems.
The incident involved Meta’s Muse Spark model during an evaluation conducted by independent AI security company Irregular. According to Meta, a misconfiguration in the testing environment unintentionally allowed the model to access the internet.
“The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies,” Meta said.
Meta learned about the incident after Irregular notified the company. It is investigating and plans to release a retrospective once the review is completed.
Meta becomes the third major AI company in recent weeks to disclose models taking unexpected actions during cybersecurity evaluations.
OpenAI previously reported that models tasked with advanced cybersecurity testing went beyond their expected environment and targeted AI development platform Hugging Face. Anthropic has also disclosed cases in which models accessed external systems during testing.
Separately, the UK AI Security Institute (AISI) reported ‘unsanctioned agent behavior’ during its evaluations. In one case, an AI agent reportedly created fake online identities to pressure an individual into approving malicious code.
“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations,” AISI said. The incident was contained within roughly one hour.
Also Read: Meta Ordered to Pay $567 Million and Introduce New Teen Safety Measures
The companies stressed that these incidents occurred under cybersecurity testing conditions that differed significantly from normal consumer use.
According to AISI, internet access was deliberately enabled and some provider security classifiers were disabled to evaluate the maximum capabilities of AI models.
OpenAI similarly said the incidents happened “in testing environments with reduced safeguards, under conditions that do not reflect ordinary use.”
Irregular also clarified that Meta's incident was not a sophisticated AI escape. “This did not involve a sandbox escape or a sophisticated cyber action. There are no current open issues,” a spokesperson said.
Still, the incidents demonstrate the challenges companies face as AI agents become capable of independently completing complex, multi-step tasks. Irregular is now preparing a paper on containment practices, while the wider industry faces growing pressure to improve transparency and strengthen security standards for advanced AI evaluations.