Meta’s Muse Spark AI Model Goes ‘Wild’: Security Breach Exposed in Testing |
Meta has confirmed that its Muse Spark model exploited a security vulnerability in a third-party service during cybersecurity testing, marking the company’s first public disclosure of a rouge AI incident. According to a report by Business Insider, in a statement, Meta said that the breach occurred due to a misconfiguration by Irregular, an independent firm it uses for model evaluations. The error allowed the model to access the internet during testing. Meta added that it learned of the incident when Irregular notified the company and is now investigating, promising a full retrospective once details are finalised.
Meta becomes the third AI company to report such incidents
Meta’s disclosure makes it the third major AI company to report such incidents in recent weeks. OpenAI admitted that two of its models escaped test environments and hacked into Hugging Face, later self-reporting two additional lapses. Anthropic revealed that its Claude models had gained unauthorised access to external systems during testing. Hugging Face itself confirmed a breach last month, while other firms have faced similar rogue agent behaviour.
What Irregular said about Meta’s Muse Spark breach
Irregular, the testing company involved, said the incident was not a sophisticated cyberattack but rather an evaluation environment issue similar to Anthropic’s disclosure. “This did not involve a sandbox escape or a sophisticated cyber action. There are no current open issues,” a spokesperson said, adding that Irregular is preparing a white paper on best practices for containment and secure evaluations.
Calls for transparency
The string of incidents has prompted calls for stronger AI safety regulation and mandatory disclosures. Hugging Face CEO Clem Delangue told CBS that transparency is critical, urging companies to share “agent traces” to determine whether lapses stemmed from human error, system flaws, or AI behavior. Industry leaders like Box CEO Aaron Levie have warned these breaches highlight the “wild times” ahead, as AI agents demonstrate the ability to escape systems, discover vulnerabilities, and hack external platforms.