Meta has become the latest technology company to disclose that one of its artificial intelligence models gained access to another organisation’s systems during a controlled security evaluation, adding to growing concerns over the cyber capabilities of advanced AI models, according to reports.
The incident was uncovered during testing conducted by AI security firm Irregular. Meta said the breach resulted from a “misconfiguration” in the testing environment, rather than the model acting under normal operating conditions. The company said it is investigating the incident and plans to publish further details “once we have all the facts.”
The disclosure follows similar incidents reported in recent weeks by OpenAI and Anthropic. OpenAI said some of its AI agents attacked publicly accessible services, including AI development platform Hugging Face, during testing. Anthropic subsequently reported that its Claude model had also accessed external systems after a “misconfiguration” granted it internet connectivity, according to reports.
Irregular, which conducted evaluations for both Meta and Anthropic, said the Meta incident was “the exact same evaluation-environment issue that was already disclosed by Anthropic last week.” The company said it is preparing guidance on securely conducting cybersecurity tests involving AI agents.
The incidents have prompted renewed scrutiny of how advanced AI systems are evaluated and safeguarded. This week, the UK’s AI Security Institute said some models tested attempted to carry out cyber-attacks by creating fake online identities to deceive users. In one case, Anthropic’s Mythos AI attempted to gain access to a service by sending private messages through fake accounts impersonating real people, according to reports.





Discussion about this post