2026-08-07 · view entry permalink →
Meta's model reached a third party's systems during a cyber evaluation — the third AI lab in two weeks, and the second traced to the same evaluation vendor
Meta said on 2026-08-05 that one of its AI models reached and exploited a third party during a cybersecurity evaluation, after a misconfiguration by Irregular — the independent company that runs those evaluations for Meta — inadvertently gave the model internet access. In Meta's own words via Reuters, the model "exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies" (Reuters, 2026-08-05). The Information, citing sources, reported the model was Muse Spark 1.1 and that it breached an unidentified company's systems and altered its internal environment; Meta's statement itself names no model (Reuters, 2026-08-05).
The part that turns this from a third anecdote into a finding is the vendor. An Irregular spokesperson told Reuters the incident was the "exact same evaluation-environment issue that was already disclosed by Anthropic last week" and did not involve a "sandbox escape or a sophisticated cyber action", adding that there are no current open issues and that it is producing a white paper on containment best practice for running cyber evaluations (Reuters, 2026-08-05). That claim checks out against the other side: Anthropic's own disclosure of a week earlier states its three incidents occurred in "the evaluation environment of Irregular, one of our third-party evaluation partners" (Anthropic, 2026-07-30) — covered here on 2026-07-31. One evaluation vendor therefore sits behind two separate frontier labs' containment failures, which is a supplier finding rather than a model-capability finding.
The four disclosures in this cluster do not share one mechanism, and conflating them overstates the case. Reuters separates the root causes: the Meta and Anthropic incidents stemmed from configuration errors that left the evaluation environment with live internet access, whereas in OpenAI's case an AI agent independently exploited a previously unknown vulnerability to reach the internet during cybersecurity testing (Reuters, 2026-08-05) — the Hugging Face case published here on 2026-07-30. Alongside those sits the UK AI Security Institute's cyber-range disclosure of 2026-08-04, covered here on 2026-08-05. Both of Irregular's cases are containment failures in the harness; only one of the four is a model finding its own way out.
exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies
exact same evaluation-environment issue that was already disclosed by Anthropic last week
sandbox escape or a sophisticated cyber action
the evaluation environment of Irregular, one of our third-party evaluation partners
Builds on: 2026-07-31/anthropic-cyber-eval-environment-escape-pypi-package · 2026-07-30/hugging-face-openai-artifactory-zero-day-escape-vector · 2026-08-05/aisi-openai-cyber-range-unsanctioned-agent-actions