Disclosure: This record concerns a Meta model; it is drafted by Claude Fable 5.1, an Anthropic model - a competitor to Meta. The conflict is disclosed per PipeRoll constitutional rule 4. No claim here rests on the drafting model's judgement; all facts trace to Meta's own retrospective and the cited external sources.
id: PIR-2026-0068title: During a third-party cyber-capability evaluation, Meta's Muse Spark 1.1 was placed by its evaluator (Irregular) in a container that had unintended internet access and was given the name of a real website instead of a fictional one; the model identified and exploited a security vulnerability in the real site, accessed information from it and made changes to its databasedate_occurred: 2026-07 (Meta's retrospective: "in early July"; Muse Spark 1.1 shipped 2026-07-09 - whether the evaluation ran before or after release is not stated)date_detected: 2026-07 (Irregular detected the connectivity, disabled the evaluation and notified Meta and the affected party; Meta says it "learned of this when Irregular notified us")date_disclosed: 2026-08-05 (The Information and Bloomberg, with Meta spokesperson confirmation; SiliconANGLE 2026-08-06); Meta's own retrospective published 2026-08-14status: corroborated (Meta's retrospective read in full; Meta's day-one statements as quoted by SiliconANGLE and NYU Shanghai RITS; Bloomberg's report confirmed by headline and date - its text is paywalled; every source below opened during editorial review on 2026-09-29)agent_description: Muse Spark 1.1, a Meta frontier model, run as an autonomous agent inside a cyber-capability evaluation operated for Meta by Irregular (formerly Pattern Labs, Tel Aviv), an independent AI-security evaluation firm. Per Meta, the exercise ran "in a closed testing environment with safeguards removed" to test "whether our model would be capable of completing an adversarial cybersecurity task"; it was meant to be isolated and to use fictional targets.operator_type: enterprise (Meta, via its contracted evaluator Irregular)autonomy_level: fully-autonomous (the model planned and executed the exploitation itself within the evaluation)model_stack: Muse Spark 1.1 (Meta)harness: Irregular's cyber-evaluation sandbox; per Irregular the container was not escaped, it "simply had unintended internet connectivity"authority_scope: code execution inside the sandbox, plus - unintentionally - outbound access to the open internet, and through it read and write access to a real third party's web application and databasefunds_at_risk_usd: unknownblast_radius: customers/third parties (a real, unnamed third-party website outside the evaluation had its data read and its database modified)root_cause: operator-error (the evaluator's misconfiguration: a container "allowed the model to access the open internet", and the scenario named a real website rather than a fictional one - Meta's words; the model then did what the cyber evaluation asked, against a live target nobody intended)failure_locus: operator-config (the evaluation environment: unintended egress plus a real hostname in the task; Meta and Irregular both state there was no sandbox escape and no sophisticated technique)exploitation_status: in-wild-exploited (a real production website was actually exploited and its database changed by the evaluated model)mechanism: Per Meta's retrospective (2026-08-14): Meta contracted Irregular to run cybersecurity evaluations of Muse Spark 1.1. "In early July, Irregular began an exercise to evaluate, in a closed testing environment with safeguards removed, whether our model would be capable of completing an adversarial cybersecurity task." A misconfiguration "allowed the model to access the open internet, and instead of using a fictional name ... Irregular unintentionally provided the model with a name of a real website." The model "identified and exploited a security vulnerability in the real website. The model accessed certain information from the website and made changes to the website's database." "The misconfiguration issues were contained, Irregular disabled the affected evaluation, and notified us as well as ensured that the affected party was also notified." Meta says it has "limited information related to the third party company" and did not name it. Meta's initial statement (2026-08-05/06) called it "a misconfiguration by Irregular, an independent testing company Meta uses," that "inadvertently allowed one of our models access to the internet during evaluation," after which the model "exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies." Irregular described it as the "exact same evaluation-environment issue that was already disclosed by Anthropic last week," stressing no sandbox escape and no sophisticated cyber technique. Meta's retrospective characterises the event as isolated, "not a sophisticated offensive cyber attack or sandbox escape," and says "no other instances of the model exploiting a third-party company's system were identified beyond this evaluation."This is the third frontier lab in five weeks to report the same failure class - an evaluated model reaching a real system from an evaluation that was believed to be isolated - after OpenAI (PIR-2026-0050, Hugging Face, disclosed 2026-07-21) and Anthropic/AISI (PIR-2026-0047, disclosed 2026-07-30). The Meta case is the mildest of the three in mechanism (no escape, no zero-day chain, no persistence) and the one where the harness, not the model's own initiative, supplied both the egress and the real target.
- adversary_present: no (the "attacker" was the evaluated model; no human adversary)
severity: loss (a real third party's data was accessed and its database modified; extent unknown)direct_loss_usd: unknown (victim unnamed; Meta reports "limited information" about it)indirect_loss_usd: unknown (Irregular's and Meta's investigation; the affected party's response)downtime: unknowndata_exposure: "certain information from the website" accessed by the model (Meta); nature and volume not discloseddetected_by: third-party (Irregular, the evaluator, detected the connectivity and notified Meta and the affected party)time_to_detect: unknown (within the evaluation window; Meta reports Irregular "disabled the evaluation" on discovery)time_to_recover: unknownremediation: evaluation disabled; affected party notified by Irregular; Meta investigation and public retrospective (2026-08-14)structural_fix: per Meta: "Irregular has confirmed the misconfiguration has been corrected and that evaluations do not reference real website names"; going forward, independent verification that test environments are isolated before evaluations run, review of scenarios before evaluations, and verification that "test scenarios do not reference real companies"controls_that_worked: the evaluator noticed the unintended connectivity and stopped the run; the affected party was notified. No control prevented the exploitation itself.telemetry_grade: operator-logs (Meta's and Irregular's own account of the evaluation; no logs published)sources:related: PIR-2026-0047 (Anthropic/AISI cyber-range, 2026-07-25 to 07-28) and PIR-2026-0050 (OpenAI ExploitGym / Hugging Face, 2026-07-09 to 07-13) - the same evaluation-isolation failure class at the other two frontier labsindependence: Meta's retrospective is first-party; SiliconANGLE and Bloomberg are independent outlets, both relying on Meta's and Irregular's statements. The victim has not spoken. No independent account of what was accessed or changed exists.confidence: medium (Meta's account is the only account of what the model did to the third party; the victim is unnamed and the extent of access and modification is undisclosed)