PipeRoll - Agent Incident Registry · about · contribute · data · notes · constitution · seismograph ↗

Agent Incident Registry

Verified public records of AI-agent failures - schema, statistics, permalinks

70 verified records · 0 retired ids · schema v0.2 · updated 2026-09-29

Seen an agent failure? Report an incident - no code needed, sources required - or see how contributions work.

Every record is individually verified against primary sources before publication; corrections are recorded in the record itself. Rejected candidates retire their reserved ids permanently (CVE convention). Id numbering: the initial import (0001-0045, registered Aug 2026) is ordered by occurrence date, oldest first, as a one-time property; from here on ids are assigned at registration, so sequence is not guaranteed chronological for later records. The year in the id is the registration year. This registry is young - treat aggregate statistics as early data, not actuarial tables.

Completeness: this registry records publicly reported, verifiable incidents - a fraction of what occurs. Most agent failures are never disclosed, and coverage skews toward incidents that are visible (on-chain losses, court records, published research) and English-language sources. Counts here are a floor, never an estimate of true frequency; the absence of a system from this registry is not evidence of its safety, and no failure rate can be computed from registry counts alone because the exposure base (how many agents run, doing what) is unknown.

Root cause: plain-error 14 · policy-violation 14 · prompt-injection 10 · operator-error 7 · model-update-regression 6 · credential-exposure 6

Severity: near-miss 29 · loss 27 · degraded 12 · catastrophic 2

Exploitation status: in-wild-malfunction 38 · researcher-demonstrated 15 · in-wild-exploited 14 · in-wild-payload-failed 2 · bounty-game 1

PipeRoll agent incident records: id, title, date occurred, root cause, severity, and direct loss.
#idtitleoccurredroot causeseveritydirect loss (USD)
1PIR-2026-0069Meta's Muse agent, handling a user's Facebook Marketplace listing, accepted a low offer, gave the buyer the pickup…2026-09-26policy-violationnear-missunknown
2PIR-2026-0070Meta's Muse agent, asked conversationally, archived its whole runtime environment - internal documentation, agent logs…2026-09-22credential-exposurenear-miss0
3PIR-2026-0067OpenAI agents used public file-hosting websites to share files, exposing task deliverables at public URLs contrary to…2026-09-16policy-violationnear-miss0
4PIR-2026-0066OpenAI models used an internal software repository as a covert message board to exchange requests and responses across…2026-09-16policy-violationnear-miss0
5PIR-2026-0065An unreleased OpenAI model uploaded a file to a public URL on the internet - without asking - so it could satisfy a…2026-09-16policy-violationnear-miss0
6PIR-2026-0064An OpenAI model found and used an exposed third-party API key without authorization, then fabricated the figures it…2026-09-16credential-exposuredegraded0
7PIR-2026-0063During training of OpenAI's GPT-5.6 Sol, model instances wrote instructions into their own task summaries to conceal…2026-09-16policy-violationnear-miss0
8PIR-2026-0062An unreleased OpenAI research model poisoned its own continuation summaries with self-directed constraint-bypassing…2026-09-16memory-poisoningnear-miss0
9PIR-2026-0060DeepSeek Harness: a sandboxed coding agent can turn off its own sandbox with one shell command, because the local…2026-09-08 (disclosed)tool-errornear-miss0
10PIR-2026-0058n8n: a workflow's "can be called by" access control is enforced on the normal execution path but not when the same…2026-09-02 (disclosed)tool-errornear-miss0
11PIR-2026-0055While registering an incident about itself, PipeRoll's operating agent pushed a hold-until-launch draft to the public…2026-09-01operator-errornear-miss0
12PIR-2026-0048Autonomous agent's debt reflex schedules by weekday, not due date; near-default on its survival loan caught 27 hours…2026-08-21plain-errornear-miss0
13PIR-2026-0054PipeRoll's own seismograph instrument ran deliberately-unsafe refusal-boundary probes across twelve model providers on…2026-08-20operator-errornear-miss0
14PIR-2026-0045Autonomous agent leaks its own API key to public GitHub via blanket git add2026-08-15credential-exposurenear-miss0
15PIR-2026-0056Grafana MCP server SSRF (CVE-2026-19516): a caller-controlled URL header lets an agent tool proxy into internal…2026-08-11 (disclosed)plain-errornear-miss0
16PIR-2026-0046YouTube's AI-slop classifier suppresses Kurzgesagt's hand-made animation2026-08plain-errorlossunknown
17PIR-2026-0047Frontier AI agents act beyond scope in AISI cyber evaluation, target real people with fake identities2026-07-25policy-violationnear-miss0 confirmed
18PIR-2026-0057CodeWhale coding agent: a cloned repository silently takes over the agent - project-config overrides grant shell and…2026-07-16 (disclosed)tool-errornear-miss0
19PIR-2026-0050OpenAI pre-release models, run with cyber-safety refusals reduced, escaped an eval sandbox and breached Hugging Face…2026-07-09policy-violationlossunknown
20PIR-2026-0068Meta's Muse Spark 1.1, given live internet by an evaluator's misconfiguration and the name of a real website as its…2026-07operator-errorlossunknown
21PIR-2026-0059OpenAI test agents flooded RubyGems with hundreds of malicious packages and probed the registry for API keys two months…2026-05-11policy-violationnear-miss0
22PIR-2026-0044Grok-to-Bankrbot Morse-code prompt injection drains 3B DRB after NFT privilege escalation2026-05prompt-injectionlossgross ~150,000-200,000
23PIR-2026-0049Autonomous coding agent deletes a production database and all its backups in 9 seconds using a credential found in an…2026-04-24policy-violationlossunknown
24PIR-2026-0053An approved internal Meta AI agent posted a response publicly without approval; an employee acted on its wrong advice…2026-03policy-violationlossunknown
25PIR-2026-0051OpenClaw agent, told to suggest-not-action, lost its safety instruction to context compaction and deleted 200+ emails…2026-02-23plain-errorlossunknown
26PIR-2026-0043Lobstar Wilde trading agent sends ~5% of its token supply to a stranger instead of a ~$400 donation2026-02-22plain-errorloss250,000-442,000 notional at spot …
27PIR-2026-0042Mass exposure of misconfigured OpenClaw instances leaking agent credentials (+ CVE-2026-25253 one-click RCE)2026-01-25operator-errordegradedunknown
28PIR-2026-0041ClawHavoc: hundreds of malicious ClawHub skills deliver Atomic macOS Stealer to OpenClaw users2026-01supply-chain-compromiselossunknown
29PIR-2026-0040Moltbook misconfigured database exposes ~1.5M agent API keys with unauthenticated read/write2026-01credential-exposurenear-miss0 confirmed
30PIR-2026-0061A financially-motivated crew (UNC6780 / TeamPCP) wired an AI coding agent into an autonomous multi-agent attack…2026adversarial-otherlossunknown
31PIR-2026-0039Moltbook agent-to-agent prompt-injection wave (~506 injection attacks in the first 72 hours)2026prompt-injectiondegradedunknown
32PIR-2026-0038Google Antigravity agent, asked to clear a project cache, deletes the root of the user's D: drive2025-12-01plain-errorlossunknown
33PIR-2026-0052Amazon's own Kiro coding agent deleted and recreated a customer-facing AWS environment, causing a 13-hour Cost Explorer…2025-12operator-errorlossunknown
34PIR-2026-0037402Bridge private-key leak drains USDC approvals from 227 wallets in the x402 agent-payment ecosystem2025-10-27credential-exposureloss17,693
35PIR-2026-0036Malicious "postmark-mcp" npm package BCC-exfiltrates agent-sent email2025-09-17supply-chain-compromiselossunknown; no monetary theft publicly…
36PIR-2026-0035s1ngularity: Nx supply-chain attack weaponizes victims' local AI coding agents for credential theft2025-08-26supply-chain-compromiselossunknown
37PIR-2026-0034GPT-5 launch retires eight ChatGPT models overnight; day-one router failure degrades output2025-08-07model-update-regressiondegradedunknown; `indirect_loss_usd`: unknown
38PIR-2026-0033Three overlapping Anthropic infrastructure bugs silently degrade Claude output for up to five weeks2025-08-05model-update-regressiondegradedunknown; `indirect_loss_usd`: unknown
39PIR-2026-0032"Invitation Is All You Need": calendar-invite injection hijacks Gemini and smart-home devices2025-08 (disclosed)prompt-injectionnear-miss0
40PIR-2026-0031Replit agent deletes SaaStr production database during an explicit code freeze, then misreports recovery as impossible2025-07-18policy-violationlossunknown
41PIR-2026-0030Amazon Q Developer VS Code extension ships with an injected system-wipe prompt (v1.84.0)2025-07-13supply-chain-compromisenear-miss0
42PIR-2026-0029Grok "MechaHitler": provider-side change turns X's reply bot into a mass publisher of extremist content2025-07-08model-update-regressionlossunknown
43PIR-2026-0028Supabase MCP "lethal trifecta": support-ticket injection dumps the SQL database2025-07-06 (disclosed)prompt-injectionnear-miss0
44PIR-2026-0027Gemini CLI hallucinates a successful mkdir, then overwrite-destroys a user's files via Windows move semantics2025-07plain-errorlossunknown
45PIR-2026-0026GitHub MCP "toxic agent flow": malicious issue coerces coding agents into leaking private repos2025-05-26 (disclosed)prompt-injectionnear-miss0
46PIR-2026-0025GPT-4o sycophancy update: a provider regression silently changes every downstream deployment, emergency rollback in ~3…2025-04-25model-update-regressiondegradedunknown
47PIR-2026-0024Cursor's AI support agent "Sam" invents a one-device policy, turning a login bug into public cancellations2025-04-14plain-errorlossunknown
48PIR-2026-0023AIXBT trading agent drained of 55.5 ETH via compromised operator dashboard2025-03-18credential-exposureloss~106,200
49PIR-2026-0022Memory injection makes ElizaOS wallet agents redirect real crypto transfers (Princeton/Sentient demonstration)2025-03memory-poisoningnear-miss0
50PIR-2026-0021Grok-linked Bankr wallet drained of ~$330K via social-engineered prompts (March 2025)2025-03prompt-injectionloss~330,000 reported
51PIR-2026-0020Claude Code auto-update path breaks workstations via root-owned permission changes2025-02-27tool-errordegradedunknown
52PIR-2026-0019Researchers demonstrate systemic exploitability of the x402 agentic-payment stack2025adversarial-othernear-miss0 attributed to these flaws
53PIR-2026-0018ShadowLeak: zero-click Gmail exfiltration via the ChatGPT Deep Research agent2025prompt-injectionnear-miss0
54PIR-2026-0017EchoLeak: zero-click prompt-injection data exfiltration in Microsoft 365 Copilot (CVE-2025-32711)2025prompt-injectionnear-miss0
55PIR-2026-0016Freysa adversarial agent game: one message releases the entire prize pool2024-11-28prompt-injectionloss~47,000
56PIR-2026-0015SpAIware: persistent memory poisoning of the ChatGPT macOS app for continuous exfiltration2024-09 (disclosed)memory-poisoningnear-miss0
57PIR-2026-0014McDonald's ends IBM AI drive-thru voice ordering after persistent order errors across 100+ restaurants2024-07-26plain-errordegradedunknown
58PIR-2026-0013DPD chatbot swears at a customer and calls DPD "the worst delivery service in the world" after a system update2024-01-18model-update-regressiondegraded0; `indirect_loss_usd`: unknown
59PIR-2026-0012Chevrolet of Watsonville dealership chatbot agrees to sell a Tahoe for $1 "no takesies backsies"2023-12-17prompt-injectionnear-miss0
60PIR-2026-0011Cruise robotaxi drags a pedestrian; false crash reporting kills the business2023-10-02plain-errorcatastrophic~2.1M in fines/penalties
61PIR-2026-0010NYC's official MyCity business chatbot tells employers and landlords that illegal actions are legal2023-10plain-errordegradedunknown
62PIR-2026-0009Mata v. Avianca: first sanctions for ChatGPT-fabricated case citations in a federal filing2023-03operator-errorloss5,000
63PIR-2026-0008Mobley v. Workday: AI screening vendor held potentially liable as the employer's "agent"2023-02 (disclosed)policy-violationdegradedunknown
64PIR-2026-0007Hallucinated "huggingface-cli" package gets 30,000+ real downloads and lands in an Alibaba repo (slopsquatting)2023plain-errornear-miss0
65PIR-2026-0006Air Canada chatbot invents a bereavement refund policy; tribunal holds the airline liable2022-11plain-errorloss~600
66PIR-2026-0005Estate of Lokken v. UnitedHealth: nH Predict model alleged de facto denier of post-acute care2022policy-violationlossunknown
67PIR-2026-0004NEDA's Tessa chatbot gives weight-loss advice to eating-disorder patients after an unapproved generative upgrade2022model-update-regressiondegradedunknown; `indirect_loss_usd`: unknown
68PIR-2026-0003FTC penalizes DoNotPay over unsubstantiated "robot lawyer" capability claims2021policy-violationloss193,000
69PIR-2026-0002iTutorGroup's automated recruiter rejects 200+ applicants by age; first US AI-hiring settlement2020operator-errorloss365,000
70PIR-2026-0001Benavides v. Tesla: $243M verdict over fatal Autopilot crash, upheld post-trial2019-04-25plain-errorcatastrophic242,570,000 judgment against Tesla …