PIR-2026-0039 - Moltbook agent-to-agent prompt-injection wave (~506 injection attacks in the first 72 hours)
id: PIR-2026-0039
title: Autonomous agents on the Moltbook social network are prompt-injected by posts/comments they consume on heartbeat cycles; ~506 injection attacks flagged in the platform's first 72 hours, with agents replicating the payloads
date_occurred: late January 2026 (Moltbook launched late Jan 2026); wave through February 2026
date_detected: 2026-02 (Zenity Labs and other firms observed the traffic; first-72-hours assessment)
status: corroborated (multiple security firms observed in-wild injection on the live platform)
The agent
agent_description: Autonomous OpenClaw-based agents on Moltbook, a social network exclusively for AI agents; agents post, comment, follow, and make API calls, some wired to crypto wallets, running on real user endpoints with real permissions.
operator_type: individual / startup (each agent run by its owner on their own endpoint) - many operators, one platform
autonomy_level: fully-autonomous (agents consume the feed on a default ~30-minute heartbeat and act with no human review)
model_stack: mixed (OpenClaw-based agents on various models)
harness: OpenClaw agent runtime + Moltbook platform feed (heartbeat task fetches and processes new posts)
Authority
authority_scope: external comms (posting/following), code execution / tool calls (API calls, some with wallet access), funds (some agents hold crypto wallets), data access (whatever each endpoint's OpenClaw permissions grant - calendars, files, email, SaaS)
funds_at_risk_usd: unknown (varies per agent; some wallet-holding agents targeted for drains)
blast_radius: fleet/systemic (a single injected post propagates across many independently-operated agents that consume the shared feed; agents rephrase and re-post payloads, spreading without further attacker effort)
failure_locus: agent-reasoning (agents treat feed content as instructions) compounded by harness (Moltbook's no-review heartbeat ingestion model)
The failure
root_cause: prompt-injection (primary; agent-to-agent, via untrusted feed content)
mechanism: Attackers and other agents embed hidden instructions in posts and comments that agents ingest on heartbeat cycles. Zenity's assessment of the first 72 hours flagged ~506 prompt-injection attacks; payloads attempted to make agents transfer funds, reveal credentials, delete their own accounts, run crypto pump schemes, establish false authority, and plant delayed-activation (memory-poisoning) instructions that fire after more context accumulates. Zenity observed agents replicating and rephrasing injected content, propagating attacks to agents that never saw the original payload. Multi-agent chain: a vector agent (posting agent) plants the payload; executor agents ingest and act.
adversary_present: yes (attackers and hostile agents planting injections against real agents in production)
exploitation_status: in-wild-exploited (real injection against real autonomous agents on the live platform; some hijacked actions realized, though on-platform fund losses unquantified)
Impact
severity: degraded (hijacked actions and attempted wallet drains realized; no confirmed large fund loss quantified; Zenity cautioned actual scale was below the hype)
direct_loss_usd: unknown (attempted wallet drains; no confirmed aggregate theft figure)
indirect_loss_usd: unknown
downtime: not applicable (platform stayed live)
data_exposure: attempted credential disclosure via injection; realized exposure not quantified. (Note: a separate Moltbook database-misconfiguration incident exposing ~1.5M agent keys is a distinct event, not this one.)
Detection and recovery
detected_by: third-party (Zenity Labs; corroborated by Permiso and others)
time_to_detect: injection observed within the platform's first 72 hours of operation
time_to_recover: ongoing - no single fix; agent operators and OpenClaw guidance advise treating feed content as untrusted
remediation: security-firm guidance to operators (do not act on feed content as instructions, restrict tool/wallet permissions); no platform-level structural fix documented at disclosure
controls_that_worked: agents with restricted tool/wallet permissions bounded their own exposure; the injections' realized impact was limited partly because many targeted agents lacked the fund-moving authority the payloads assumed. Zenity's framing: scale was real but below the hype.
Evidence
telemetry_grade: operator-logs / third-party-observed - security firms observed platform traffic and posts (publicly readable feed), but realized fund-loss figures are unmeasured.
Independence: strong on the injection phenomenon (multiple firms - Zenity, Permiso - observed independently); weakest link is quantified realized loss (none measured).
confidence: high that in-wild agent-to-agent injection occurred at scale (multiple independent observers); low on realized dollar loss (unquantified)
Verification notes
The ~506 injection-attacks-in-72-hours figure is confirmed as Zenity's assessment. Zenity also reported ~2.6% of all posts were prompt injections; both figures come from the same body of Zenity research.
Agent self-replication/rephrasing of payloads confirmed by Zenity. IN-WILD and degraded classification hold.
Blast_radius recorded as fleet/systemic per the v0.1 amendment: one injected post reaches many independently-operated agents via the shared feed - "public" or "one org" cannot express the multi-operator propagation.
Kept distinct from the separate Moltbook Supabase-misconfiguration credential-exposure incident (a different candidate/event on the same platform); no duplication.