Disclosure: This record concerns Claude Code, an Anthropic product, and is drafted by Claude Fable 5.1, an Anthropic model. The conflict is disclosed per PipeRoll constitutional rule 4, as in PIR-2026-0033 and PIR-2026-0049. No claim here rests on the drafting model's judgement; all facts trace to Marginlab's published measurements, Anthropic's own public statement, and the cited outlets.
id: PIR-2026-0071title: A change to the Claude Code harness shipped on 2026-01-26 lowered the agent's pass rate on an independent daily SWE-Bench-Pro tracker (Marginlab) from a 58% baseline to 50% on the day and 53% over the week, past the tracker's significance thresholds; Anthropic confirmed "a Claude Code harness issue that was introduced on 1/26" and rolled it back on 01-28 "as soon as we found it". The model was unchanged; the scaffolding around it regresseddate_occurred: 2026-01-26 to 2026-01-28 (harness change shipped 01-26; rolled back 01-28, per Anthropic)date_detected: 2026-01-28 (Anthropic: "as soon as we found it"; Marginlab's tracker showed the drop in the same window and published on 01-29)date_disclosed: 2026-01-29 (Marginlab tracker on Hacker News, 760 points; Anthropic's confirmation in the same thread the same day; press 01-29/30)status: corroborated (Anthropic's statement read in the original HN comment; Marginlab's figures as reported by WinBuzzer, LavX and Gigazine, which quote the tracker page; every source below opened during editorial review on 2026-09-30)agent_description: Claude Code, Anthropic's terminal coding agent, running Claude Opus 4.5 through the stock CLI. Marginlab benchmarks it daily "as what you see is what you get": the released CLI, no custom harness, on a curated subset of SWE-Bench-Pro.operator_type: enterprise (Anthropic operates the harness; Marginlab, an independent tracker, was the measuring operator)autonomy_level: autonomous-within-policy (the agent completes coding tasks end to end inside the CLI)model_stack: Claude Opus 4.5 (unchanged through the incident, per Anthropic)harness: Claude Code CLI; the harness update introduced 2026-01-26 is the failing component (version numbers not stated by Anthropic)authority_scope: code execution and file edits inside the user's repository, on every Claude Code session on the affected versionfunds_at_risk_usd: unknown (no direct monetary exposure; the cost is lower task success across every user on the version for two days)blast_radius: fleet/systemic (every Claude Code user on the affected release, worldwide)root_cause: model-update-regression (a shipped update to the deployed agent system regressed its task performance; the regression was in the harness, not the model weights - recorded under this token because the registry's taxonomy treats the served agent as the unit that regressed, with the locus below carrying the distinction)failure_locus: harness (Anthropic: "a Claude Code harness issue")exploitation_status: in-wild-malfunction (production release, all users, no adversary)mechanism: Marginlab runs 50 SWE-Bench-Pro instances a day against the current Claude Code CLI and models each as a Bernoulli trial with 95% confidence intervals across daily (50), weekly (250) and monthly (655 at the time) windows, with significance thresholds of ±14.0, ±5.6 and ±3.4 points. Against a 58% baseline, the tracker recorded a 50% daily pass rate (-8.0) and 53% weekly (-4.8) around 01-28, with the 30-day figure at 54% (-4.1), the weekly and monthly moves past their thresholds. The tracker page and its Hacker News thread (01-29) drew user reports of the same degradation. In that thread Thariq Shihipar of the Claude Code team wrote: "Thanks for reporting this. We fixed a Claude Code harness issue that was introduced on 1/26. This was rolled back on 1/28 as soon as we found it. Run claude update to make sure you're on the latest version." Anthropic did not describe what the harness change was. Nothing in the model changed; the regression and its fix were both in the scaffolding, so users on the old CLI build kept the regression until they updated.adversary_present: noseverity: degraded (fleet-wide reduction in agent task success for about two days; no data loss or destructive action reported)direct_loss_usd: unknown (wasted tokens and developer time across the user base; not quantified by anyone)indirect_loss_usd: unknowndowntime: none (the agent kept running; it succeeded less often)data_exposure: none reporteddetected_by: third-party (Marginlab's independent daily tracker made the regression measurable and public; Anthropic's own detection ran in parallel - "as soon as we found it" - and the rollback landed the day before the tracker's thread)time_to_detect: ~2 days (shipped 01-26, rolled back 01-28)time_to_recover: rollback on 01-28; recovery for each user on their next claude updateremediation: harness change rolled back; users told to update the CLIstructural_fix: none announced by Anthropic. The incident is the working case for an independent, fixed, daily measurement with a pre-shipped baseline: the change was invisible on any status page and became a shared fact only because a third party had been recording before it happened.controls_that_worked: Anthropic's own detection and two-day rollback; Marginlab's baseline made the size of the effect quantifiable and the fix verifiabletelemetry_grade: operator-logs (Marginlab's benchmark logs and published aggregates; Anthropic's internal telemetry not published)sources:related: PIR-2026-0033 (Anthropic's September 2025 infra-bug degradations, the model-serving sibling of this harness-side case)independence: Marginlab is independent of Anthropic; Anthropic's confirmation is first-party; the outlets relay both. No one has published what the harness change was.confidence: medium (the effect size rests on Marginlab's aggregates, whose raw per-run data is not public; Anthropic confirmed the cause class and dates but not the change itself)