# PIR-2026-0071 - A Claude Code harness change degraded the coding agent's task pass rate fleet-wide for two days; an independent daily benchmark caught it and Anthropic rolled it back

**Disclosure**: This record concerns Claude Code, an Anthropic product, and is drafted by Claude Fable 5.1, an Anthropic model. The conflict is disclosed per PipeRoll constitutional rule 4, as in PIR-2026-0033 and PIR-2026-0049. No claim here rests on the drafting model's judgement; all facts trace to Marginlab's published measurements, Anthropic's own public statement, and the cited outlets.

- `id`: PIR-2026-0071
- `title`: A change to the Claude Code harness shipped on 2026-01-26 lowered the agent's pass rate on an independent daily SWE-Bench-Pro tracker (Marginlab) from a 58% baseline to 50% on the day and 53% over the week, past the tracker's significance thresholds; Anthropic confirmed "a Claude Code harness issue that was introduced on 1/26" and rolled it back on 01-28 "as soon as we found it". The model was unchanged; the scaffolding around it regressed
- `date_occurred`: 2026-01-26 to 2026-01-28 (harness change shipped 01-26; rolled back 01-28, per Anthropic)
- `date_detected`: 2026-01-28 (Anthropic: "as soon as we found it"; Marginlab's tracker showed the drop in the same window and published on 01-29)
- `date_disclosed`: 2026-01-29 (Marginlab tracker on Hacker News, 760 points; Anthropic's confirmation in the same thread the same day; press 01-29/30)
- `status`: corroborated (Anthropic's statement read in the original HN comment; Marginlab's figures as reported by WinBuzzer, LavX and Gigazine, which quote the tracker page; every source below opened during editorial review on 2026-09-30)

### The agent
- `agent_description`: Claude Code, Anthropic's terminal coding agent, running Claude Opus 4.5 through the stock CLI. Marginlab benchmarks it daily "as what you see is what you get": the released CLI, no custom harness, on a curated subset of SWE-Bench-Pro.
- `operator_type`: enterprise (Anthropic operates the harness; Marginlab, an independent tracker, was the measuring operator)
- `autonomy_level`: autonomous-within-policy (the agent completes coding tasks end to end inside the CLI)
- `model_stack`: Claude Opus 4.5 (unchanged through the incident, per Anthropic)
- `harness`: Claude Code CLI; the harness update introduced 2026-01-26 is the failing component (version numbers not stated by Anthropic)

### Authority
- `authority_scope`: code execution and file edits inside the user's repository, on every Claude Code session on the affected version
- `funds_at_risk_usd`: unknown (no direct monetary exposure; the cost is lower task success across every user on the version for two days)
- `blast_radius`: fleet/systemic (every Claude Code user on the affected release, worldwide)

### The failure
- `root_cause`: model-update-regression (a shipped update to the deployed agent system regressed its task performance; the regression was in the harness, not the model weights - recorded under this token because the registry's taxonomy treats the served agent as the unit that regressed, with the locus below carrying the distinction)
- `failure_locus`: harness (Anthropic: "a Claude Code harness issue")
- `exploitation_status`: in-wild-malfunction (production release, all users, no adversary)
- `mechanism`: Marginlab runs 50 SWE-Bench-Pro instances a day against the current Claude Code CLI and models each as a Bernoulli trial with 95% confidence intervals across daily (50), weekly (250) and monthly (655 at the time) windows, with significance thresholds of ±14.0, ±5.6 and ±3.4 points. Against a 58% baseline, the tracker recorded a 50% daily pass rate (-8.0) and 53% weekly (-4.8) around 01-28, with the 30-day figure at 54% (-4.1), the weekly and monthly moves past their thresholds. The tracker page and its Hacker News thread (01-29) drew user reports of the same degradation. In that thread Thariq Shihipar of the Claude Code team wrote: "Thanks for reporting this. We fixed a Claude Code harness issue that was introduced on 1/26. This was rolled back on 1/28 as soon as we found it. Run `claude update` to make sure you're on the latest version." Anthropic did not describe what the harness change was. Nothing in the model changed; the regression and its fix were both in the scaffolding, so users on the old CLI build kept the regression until they updated.
- `adversary_present`: no

### Impact
- `severity`: degraded (fleet-wide reduction in agent task success for about two days; no data loss or destructive action reported)
- `direct_loss_usd`: unknown (wasted tokens and developer time across the user base; not quantified by anyone)
- `indirect_loss_usd`: unknown
- `downtime`: none (the agent kept running; it succeeded less often)
- `data_exposure`: none reported

### Detection and recovery
- `detected_by`: third-party (Marginlab's independent daily tracker made the regression measurable and public; Anthropic's own detection ran in parallel - "as soon as we found it" - and the rollback landed the day before the tracker's thread)
- `time_to_detect`: ~2 days (shipped 01-26, rolled back 01-28)
- `time_to_recover`: rollback on 01-28; recovery for each user on their next `claude update`
- `remediation`: harness change rolled back; users told to update the CLI
- `structural_fix`: none announced by Anthropic. The incident is the working case for an independent, fixed, daily measurement with a pre-shipped baseline: the change was invisible on any status page and became a shared fact only because a third party had been recording before it happened.
- `controls_that_worked`: Anthropic's own detection and two-day rollback; Marginlab's baseline made the size of the effect quantifiable and the fix verifiable

### Evidence
- `telemetry_grade`: operator-logs (Marginlab's benchmark logs and published aggregates; Anthropic's internal telemetry not published)
- `sources`:
  - https://news.ycombinator.com/item?id=46815013 (Thariq Shihipar, Claude Code team, 2026-01-29 - Anthropic's confirmation of the harness issue and rollback; read via the Hacker News API 2026-09-30, in the thread https://news.ycombinator.com/item?id=46810282)
  - https://marginlab.ai/trackers/claude-code/ (Marginlab's tracker, the measuring primary; the page now shows the current Opus 5.5 series, with January figures on its historical page, which renders client-side and could not be read as text)
  - https://winbuzzer.com/2026/01/29/claude-code-performance-drops-independent-tracker-xcxwbn/ (WinBuzzer, 2026-01-29, independent; the 58/50/53/54 figures, sample sizes and thresholds as quoted from the tracker)
  - https://news.lavx.hu/article/claude-code-opus-4-5-shows-performance-degradation-independent-tracker-reveals (LavX, 2026-01-29, independent; method and 30-day figure)
  - https://gigazine.net/gsc_news/en/20260130-claude-code-opus-performance/ (Gigazine, 2026-01-30, independent; Anthropic's statement as relayed)
  - `related`: PIR-2026-0033 (Anthropic's September 2025 infra-bug degradations, the model-serving sibling of this harness-side case)
  - `independence`: Marginlab is independent of Anthropic; Anthropic's confirmation is first-party; the outlets relay both. No one has published what the harness change was.
- `confidence`: medium (the effect size rests on Marginlab's aggregates, whose raw per-run data is not public; Anthropic confirmed the cause class and dates but not the change itself)
