TL;DR
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A pre-registered benchmark called livenerf is running daily tests on Claude Opus 5.5 to detect whether the model degrades after launch. As of September 29, 2026, six of 30 days are collected with no results yet; the first verdict on possible degradation is possible around October 24, 2026.
An independent benchmark project called livenerf is roughly a fifth of the way through a 30-day effort to determine, with statistics rather than anecdote, whether Anthropic’s Claude Opus 5.5 gets quietly worse after launch. As of September 29, 2026, the project had collected six of 30 planned daily runs with none missed, according to its GitHub repository — but no verdict on the model’s performance is possible until around October 24, 2026.
Livenerf, created by a developer using the handle ninjahawk, is an append-only, pre-registered benchmark built to answer a single question, as the repository puts it: “does a model get worse after it ships?” The project was motivated by months of community claims that Anthropic “nerfs” models days or weeks after release — through quantization, model swaps, reduced effort, or routing changes — claims the project notes could equally be people “pattern-matching on noise.” Claude Opus 5.5 launched on September 22, 2026, and the benchmark’s clock started on September 24, 2026 at 22:10 UTC, about 2.5 days after launch.
The methodology is designed to remove subjectivity. The test panel consists of 78 questions selected from 2,336 screened GPQA Diamond, MMLU-Pro, competition-math and AIME 2025–26 items — specifically questions Opus 5.5 answers inconsistently. The harness uses frozen prompts, a pinned CLI version (2.1.280), and a fixed harness hash, running 90 samples per day through headless Claude Code on a Max subscription. The statistical approach follows Anthropic’s own published methodology (“Adding Error Bars to Evals”) on top of the UK AI Security Institute’s open-source Inspect framework, which the project says means “there’s nothing homebrew to argue about.”
All six completed days ran the full 90 samples on the same harness and CLI. The project disclosed one deviation: day 5 ran with the budget guard overridden once, logged in a deviations file. The first results row lands after day 20, and days 1–10 form the baseline against which two subsequent 10-day windows are compared.
Why a Statistical Nerf Detector Matters
Debates over whether AI labs secretly degrade shipped models have historically been unresolvable, as the project frames it, because nobody holds a clean day-0 baseline: “every argument ends up as vibes versus vibes.” Livenerf matters because it converts that dispute into a measurable question with a pre-registered decision rule — improvements are reported “just as loudly as regressions,” which cuts against confirmation bias in either direction.
The project is also unusually transparent about its own limits. Its validation run showed the instrument cannot distinguish Opus 5.5 from Opus 5 at 99% confidence (−3.8 ± 6.3 points accuracy, −23% tokens) at validation sample sizes, meaning a same-family model swap of that magnitude would likely go undetected. It can reliably detect a change of about 7.5 accuracy points per 10-day window. Notably, the project found that reduced effort shows up in output token counts before accuracy moves: a simulated low-effort condition produced −62% output tokens against −8.3 points of accuracy.
Top picks for "livenerf opus nerf"
As an affiliate, we earn on qualifying purchases.
Months of Quiet-Nerf Allegations
Claims that Anthropic quietly degrades models after release have circulated in AI developer communities for months, with suggested mechanisms including quantization, substituting a smaller model behind the same name, or routing changes. Livenerf’s premise is that none of these claims had a controlled baseline to test against, making them impossible to adjudicate.
The panel construction itself was calibrated carefully: of 2,336 screened questions, Opus 5.5 scored about 93% correct on first attempt, and 97% of questions were answered consistently right or wrong. The 78 inconsistent questions form the panel. The project measured its own selection bias — fresh-sample pass rates rose from 54.7% to 62.0% on selected items — and used the fresh rates for power calculations. A report-only audit found 8 likely-wrong answer keys and 30 ambiguous questions among the 78; none were dropped, but a pre-registered sensitivity analysis will rerun results without them. Samples touched by an apparent safety classifier routing to Opus 5 are rejected and excluded.
“For months there have been reports that Anthropic ‘nerfs’ models some days or weeks after release… It could also mean nothing happened and people are pattern-matching on noise.”
— livenerf GitHub repository (ninjahawk)
No Answer Yet — and Built-In Blind Spots
No result exists yet: the baseline is still collecting, and the first possible nerf call arrives around October 24, 2026. Any current claim that Opus 5.5 has or has not been degraded is unsupported by this benchmark.
Even when results arrive, detection limits apply. A same-family model swap producing Opus 5-sized changes may be undetectable despite the larger 10-day windows containing about 2.5 times the validation sample count — the project states this “hasn’t been shown to be enough.” The panel also contains acknowledged question-quality issues (8 wrong answer keys, 30 ambiguous items), and the primary metric’s practical sensitivity is roughly 7.5 accuracy points per window. Anthropic has not, in the source material, commented on the benchmark or the nerf allegations.
The Road to October 24
Four more baseline days complete the days 1–10 baseline window, after which two 10-day comparison windows run. The first results table row appears after day 20, and the first formal decision on whether Opus 5.5 changed post-launch is possible around October 24, 2026. The pre-registered sensitivity analysis excluding flagged questions will run alongside the main result. The repository says it will maintain a running 10-day table of scores and token counts against the launch-week baseline, and invites scrutiny via its GitHub Discussions tab.
Key Questions
Has Claude Opus 5.5 been nerfed according to livenerf?
No verdict exists yet. As of September 29, 2026, only 6 of 30 days were collected, all within the baseline window. The first possible determination is around October 24, 2026.
How does livenerf decide a model got worse?
It compares paired per-item scores on a locked 78-question panel against the days 1–10 launch baseline, using clustered standard errors and Anthropic’s published eval statistics. Output token counts serve as a secondary early-warning signal.
What changes can the benchmark actually detect?
Per the project’s power calculation, one daily run detects accuracy changes of about 7.5 points per 10-day window. Its validation showed it could not distinguish Opus 5 from Opus 5.5 at 99% confidence at validation sample sizes, so a same-family swap of that size may go undetected.
Why can’t the tests be fully deterministic?
The project states sampling parameters are unavailable and thinking cannot be turned off in the Claude Code interface, so it instead freezes everything controllable — prompts, CLI version, harness, and graders — and measures drift statistically over thousands of samples.
Is the benchmark independent of Anthropic?
Yes, it is a community-run project on a Claude Max subscription, though it deliberately uses Anthropic’s own published statistical methodology and the UK AI Security Institute’s Inspect framework to avoid methodological disputes.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
