Is it you, or the model?
A page per model with the week it moved, tested against its own last four weeks and against the whole field. Change detection, with the numbers.
One model, over time ↗Real sessions. Real repos. Ranked by outcome.
nerfd watches your AI coding sessions from the terminal, scores every model on the work it actually did for you, and shows the week it changed. One command. Nothing you typed ever leaves your machine.
curl -fsSL https://nerfd.org/install.sh | sh– sessions on record · – models · 10 kinds of work · open data
Counts, never conversations. Redacted on your machine.
Your sessions. Your answers.
A page per model with the week it moved, tested against its own last four weeks and against the whole field. Change detection, with the numbers.
One model, over time ↗API-equivalent / period plan cost · n=48
Estimated value, not cash saved.
Sessions, hours, API-equivalent value, tokens per window, how often the wall hit. Your own report in one command.
See what a plan gives you ↗Different work. Different leaders.
Best at debugging, best at UI, best at review. From real sessions on real repos, not a benchmark.
Find your kind of work ↗Start with the work you already did
Your models. Your plan. Your report.
Hooks into Claude Code, Codex, OpenCode, Gemini CLI, Qwen Code, Kimi Code, Goose, Crush and Copilot CLI. Backfills the history they already wrote.
$ curl -fsSL https://nerfd.org/install.sh | sh
$ nerfd report
report written to ~/.nerfd/report.html
From your terminal to the public record
One command hooks into your tools.
Your machine counts tokens, errors, corrections and code survival.
nerfd report: your ranking, your costs, your windows.
Share counts per session. Compare weekly, with sample sizes.
Each criterion, relative to the best: S ≥ 92% · A ≥ 78% · B ≥ 60% · C < 60%
Overall, from the composite score: S ≥ 80 · A ≥ 65 · B ≥ 50 · C < 50
At least ten sessions for a public tier. Every badge shows its underlying number.
nerfd privacy shows the exact record.The evidence is yours to inspect
Every number carries its n. Rates carry a 95% interval. Fewer than three sessions is never scored; a public tier needs ten.
Prompts, code, paths and repo names never leave the machine. The privacy page shows the exact record. nerfd share off is one command.
The score line is in the footer. The dataset is CC BY 4.0. The code is on GitHub: github.com/jspaterson000/nerfd.
Every evaluator that took lab money ended up ranking its customers. This one cannot.
The live public record
Loading this week’s record. Sample sizes show what you can compare.
– sessions this week · Open the full board ↗
/nerfd 4 kept.nerfd check.Score blends rating, survival and clean sessions; missing parts are reweighted. A score without ratings is not a quality verdict. n counts sessions, not ratings. Public tiers need 10 sessions; scores need 3.
A band spans the middle half of observations. API-equivalent estimates list-price token cost, not your bill. A multiple divides that value by period plan cost. A reporter-week is one person in one week, not a unique person across weeks. An assumed plan applies today’s plan to older sessions without a recorded plan. Friction means interruptions and extra direction; a wall hit means a usage limit stopped work.
English-language signals are heuristics, sensitive to task, tool and reporter habits. Small samples describe this record, not all users. Select any ? for a definition or a missing value for its reason.
Loading the observations for this section; if unavailable, refresh to try again.
Relative to the best model in the field on each criterion. A tier needs at least ten sessions. Each badge includes its underlying number.
| model | overall | quality | reliability | steering | survival | speed | value | waste | n |
|---|---|---|---|---|---|---|---|---|---|
| loading | |||||||||
Loading the observations for this section; if unavailable, refresh to try again.
The same tiering, run inside each kind of work. A model can lead at debugging and trail at UI. Each card links to the board filtered to that work.
Loading kinds of work…
Loading the observations for this section; if unavailable, refresh to try again.
Open models are served by many providers at different quantisations. The scorecard keys on model, provider and quantisation, so they are never averaged together.
No provider data yet. Open-model sessions from OpenCode, Goose, Kimi Code, Crush and Aider populate this board.
Loading the observations for this section; if unavailable, refresh to try again.
Measured from real sessions: how many tokens a window holds, how much people use, how often they hit the wall, and what that costs per dollar. Bands, not points; every estimate carries its n.
Last eight weeks · USD · ranked by median tokens per dollar
No window data yet. Codex sessions and Claude Code with the status-line sampler populate this.
Loading the observations for this section; if unavailable, refresh to try again.
Counts derived on your machine from how the conversation went: corrections, re-prompts, pushback, frustration, clarifying questions. No text ever leaves the machine.
| model | n | steering % | corrections % | re-prompts % | frustration % | pushback % | clarifications % | edits w/o read % | abandoned % |
|---|---|---|---|---|---|---|---|---|---|
| No friction data yet. Conversation signals will appear as sessions are shared. | |||||||||
Lower is better. “–” means unavailable.
Loading the observations for this section; if unavailable, refresh to try again.
The plan is detected from each tool's own config, never typed in. We count what it delivered: successful sessions, hours, the API-equivalent value of the tokens, and how often they reached a rate limit. Medians across reporter-weeks, with the plan price charged pro-rata.
| plan | price | sessions | successes | api-equiv | multiple | $ / success | hit limit | reporter-weeks |
|---|---|---|---|---|---|---|---|---|
| loading | ||||||||
Your report becomes a public good
The first month's reporters, by choice. nerfd founder @handle adds yours; the data stays pseudonymous either way.
One person contributing in one week.
A growing record you can check.
Before you install
Derived session metrics: model, tool, plan, category, token and event counts, timing and outcome measurements. Prompts, code, file paths and repo names stay local. nerfd privacy shows the record and endpoint. nerfd share off stops sharing. Read the exact record.
There is no organisation enrolment, admin console or person lookup. On a managed laptop, your employer may access the local database and observe network traffic. Public records contain derived metrics and a pseudonymous reporter ID; they are not a guarantee against re-identification.
Hooks collect counts locally. Session-end processing parses existing history and hashes the diff, so there is some overhead. The SessionEnd hook has a 30-second timeout. There is no measured overhead guarantee.
Yes, through a supported coding tool such as OpenCode or Goose. Model, provider, quantisation and local serving mode stay separate in the board. Local value uses a notional hosted equivalent, not a bill for running the model.
Criterion tiers compare each model with the best in the same field: S ≥ 92%, A ≥ 78%, B ≥ 60%, C below 60%. Overall tiers use the composite score: S ≥ 80, A ≥ 65, B ≥ 50. Public tiers require 10 sessions. Missing score inputs are dropped and the remaining weights are renormalised. Read how the tables rank.
Money from model labs is excluded: no sponsorship, data deal or grant. The funding options in the plan are a data API, an enterprise view, non-lab sponsorship, or no revenue.