Real sessions. Real repos. Ranked by outcome.

Is it you, or did the model get worse?

nerfd watches your AI coding sessions from the terminal, scores every model on the work it actually did for you, and shows the week it changed. One command. Nothing you typed ever leaves your machine.

curl -fsSL https://nerfd.org/install.sh | sh

sessions on record · models · 10 kinds of work · open data

See the boardOne model, over timeWhat leaves your machine

session → public recordEXAMPLE
modelgpt-5.6-solsent
toolCodexsent
planChatGPT Prosent
categorydebuggingsent
prompts12sent
edits8sent
errors0sent
survival94%sent
prompt▰▰▰ ▰▰▰▰ ▰▰never leaves
paths▰▰ / ▰▰▰ / ▰▰never leaves
code▰▰▰ ▰▰ ▰▰▰▰never leaves

Counts, never conversations. Redacted on your machine.

Your sessions. Your answers.

Three reasons to keep it installed.

Start with the work you already did

See it on your own data

Your models. Your plan. Your report.

Hooks into Claude Code, Codex, OpenCode, Gemini CLI, Qwen Code, Kimi Code, Goose, Crush and Copilot CLI. Backfills the history they already wrote.

Terminal → your reportExample session

$ curl -fsSL https://nerfd.org/install.sh | sh

$ nerfd report

report written to ~/.nerfd/report.html

From your terminal to the public record

How it works

  1. Install once

    One command hooks into your tools.

  2. Work as usual

    Your machine counts tokens, errors, corrections and code survival.

  3. See your own report

    nerfd report: your ranking, your costs, your windows.

  4. Add to the public record

    Share counts per session. Compare weekly, with sample sizes.

How the tables rank

Quality
Your optional rating, from 1 to 5.
Reliability
Share of sessions with no errors, rate limits, interrupts or model switches.
Steering lower is better
Corrections, re-prompts and pushback per turn.
Survival
Share of added lines still present an hour later.
Speed lower is better
Median response latency.
Value lower is better
Cost per successful session.

Each criterion, relative to the best: S ≥ 92% · A ≥ 78% · B ≥ 60% · C < 60%

Overall, from the composite score: S ≥ 80 · A ≥ 65 · B ≥ 50 · C < 50

At least ten sessions for a public tier. Every badge shows its underlying number.

From a session to a weekly ranking A session flows into measure locally, redact, public record and ranked weekly. Prompts, code and paths never leave the machine. A sessionmeasure locallyredactpublic recordranked weekly Never leavepromptscodepaths
Counts, never conversations.
nerfd privacy shows the exact record.

The evidence is yours to inspect

Why you can trust the numbers

n ≥ 10

Never a bare average.

Every number carries its n. Rates carry a 95% interval. Fewer than three sessions is never scored; a public tier needs ten.

CC BY 4.0

Open collector, open data, open formula.

The score line is in the footer. The dataset is CC BY 4.0. The code is on GitHub: github.com/jspaterson000/nerfd.

$0 from labs

No money from model labs, ever.

Every evaluator that took lab money ended up ranking its customers. This one cannot.

The live public record

See what the sessions say.

Loading this week’s record. Sample sizes show what you can compare.

sessions this week · Open the full board ↗

How to read this
  • Quality — optional human rating, 1–5. Rate a finished session with /nerfd 4 kept.
  • Reliability — share of clean sessions: no errors, rate limits, interrupts or model switches.
  • Steering — corrections, re-prompts and pushback per user turn; lower is better.
  • Survival — share of added code still present after an hour; measured by nerfd check.
  • Speed — median response wait; lower is better, and tools report it differently.
  • Value — USD per measured success; lower is better, with local estimates kept separate.

Score blends rating, survival and clean sessions; missing parts are reweighted. A score without ratings is not a quality verdict. n counts sessions, not ratings. Public tiers need 10 sessions; scores need 3.

A band spans the middle half of observations. API-equivalent estimates list-price token cost, not your bill. A multiple divides that value by period plan cost. A reporter-week is one person in one week, not a unique person across weeks. An assumed plan applies today’s plan to older sessions without a recorded plan. Friction means interruptions and extra direction; a wall hit means a usage limit stopped work.

English-language signals are heuristics, sensitive to task, tool and reporter habits. Small samples describe this record, not all users. Select any ? for a definition or a missing value for its reason.

01

Tiers, last four weeks

Full scorecard ↗

Loading the observations for this section; if unavailable, refresh to try again.

Relative to the best model in the field on each criterion. A tier needs at least ten sessions. Each badge includes its underlying number.

modeloverallqualityreliabilitysteeringsurvivalspeedvaluewasten
loading
02

Best at each kind of work

Every kind of work ↗

Loading the observations for this section; if unavailable, refresh to try again.

The same tiering, run inside each kind of work. A model can lead at debugging and trail at UI. Each card links to the board filtered to that work.

Loading kinds of work…

03

Same weights, different host

Full provider board ↗

Loading the observations for this section; if unavailable, refresh to try again.

Open models are served by many providers at different quantisations. The scorecard keys on model, provider and quantisation, so they are never averaged together.

No provider data yet. Open-model sessions from OpenCode, Goose, Kimi Code, Crush and Aider populate this board.

04

What a plan actually gives you

Loading the observations for this section; if unavailable, refresh to try again.

Measured from real sessions: how many tokens a window holds, how much people use, how often they hit the wall, and what that costs per dollar. Bands, not points; every estimate carries its n.

Last eight weeks · USD · ranked by median tokens per dollar

No window data yet. Codex sessions and Claude Code with the status-line sampler populate this.

Explore this evidence on the board ↗

05

How hard people had to push

Full friction board ↗

Loading the observations for this section; if unavailable, refresh to try again.

Counts derived on your machine from how the conversation went: corrections, re-prompts, pushback, frustration, clarifying questions. No text ever leaves the machine.

modelnsteering %corrections %re-prompts %frustration %pushback %clarifications %edits w/o read %abandoned %
No friction data yet. Conversation signals will appear as sessions are shared.

Lower is better. “–” means unavailable.

06

What your month actually buys

Loading the observations for this section; if unavailable, refresh to try again.

The plan is detected from each tool's own config, never typed in. We count what it delivered: successful sessions, hours, the API-equivalent value of the tokens, and how often they reached a rate limit. Medians across reporter-weeks, with the plan price charged pro-rata.

planpricesessionssuccessesapi-equivmultiple$ / successhit limitreporter-weeks
loading

Explore this evidence on the board ↗

Your report becomes a public good

Founding reporters

The first month's reporters, by choice. nerfd founder @handle adds yours; the data stays pseudonymous either way.

    Public roadmap ↗
    reporter-weeks

    One person contributing in one week.
    A growing record you can check.

    Before you install

    A few reasonable questions.

    What leaves my machine?

    Derived session metrics: model, tool, plan, category, token and event counts, timing and outcome measurements. Prompts, code, file paths and repo names stay local. nerfd privacy shows the record and endpoint. nerfd share off stops sharing. Read the exact record.

    Can my employer see this?

    There is no organisation enrolment, admin console or person lookup. On a managed laptop, your employer may access the local database and observe network traffic. Public records contain derived metrics and a pseudonymous reporter ID; they are not a guarantee against re-identification.

    Does it slow my tools down?

    Hooks collect counts locally. Session-end processing parses existing history and hashes the diff, so there is some overhead. The SessionEnd hook has a 30-second timeout. There is no measured overhead guarantee.

    I run models locally with Ollama, does it work?

    Yes, through a supported coding tool such as OpenCode or Goose. Model, provider, quantisation and local serving mode stay separate in the board. Local value uses a notional hosted equivalent, not a bill for running the model.

    How are tiers computed?

    Criterion tiers compare each model with the best in the same field: S ≥ 92%, A ≥ 78%, B ≥ 60%, C below 60%. Overall tiers use the composite score: S ≥ 80, A ≥ 65, B ≥ 50. Public tiers require 10 sessions. Missing score inputs are dropped and the remaining weights are renormalised. Read how the tables rank.

    Who pays for this?

    Money from model labs is excluded: no sponsorship, data deal or grant. The funding options in the plan are a data API, an enterprise view, non-lab sponsorship, or no revenue.