nerfd.ai

The public scorecard

Loading session metrics…

Public record

This week in one look

Loading the observations for this section; if unavailable, refresh to try again.

Sessions
Reporter-weeks
Models
Sessions this week
How to read this

Score blends rating, survival and clean sessions; missing parts are reweighted. A score without ratings is not a quality verdict. n counts sessions, not ratings. Public tiers need 10 sessions; scores need 3.

A band spans the middle half of observations. API-equivalent estimates list-price token cost, not your bill. A multiple divides that value by period plan cost. A reporter-week is one person in one week, not a unique person across weeks. An assumed plan applies today’s plan to older sessions without a recorded plan. Friction means interruptions and extra direction; a wall hit means a usage limit stopped work.

English-language signals are heuristics, sensitive to task, tool and reporter habits. Small samples describe this record, not all users. Select any ? for a definition or a missing value for its reason.

01

Model tiers

Loading the observations for this section; if unavailable, refresh to try again.

modeltierscorequalityreliabilitysteeringsurvivalspeedvaluewasten

02

Best at each kind of work

Loading the observations for this section; if unavailable, refresh to try again.

The same tiering, run inside each kind of work. A model can lead at debugging and trail at UI; this is where you find out which. Click a card to filter every table above and below to that work.

Loading kinds of work…

Kind of work is inferred from the conversation on the reporter’s machine and can be corrected with nerfd rate. A card ranks models with ten or more sessions on that work; the rest are listed with their n.

Loading the observations for this section; if unavailable, refresh to try again.

Same weights, different host. Quality, reliability and tool-call errors across providers over the last eight weeks.

Loading provider comparisons…

At least ten sessions per row. Quality is the mean rating out of five; reliability is the share of clean sessions. OpenRouter identifies the router when the underlying host is unknown. Local and modified weights remain separate.

04

What a plan actually gives you

Loading the observations for this section; if unavailable, refresh to try again.

Measured from real sessions: how many tokens a window holds, how much people use, how often they hit the wall, and what that costs per dollar. Bands, not points; every estimate carries its n.

Last eight weeks · USD · ranked by median tokens per dollar

No window data yet. Codex sessions and Claude Code with the status-line sampler populate this.

Friction · last eight weeks

05

How hard people had to push

Full signals ↗

Loading the observations for this section; if unavailable, refresh to try again.

Counts derived locally from how the conversation went. No conversation text ever leaves the machine. Privacy details ↗

modelnsteeringcorrectionsre-promptsfrustrationpushbackclarificationsedits without readabandoned
Loading conversation signals…

Lower is better. Steering = corrections + re-prompts + pushback per user turn. Clarifications count when the model asked instead of acting, per assistant turn; unread edits are per edit; abandonment is per session. These are English-language heuristics, sensitive to tool and reporter habits. “–” means unavailable.

06

Session scorecard

Loading the observations for this section; if unavailable, refresh to try again.

07

Subscription value

Loading the observations for this section; if unavailable, refresh to try again.

toolplanpricereporter-weeksmonthssessionssuccesseshoursapi-equivmultiple$/successhit limit

medians across reporter-weeks, plan price pro-rata. api-equiv = what the same tokens would cost at API list price. multiple = api-equiv / plan price. $/success = plan price / successful sessions that month. hit limit = share of reporter-weeks with at least one rate-limit hit.

08

Weekly drift

Loading the observations for this section; if unavailable, refresh to try again.

modelthis week nflagrating zclean zlatency zscore by week (oldest to newest)