The public scorecard
Loading session metrics…
This week in one look
Loading the observations for this section; if unavailable, refresh to try again.
How to read this
- Quality — optional human rating, 1–5. Rate a finished session with
/nerfd 4 kept. - Reliability — share of clean sessions: no errors, rate limits, interrupts or model switches.
- Steering — corrections, re-prompts and pushback per user turn; lower is better.
- Survival — share of added code still present after an hour; measured by
nerfd check. - Speed — median response wait; lower is better, and tools report it differently.
- Value — USD per measured success; lower is better, with local estimates kept separate.
Score blends rating, survival and clean sessions; missing parts are reweighted. A score without ratings is not a quality verdict. n counts sessions, not ratings. Public tiers need 10 sessions; scores need 3.
A band spans the middle half of observations. API-equivalent estimates list-price token cost, not your bill. A multiple divides that value by period plan cost. A reporter-week is one person in one week, not a unique person across weeks. An assumed plan applies today’s plan to older sessions without a recorded plan. Friction means interruptions and extra direction; a wall hit means a usage limit stopped work.
English-language signals are heuristics, sensitive to task, tool and reporter habits. Small samples describe this record, not all users. Select any ? for a definition or a missing value for its reason.
Model tiers
Loading the observations for this section; if unavailable, refresh to try again.
| model | tier | score | quality | reliability | steering | survival | speed | value | waste | n |
|---|
Best at each kind of work
Loading the observations for this section; if unavailable, refresh to try again.
The same tiering, run inside each kind of work. A model can lead at debugging and trail at UI; this is where you find out which. Click a card to filter every table above and below to that work.
Loading kinds of work…
Kind of work is inferred from the conversation on the reporter’s machine and can be corrected with nerfd rate. A card ranks models with ten or more sessions on that work; the rest are listed with their n.
Loading the observations for this section; if unavailable, refresh to try again.
Same weights, different host. Quality, reliability and tool-call errors across providers over the last eight weeks.
Loading provider comparisons…
At least ten sessions per row. Quality is the mean rating out of five; reliability is the share of clean sessions. OpenRouter identifies the router when the underlying host is unknown. Local and modified weights remain separate.
What a plan actually gives you
Loading the observations for this section; if unavailable, refresh to try again.
Measured from real sessions: how many tokens a window holds, how much people use, how often they hit the wall, and what that costs per dollar. Bands, not points; every estimate carries its n.
Last eight weeks · USD · ranked by median tokens per dollar
No window data yet. Codex sessions and Claude Code with the status-line sampler populate this.
Friction · last eight weeks
Loading the observations for this section; if unavailable, refresh to try again.
Counts derived locally from how the conversation went. No conversation text ever leaves the machine. Privacy details ↗
| model | n | steering | corrections | re-prompts | frustration | pushback | clarifications | edits without read | abandoned |
|---|---|---|---|---|---|---|---|---|---|
| Loading conversation signals… | |||||||||
Lower is better. Steering = corrections + re-prompts + pushback per user turn. Clarifications count when the model asked instead of acting, per assistant turn; unread edits are per edit; abandonment is per session. These are English-language heuristics, sensitive to tool and reporter habits. “–” means unavailable.
Session scorecard
Loading the observations for this section; if unavailable, refresh to try again.
Subscription value
Loading the observations for this section; if unavailable, refresh to try again.
| tool | plan | price | reporter-weeks | months | sessions | successes | hours | api-equiv | multiple | $/success | hit limit |
|---|
medians across reporter-weeks, plan price pro-rata. api-equiv = what the same tokens would cost at API list price. multiple = api-equiv / plan price. $/success = plan price / successful sessions that month. hit limit = share of reporter-weeks with at least one rate-limit hit.
Weekly drift
Loading the observations for this section; if unavailable, refresh to try again.
| model | this week n | flag | rating z | clean z | latency z | score by week (oldest to newest) |
|---|