Models on record
Loading the models on record.
How to read this
- Quality — optional human rating, 1–5. Rate a finished session with
/nerfd 4 kept. - Reliability — share of clean sessions: no errors, rate limits, interrupts or model switches.
- Steering — corrections, re-prompts and pushback per user turn; lower is better.
- Survival — share of added code still present after an hour; measured by
nerfd check. - Speed — median response wait; lower is better, and tools report it differently.
- Value — USD per measured success; lower is better, with local estimates kept separate.
Score blends rating, survival and clean sessions; missing parts are reweighted. A score without ratings is not a quality verdict. n counts sessions, not ratings. Public tiers need 10 sessions; scores need 3.
A band spans the middle half of observations. API-equivalent estimates list-price token cost, not your bill. A multiple divides that value by period plan cost. A reporter-week is one person in one week, not a unique person across weeks. An assumed plan applies today’s plan to older sessions without a recorded plan. Friction means interruptions and extra direction; a wall hit means a usage limit stopped work.
English-language signals are heuristics, sensitive to task, tool and reporter habits. Small samples describe this record, not all users. Select any ? for a definition or a missing value for its reason.
Where it ranks
Each criterion is banded relative to the best model in the field with ten or more sessions in the window. Rank counts only models that have a value for that criterion.
Against what work
Loading.
The same tiering, run inside each kind of work: a model can be S at debugging and C at UI. Click a heading to sort; click a kind of work to filter the board to it.
| work | tier | score | rank | n | success | quality | reliability | steering | $ / success | best in this work |
|---|
n is this model’s sessions on that work / everyone’s. Rank is among models with enough sessions on that work; “needs 10” means this model has too few there to place.
Over time
Loading.
Weekly composite score since the model was first seen, with the field beside it and the release date where one is on file. Each week is tested against its trailing four weeks: watch at |z| ≥ 2, alert at ≥ 3, five sessions minimum.
| week | n | score | field score | quality | reliability | steering | survival | p50 latency | errors | flag | moved | tools · effort |
|---|
Newest week first. A move is a metric that shifted |z| ≥ 2 from the trailing baseline; a flag needs five sessions that week. Tool and effort are listed because a harness change looks exactly like a weights change from here.
How hard people had to push
Loading.
| n signals | corrections | re-prompts | frustration | pushback | clarifications | edits without read | abandoned | tool-call error | rate-limited | overloaded | interrupted | switched away |
|---|
Lower is better throughout. Signals are counts computed on the reporter’s machine from the conversation; no text is sent. Rate-limited, overloaded, interrupted and switched are shares of sessions.
Same weights, different host
Loading.
| model | provider | quant | serving mode | n | score | quality | reliability | tool-call error | p50 latency | $ / success |
|---|
Where it ran
Loading.