THE MODEL LIBRARY
Get to know the models.
Community results, task by task. Find a model, follow its history, and see the evidence behind its reputation.
Loading the models on record.
How to read this
- Quality — optional human rating, 1–5. Rate a finished session with
/nerfd 4 kept. - Reliability — share of clean sessions: no errors, rate limits, interrupts or model switches.
- Steering — corrections, re-prompts and pushback per user turn; lower is better.
- Survival — share of added code still present after an hour; measured by
nerfd check. - Speed — median response wait; lower is better, and tools report it differently.
- Value — USD per measured success; lower is better, with local estimates kept separate.
Score blends rating, survival and clean sessions; missing parts are reweighted. A score without ratings is not a quality verdict. n counts sessions, not ratings. Public tiers need 10 sessions; scores need 3.
A band spans the middle half of observations. API-equivalent estimates list-price token cost, not your bill. A multiple divides that value by period plan cost. A reporter-week is one person in one week, not a unique person across weeks. An assumed plan applies today’s plan to older sessions without a recorded plan. Friction means interruptions and extra direction; a wall hit means a usage limit stopped work.
English-language signals are heuristics, sensitive to task, tool and reporter habits. Small samples describe this record, not all users. Select any ? for a definition or a missing value for its reason.
Performance at a glance.
Six ways to judge the work. Each tier compares models with ten or more sessions and an available measurement. Steering, speed and value are separate from the overall score.
Follow the changes.
Loading.
Track this model alongside the whole field. Inspect a week for its score and sample size. Changes in tools, effort and task mix can also move the results.
Explore the weekly evidence
| week | n | score | field score | quality | reliability | steering | survival | p50 latency | errors | flag | moved | tools · effort |
|---|
Newest week first. A move is a metric that shifted |z| ≥ 2 from the trailing baseline; a flag needs five sessions that week. Tool and effort are listed because a harness change looks exactly like a weights change from here.
Find its kind of work.
Loading.
Results vary by task. Select a kind of work to compare this model with the rest of the community record.
| work | tier | score | rank | n | success | quality | reliability | steering | $ / success | best in this work |
|---|
n is this model’s sessions on that work / everyone’s. Rank is among models with enough sessions on that work; “needs 10” means this model has too few there to place.
The human effort behind the result.
Loading.
Explore all conversation signals
| n signals | corrections | re-prompts | frustration | pushback | clarifications | edits without read | abandoned | tool-call error | rate-limited | overloaded | interrupted | switched away |
|---|
Signals describe patterns, not intent: a clarifying question can be useful, and frustration markers vary by person. Phrase detection is English-only. Rate limits, overloads, interruptions and model switches are shares of sessions. No conversation text is shared.
Same model. Different host.
Loading.
| model | provider | quant | serving mode | n | score | quality | reliability | tool-call error | p50 latency | $ / success |
|---|
The tools and settings behind it.
Loading.