nerfd.ai

BoardModels

Models on record

Loading the models on record.

Steered sessions only; automated runs are excluded from rankings.
How to read this

Score blends rating, survival and clean sessions; missing parts are reweighted. A score without ratings is not a quality verdict. n counts sessions, not ratings. Public tiers need 10 sessions; scores need 3.

A band spans the middle half of observations. API-equivalent estimates list-price token cost, not your bill. A multiple divides that value by period plan cost. A reporter-week is one person in one week, not a unique person across weeks. An assumed plan applies today’s plan to older sessions without a recorded plan. Friction means interruptions and extra direction; a wall hit means a usage limit stopped work.

English-language signals are heuristics, sensitive to task, tool and reporter habits. Small samples describe this record, not all users. Select any ? for a definition or a missing value for its reason.

01

Where it ranks

Each criterion is banded relative to the best model in the field with ten or more sessions in the window. Rank counts only models that have a value for that criterion.

02

Against what work

Loading.

The same tiering, run inside each kind of work: a model can be S at debugging and C at UI. Click a heading to sort; click a kind of work to filter the board to it.

worktierscoreranknsuccessqualityreliabilitysteering$ / successbest in this work

n is this model’s sessions on that work / everyone’s. Rank is among models with enough sessions on that work; “needs 10” means this model has too few there to place.

03

Over time

Loading.

Weekly composite score since the model was first seen, with the field beside it and the release date where one is on file. Each week is tested against its trailing four weeks: watch at |z| ≥ 2, alert at ≥ 3, five sessions minimum.

weeknscorefield scorequalityreliabilitysteeringsurvivalp50 latencyerrorsflagmovedtools · effort

Newest week first. A move is a metric that shifted |z| ≥ 2 from the trailing baseline; a flag needs five sessions that week. Tool and effort are listed because a harness change looks exactly like a weights change from here.

04

How hard people had to push

Loading.

n signalscorrectionsre-promptsfrustrationpushbackclarificationsedits without readabandonedtool-call errorrate-limitedoverloadedinterruptedswitched away

Lower is better throughout. Signals are counts computed on the reporter’s machine from the conversation; no text is sent. Rate-limited, overloaded, interrupted and switched are shares of sessions.

05

Same weights, different host

Loading.

modelproviderquantserving modenscorequalityreliabilitytool-call errorp50 latency$ / success
06

Where it ran

Loading.