Board › Models

THE MODEL LIBRARY

Get to know the models.

Community results, task by task. Find a model, follow its history, and see the evidence behind its reputation.

Loading the models on record.

Human-directed sessions. Automated runs are excluded.
How to read this

Score blends rating, survival and clean sessions; missing parts are reweighted. A score without ratings is not a quality verdict. n counts sessions, not ratings. Public tiers need 10 sessions; scores need 3.

A band spans the middle half of observations. API-equivalent estimates list-price token cost, not your bill. A multiple divides that value by period plan cost. A reporter-week is one person in one week, not a unique person across weeks. An assumed plan applies today’s plan to older sessions without a recorded plan. Friction means interruptions and extra direction; a wall hit means a usage limit stopped work.

English-language signals are heuristics, sensitive to task, tool and reporter habits. Small samples describe this record, not all users. Select any ? for a definition or a missing value for its reason.

01

Performance at a glance.

Six ways to judge the work. Each tier compares models with ten or more sessions and an available measurement. Steering, speed and value are separate from the overall score.

02

Follow the changes.

Loading.

Track this model alongside the whole field. Inspect a week for its score and sample size. Changes in tools, effort and task mix can also move the results.

Explore the weekly evidence
weeknscorefield scorequalityreliabilitysteeringsurvivalp50 latencyerrorsflagmovedtools · effort

Newest week first. A move is a metric that shifted |z| ≥ 2 from the trailing baseline; a flag needs five sessions that week. Tool and effort are listed because a harness change looks exactly like a weights change from here.

03

Find its kind of work.

Loading.

Results vary by task. Select a kind of work to compare this model with the rest of the community record.

worktierscoreranknsuccessqualityreliabilitysteering$ / successbest in this work

n is this model’s sessions on that work / everyone’s. Rank is among models with enough sessions on that work; “needs 10” means this model has too few there to place.

04

The human effort behind the result.

Loading.

Explore all conversation signals
n signalscorrectionsre-promptsfrustrationpushbackclarificationsedits without readabandonedtool-call errorrate-limitedoverloadedinterruptedswitched away

Signals describe patterns, not intent: a clarifying question can be useful, and frustration markers vary by person. Phrase detection is English-only. Rate limits, overloads, interruptions and model switches are shares of sessions. No conversation text is shared.

05

Same model. Different host.

Loading.

modelproviderquantserving modenscorequalityreliabilitytool-call errorp50 latency$ / success
06

The tools and settings behind it.

Loading.

Explore results by tool, plan, language and effort

Use these models? Add your outcomes.Good sessions and bad ones both make the community record more useful.

Contribute with nerfd ↗