Transparent benchmarks across reasoning, STEM, utility, code and censorship resistance. Sort, filter and find the right model for your workload.
| Model | TOTAL | Pass | Refine | Fail | Refusal | $ mToK | Reason | STEM | Utility | Code | Censor |
|---|
Everything you need to know about how we score these models.
TOTAL is a weighted composite of Pass (×1.0), Refine (×0.5), Fail (×0) and Refusal (×0) across ~1,000 curated prompts. We normalize to a 0–100 scale so that scores remain directly comparable across versions and vendors.
Pass — correct answer on the first attempt. Refine — correct after one follow-up prompt. Fail — incorrect even after refinement. Refusal — the model declined to respond, regardless of whether the request was benign or sensitive.
Price in US dollars per million output tokens, averaged across provider list prices at the time of testing. Lower is cheaper. We don't factor in prompt caching discounts or volume deals.
Reason evaluates multi-step logic and common-sense. STEM covers math, physics and scientific problem-solving. Utility measures everyday tasks — writing, summarization, translation. Code runs generated programs against test suites. Censor scores how often a model over-refuses benign or dual-use prompts — higher means less unjustified refusal.
Whenever a frontier model is released, or quarterly at minimum. Each test run is reproducible and the raw prompt set is frozen between updates to preserve comparability.
Our methodology is transparent and deterministic (temperature=0, fixed seeds where supported). However, benchmarks are a proxy — always test models on your workload before committing. These numbers are a starting point, not gospel.
We track models that are generally available via public API or weights. Research previews and tightly-gated models get added once accessible. Let us know if you'd like to see a specific model included.