Updated April 2026

Compare every major language model
in one interactive table

Transparent benchmarks across reasoning, STEM, utility, code and censorship resistance. Sort, filter and find the right model for your workload.

Models tested
Avg. TOTAL score
Top scorer
Best value
Pass — task solved correctly
Refine — solved after iteration
Fail — incorrect output
Refusal — declined to respond
Click any column header to sort
Model TOTAL Pass Refine Fail Refusal $ mToK Reason STEM Utility Code Censor

Frequently Asked Questions

Everything you need to know about how we score these models.

How is the TOTAL score calculated?

TOTAL is a weighted composite of Pass (×1.0), Refine (×0.5), Fail (×0) and Refusal (×0) across ~1,000 curated prompts. We normalize to a 0–100 scale so that scores remain directly comparable across versions and vendors.

What do Pass, Refine, Fail, and Refusal mean?

Pass — correct answer on the first attempt. Refine — correct after one follow-up prompt. Fail — incorrect even after refinement. Refusal — the model declined to respond, regardless of whether the request was benign or sensitive.

What does "$ mToK" mean?

Price in US dollars per million output tokens, averaged across provider list prices at the time of testing. Lower is cheaper. We don't factor in prompt caching discounts or volume deals.

What's the difference between Reason, STEM, Utility, Code and Censor?

Reason evaluates multi-step logic and common-sense. STEM covers math, physics and scientific problem-solving. Utility measures everyday tasks — writing, summarization, translation. Code runs generated programs against test suites. Censor scores how often a model over-refuses benign or dual-use prompts — higher means less unjustified refusal.

How often is the benchmark updated?

Whenever a frontier model is released, or quarterly at minimum. Each test run is reproducible and the raw prompt set is frozen between updates to preserve comparability.

Can I trust these numbers?

Our methodology is transparent and deterministic (temperature=0, fixed seeds where supported). However, benchmarks are a proxy — always test models on your workload before committing. These numbers are a starting point, not gospel.

Why isn't model X listed?

We track models that are generally available via public API or weights. Research previews and tightly-gated models get added once accessible. Let us know if you'd like to see a specific model included.