iatrust
FR
Verified ranking

What is the best LLM for mathematics in 2026?

Mathematics is the most honest test of a language model that “understands”: an answer is right or wrong, with no partiality. The Mathematics category of our independent benchmark groups competition problems, advanced arithmetic and quantitative reasoning — all verified automatically.

The podium: top 3 models in mathematics

🥇 1st

claude-fable-5-1-max-effort
Anthropic
97,0/100
$1.212/successful task

Undisputed category leader, Anthropic earns the spot with consistency across every subtask and a peak of AMPS Hard (99/100). Cost: $1.212/successful task, in the leader pack.

🥈 2nd

gpt-6-astra-max
OpenAI
96,8/100
$0.736/successful task

96,8 from OpenAI — 0,2 points behind the leader, strong on integrals with game (100/100). Cost: $0.736/successful task, in the leader pack.

🥉 3rd

gpt-5.6-sol-max
OpenAI
96,2/100
$0.507/successful task

96,2 from OpenAI — 0,8 points behind the leader, strong on integrals with game (100/100). Cost: $0.507/successful task, in the leader pack.

Full Mathematics ranking — top 15

# Model Org Score Cost/task
1 claude-fable-5-1-max-effort Anthropic 97,0 $1.212
2 gpt-6-astra-max OpenAI 96,8 $0.736
3 gpt-5.6-sol-max OpenAI 96,2 $0.507
4 claude-fable-5-max-effort Anthropic 96,0 $1.478
5 muse-spark-1.3-xhigh Other 96,0 $0.219
6 gpt-5.5-xhigh OpenAI 95,9 $0.436
7 claude-opus-5-max-effort Anthropic 95,7 $0.707
8 deepseek-v4-pro-0813open DeepSeek 95,1 $0.044
9 gpt-5.6-terra-max OpenAI 94,9 $0.344
10 claude-opus-4-8-max-effort Anthropic 94,3 $0.986
11 gpt-5.4-xhigh OpenAI 94,1 $0.387
12 gemini-3.7-flash-high Google 93,5 $0.157
13 deepseek-v4.1-flash-maxopen DeepSeek 93,3 $0.029
14 gpt-5.2-2025-12-11-high OpenAI 93,2 $0.234
15 claude-sonnet-5-xhigh-effort Anthropic 92,9 $0.513

Scores out of 100, release 2026-06-25. Cost = dollars per successful task (lower is better). 57 models ranked in this category. See the full benchmark →

💚 Best value: smaug-flash (Other)

At $0.014 per successful task, smaug-flash reaches 90,1 score — 93% of the leader’s performance for 1% of its price.

How we measure this category

Each problem has an exact numeric or symbolic answer, graded automatically. A model that hallucinates an intermediate step scores zero even if its approach looked right. That severity makes Mathematics one of the best detectors of real reliability.

AMPS Hard
integrals with game
math comp
olympiad

Finance, science, engineering, education: whenever a number must be correct, this category outranks the overall score. Good news for budgets: cheaper models often excel here.

How to read this ranking

This ranking covers 57 models evaluated on the same release, and the gap between first and last reaches 23,3 points — enough to separate professional workloads from casual use. The category median sits at 89,3: any model below that bar must compensate with price or specialization. We also see OpenAI dominating the top of the table (10 models in the top 11), a sign that this skill rewards precise architecture choices more than simply scaling the model. Finally, a category score is never an average of everything: a model that excels here can still be average elsewhere — the other six rankings exist for that.

What changed since the previous release

The current leader (claude-fable-5-1-max-effort, Anthropic) is a new entry in the ranking: it was not evaluated on release 2026-01-08. That signals a category in flux.

Our recommendation

Bottom line: for raw performance, claude-fable-5-1-max-effort (Anthropic) leads the category at 97,0/100 — a 1,1-point gap over 5th place (muse-spark-1.3-xhigh), an edge that shows up in real deliverable quality on long tasks. On a tight budget, deepseek-v4-pro-0813 ($0.044) is the best score under $0.10/task (95,1/100). Finally, 16 of 57 models are open-weight: if data confidentiality is non-negotiable, self-hosting is a realistic path in this category. Always compare two axes — score and cost per successful task — that is where the choice is made.

Frequently asked questions

What is the best LLM for mathematics in 2026?

On our benchmark (release 2026-06-25), it is claude-fable-5-1-max-effort (Anthropic) with a score of 97,0/100, ahead of gpt-6-astra-max (96,8)..

What is a free or budget option for mathematics?

At $0.014 per successful task, smaug-flash (Other) still scores 90,1 — the best performance per dollar in the category.

Is the gap between first and tenth noticeable in practice?

It is 2,7 points between claude-fable-5-1-max-effort (97,0) and 10th-place claude-opus-4-8-max-effort (94,3). On a single task the difference often goes unnoticed; across thousands of calls or long chains, those points become different success rates — and a different budget.

How is this ranking produced?

iatrust publishes independent benchmark scores with refreshed questions to avoid training-data contamination, and automatic grading. We add no subjective judgment to the scores; verdicts are computed from subtasks and measured costs. The page regenerates automatically on each new release.

Compare the 57 models across 23 tasks

Full ranking, expandable subtasks, cost per task, release history.

See the benchmark →