iatrust
FR
Independent benchmark · release 2026-06-25

I evaluated 57 LLMs so you don’t have to.

Independent language-model benchmark: 23 tasks, 7 categories, real cost per successful task. No contamination, no partnerships.

57
models evaluated
83.4
Top overall score
claude-fable-5-1-max-effort
23
verified tasks
automatically
$0.0000
cheapest on
ox-alpha-max

AI ranking 2026: top 5 LLMs

Full ranking →

Models leading the board: coding, reasoning, math — public, reproducible scores updated each release.

1. claude-fable-5-1-max-effortAnthropic · #1 Mathematics
83.4
2. claude-fable-5-max-effortAnthropic · #1 Mathematics
83.0
3. gpt-6-astra-maxOpenAI · #1 Mathematics
82.2
4. muse-spark-1.3-xhighOther · #1 Mathematics
81.6
5. deepseek-v4.1-flash-maxDeepSeek · #1 Mathematics
81.1

Overall score for each model (average of 7 categories) — release 2026-06-25.

How to find the AI that fits your job

No need to read everything: every model is benchmarked, scored, filterable by metric — then our detailed rankings recommend the AI for your role.

01

Benchmark

Every listed AI is evaluated on the same 23 fresh, verifiable tasks

02

Scores

Overall grade + 7 categories: coding, reasoning, math, agents…

03

Filters

Isolate the metric that matters: performance, cost, open weights, org

04

Real cost

Dollars per successful task — the true price of your usage, not the sticker rate

05

Skill report

A dedicated ranking per skill with analysis, podium and FAQ

06

Recommendation

Our data-driven verdict: which AI to pick for your job

The reference

“A good LLM benchmark must ask fresh questions, uncontaminated by training data — automatic grading, regular updates.”

iatrustIndependent LLM benchmark

By the numbers

57
models ranked
23
tasks evaluated
11
historical releases
100%
independent & free