iatrust
FR
Verified ranking

What is the best LLM for following instructions in 2026?

Following a brief to the letter — right format, right constraints, right steps in order — is the invisible skill that makes or breaks a production pipeline. The Instruction Following category of our independent benchmark measures that discipline.

The podium: top 3 models in instruction following

🥇 1st

gemini-3.8-flash-high
Google
81,4/100
$0.307/successful task

Undisputed category leader, Google earns the spot with consistency across every subtask and a peak of summarize (88/100). Cost: $0.307/successful task, in the leader pack.

🥈 2nd

gemini-3.7-flash-high
Google
79,9/100
$0.157/successful task

79,9 from Google — 1,5 points behind the leader, strong on summarize (85/100). Cost: $0.157/successful task, in the leader pack.

🥉 3rd

gemini-3.1-pro-preview-high
Google
79,1/100
$0.285/successful task

79,1 from Google — 2,3 points behind the leader, strong on summarize (88/100). Cost: $0.285/successful task, in the leader pack.

Full Instruction Following ranking — top 15

# Model Org Score Cost/task
1 gemini-3.8-flash-high Google 81,4 $0.307
2 gemini-3.7-flash-high Google 79,9 $0.157
3 gemini-3.1-pro-preview-high Google 79,1 $0.285
4 muse-spark-1.3-xhigh Other 78,0 $0.219
5 qwen3.8-flash-nextopen Alibaba 77,1 $0.042
6 claude-fable-5-max-effort Anthropic 75,8 $1.478
7 gemini-3.5-flash-high Google 75,6 $0.249
8 gpt-6-astra-max OpenAI 75,6 $0.736
9 gemini-3.6-flash-high Google 75,4 $0.235
10 muse-spark-1.2-xhigh Other 74,3 $0.375
11 qwen3.8-maxopen Alibaba 74,1 $0.275
12 qwen3.7-maxopen Alibaba 74,0 $0.182
13 smaug-mini Other 73,9 $0.099
14 nemotron-3-ultra-550b-a55bopen NVIDIA 73,4 $0.371
15 claude-fable-5-1-max-effort Anthropic 73,0 $1.212

Scores out of 100, release 2026-06-25. Cost = dollars per successful task (lower is better). 57 models ranked in this category. See the full benchmark →

💚 Best value: qwen3.8-flash-next (Alibaba)

At $0.042 per successful task, qwen3.8-flash-next reaches 77,1 score — 95% of the leader’s performance for 14% of its price. And its open weights let you self-host.

How we measure this category

Tasks give compound instructions and automatically score each constraint: length, format, required or forbidden words, structure. Every violation counts.

paraphrase
simplify
story generation
summarize

If you generate JSON, forms, structured documents or chained prompts, make this your first criterion — even before the overall score.

How to read this ranking

This ranking covers 57 models evaluated on the same release, and the gap between first and last reaches 28,6 points — enough to separate professional workloads from casual use. The category median sits at 69,3: any model below that bar must compensate with price or specialization. We also see OpenAI dominating the top of the table (10 models in the top 11), a sign that this skill rewards precise architecture choices more than simply scaling the model. Finally, a category score is never an average of everything: a model that excels here can still be average elsewhere — the other six rankings exist for that.

What changed since the previous release

The current leader (gemini-3.8-flash-high, Google) is a new entry in the ranking: it was not evaluated on release 2026-01-08. That signals a category in flux.

Our recommendation

Bottom line: for raw performance, gemini-3.8-flash-high (Google) leads the category at 81,4/100 — a 4,3-point gap over 5th place (qwen3.8-flash-next), an edge that shows up in real deliverable quality on long tasks. On a tight budget, qwen3.8-flash-next ($0.042) is the best score under $0.10/task (77,1/100). Finally, 16 of 57 models are open-weight: if data confidentiality is non-negotiable, self-hosting is a realistic path in this category. Always compare two axes — score and cost per successful task — that is where the choice is made.

Frequently asked questions

What is the best LLM for instruction following in 2026?

On our benchmark (release 2026-06-25), it is gemini-3.8-flash-high (Google) with a score of 81,4/100, ahead of gemini-3.7-flash-high (79,9)..

What is a free or budget option for instruction following?

At $0.042 per successful task, qwen3.8-flash-next (Alibaba) still scores 77,1 — the best performance per dollar in the category, plus open weights you can self-host.

Is the gap between first and tenth noticeable in practice?

It is 7,1 points between gemini-3.8-flash-high (81,4) and 10th-place muse-spark-1.2-xhigh (74,3). On a single task the difference often goes unnoticed; across thousands of calls or long chains, those points become different success rates — and a different budget.

How is this ranking produced?

iatrust publishes independent benchmark scores with refreshed questions to avoid training-data contamination, and automatic grading. We add no subjective judgment to the scores; verdicts are computed from subtasks and measured costs. The page regenerates automatically on each new release.

Compare the 57 models across 23 tasks

Full ranking, expandable subtasks, cost per task, release history.

See the benchmark →