Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

Humanity's Last Exam (tools) leaderboard

Humanity's Last Exam is a multidisciplinary set of expert-level questions built to resist the saturation that overtook MMLU. This column reports the tool-augmented mode, where the model may search and run tools while answering. That distinction is the whole story: the official maintainer leaderboards are closed-book by construction, evaluated text-only at temperature zero with searchable questions deliberately removed from the set so retrieval cannot substitute for knowledge. There is no separate with-tools dataset, paper, repository, or maintainer. It is a reporting mode, not a distinct benchmark, and the figures come from the labs themselves.

Current leader
Claude Fable 5.1(Anthropic)65%

Last refreshed 2026-09-05. 12 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1Claude Fable 5.1Anthropic65%vendor2026-09
2Claude Opus 5Anthropic64.7%2026-07
3Claude Fable 5Anthropic63.9%2026-06
4GLM-5.3Z.ai62.5%vendor2026-08
5Claude Opus 4.8Anthropic57.9%2026-05
6Claude Sonnet 5Anthropic57.4%vendor2026-06
7Qwen3.8 2.4T-A95BAlibaba56.2%vendor2026-08
8Qwen3.8-MaxAlibaba56.2%vendor2026-08
9Kimi K3Moonshot AI56%vendor2026-07
10GLM-5.3 FlashZ.ai55.3%vendor2026-08
11GLM-5.2Z.ai54.7%vendor2026-06
12InklingThinking Machines46%vendor2026-07

Score interpretation

Never compare a with-tools score against a closed-book one. The gap is roughly ten points at the frontier, and it measures the tools, not the model. Because the HLE dataset is public, granting a model search access also lets it retrieve discussion of the questions themselves, which is exactly why the maintainers keep the official board closed-book. Treat every number in this column as vendor-reported and contamination-exposed, and treat the closed-book board as the cleaner measurement of what a model actually knows.

60%+
Frontier with tools. Expert-level answers across most disciplines.
45-60%
Strong. Handles hard questions when retrieval cooperates.
25-45%
Mixed. Reliable on mainstream topics, weak at the edges.
< 25%
Below the useful bar for expert work, even with tools.

Why this matters for AI agents

Tool-augmented reasoning is how agents actually run in production, so this mode is closer to real deployment than the closed-book board. It is simply not a knowledge measurement. Read this column as an upper bound on what a model plus its retrieval stack can do on expert questions, and read closed-book HLE when you want to know what the weights themselves contain.

Other benchmarks

Premium API: time-series for Humanity's Last Exam (tools)

The leaderboard above is a snapshot. Want to see how a model's Humanity's Last Exam (tools) score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

Humanity's Last Exam (tools) source ·Last refreshed 2026-09-05·Max score 100