Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

Humanity's Last Exam (tools) leaderboard

Humanity's Last Exam is a multidisciplinary set of expert-level questions built to resist the saturation that overtook MMLU. This column reports the tool-augmented mode, where the model may search and run tools while answering. That distinction is the whole story: the official maintainer leaderboards are closed-book by construction, evaluated text-only at temperature zero with searchable questions deliberately removed from the set so retrieval cannot substitute for knowledge. There is no separate with-tools dataset, paper, repository, or maintainer. It is a reporting mode, not a distinct benchmark, and the figures come from the labs themselves.

Current leader
Claude Opus 5(Anthropic)64.7%

Last refreshed 2026-08-26. 9 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1Claude Opus 5Anthropic64.7%2026-07
2Claude Fable 5Anthropic63.9%2026-06
3GLM-5.3Z.ai62.5%vendor2026-08
4Claude Opus 4.8Anthropic57.9%2026-05
5Claude Sonnet 5Anthropic57.4%vendor2026-06
6Qwen3.8-MaxAlibaba56.2%vendor2026-08
7Kimi K3Moonshot AI56%vendor2026-07
8GLM-5.2Z.ai54.7%vendor2026-06
9InklingThinking Machines46%vendor2026-07

Score interpretation

Never compare a with-tools score against a closed-book one. The gap is roughly ten points at the frontier, and it measures the tools, not the model. Because the HLE dataset is public, granting a model search access also lets it retrieve discussion of the questions themselves, which is exactly why the maintainers keep the official board closed-book. Treat every number in this column as vendor-reported and contamination-exposed, and treat the closed-book board as the cleaner measurement of what a model actually knows.

60%+
Frontier with tools. Expert-level answers across most disciplines.
45-60%
Strong. Handles hard questions when retrieval cooperates.
25-45%
Mixed. Reliable on mainstream topics, weak at the edges.
< 25%
Below the useful bar for expert work, even with tools.

Why this matters for AI agents

Tool-augmented reasoning is how agents actually run in production, so this mode is closer to real deployment than the closed-book board. It is simply not a knowledge measurement. Read this column as an upper bound on what a model plus its retrieval stack can do on expert questions, and read closed-book HLE when you want to know what the weights themselves contain.

Other benchmarks

Premium API: time-series for Humanity's Last Exam (tools)

The leaderboard above is a snapshot. Want to see how a model's Humanity's Last Exam (tools) score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

Humanity's Last Exam (tools) source ·Last refreshed 2026-08-26·Max score 100