Humanity's Last Exam (tools) leaderboard
Humanity's Last Exam is a multidisciplinary set of expert-level questions built to resist the saturation that overtook MMLU. This column reports the tool-augmented mode, where the model may search and run tools while answering. That distinction is the whole story: the official maintainer leaderboards are closed-book by construction, evaluated text-only at temperature zero with searchable questions deliberately removed from the set so retrieval cannot substitute for knowledge. There is no separate with-tools dataset, paper, repository, or maintainer. It is a reporting mode, not a distinct benchmark, and the figures come from the labs themselves.
Last refreshed 2026-08-26. 9 models scored on this benchmark.
Full leaderboard
| # | Model | Provider | Score | Released |
|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 64.7% | 2026-07 |
| 2 | Claude Fable 5 | Anthropic | 63.9% | 2026-06 |
| 3 | GLM-5.3 | Z.ai | 62.5%vendor | 2026-08 |
| 4 | Claude Opus 4.8 | Anthropic | 57.9% | 2026-05 |
| 5 | Claude Sonnet 5 | Anthropic | 57.4%vendor | 2026-06 |
| 6 | Qwen3.8-Max | Alibaba | 56.2%vendor | 2026-08 |
| 7 | Kimi K3 | Moonshot AI | 56%vendor | 2026-07 |
| 8 | GLM-5.2 | Z.ai | 54.7%vendor | 2026-06 |
| 9 | Inkling | Thinking Machines | 46%vendor | 2026-07 |
Score interpretation
Never compare a with-tools score against a closed-book one. The gap is roughly ten points at the frontier, and it measures the tools, not the model. Because the HLE dataset is public, granting a model search access also lets it retrieve discussion of the questions themselves, which is exactly why the maintainers keep the official board closed-book. Treat every number in this column as vendor-reported and contamination-exposed, and treat the closed-book board as the cleaner measurement of what a model actually knows.
Why this matters for AI agents
Tool-augmented reasoning is how agents actually run in production, so this mode is closer to real deployment than the closed-book board. It is simply not a knowledge measurement. Read this column as an upper bound on what a model plus its retrieval stack can do on expert questions, and read closed-book HLE when you want to know what the weights themselves contain.
Other benchmarks
Premium API: time-series for Humanity's Last Exam (tools)
The leaderboard above is a snapshot. Want to see how a model's Humanity's Last Exam (tools) score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:
/api/premium/history/benchmarks/series?model=&benchmark=hle_tools: daily score evolution for one model on this benchmark, 1 credit per call