Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

BrowseComp leaderboard

BrowseComp tests persistent web research. Its 1,266 questions have short, verifiable answers that are genuinely difficult to locate, requiring an agent to chase entangled information across many pages rather than retrieve a single fact. OpenAI built it and released it in 2025, and it has become the standard reference for deep-research and agentic-search products.

Current leader
Kimi K3(Moonshot AI)91.2%

Last refreshed 2026-08-26. 11 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1Kimi K3Moonshot AI91.2%vendor2026-07
2Claude Opus 5Anthropic90.8%vendor2026-07
3GPT-5.6 SolOpenAI90.4%vendor2026-07
4GPT-5.6 TerraOpenAI87.5%vendor2026-07
5Claude Fable 5Anthropic87.4%2026-06
6Claude Sonnet 5Anthropic84.7%vendor2026-06
7Claude Opus 4.8Anthropic84.3%2026-05
8MiniMax M3MiniMax83.5%vendor2026-06
9GPT-5.6 LunaOpenAI83.3%vendor2026-07
10LongCat-2.0Meituan79.9%vendor2026-06
11InklingThinking Machines77.1%vendor2026-07

Score interpretation

Two caveats that matter more here than on any other benchmark we track. First, there is no official leaderboard and no independent evaluator: every published BrowseComp figure is self-reported by the lab that produced it. Second, scores move sharply with context-management strategy, so a single-agent run, a multi-agent run, and a run using context compaction are not the same measurement even for the same model. What a BrowseComp number really ranks is a full system, the model plus its tools plus its context policy, not a bare model. Frontier systems now cluster within a couple of points of each other, which suggests the benchmark is approaching saturation at the top.

85%+
Frontier research agent. Finds answers most humans would give up on.
70-85%
Strong. Reliable on hard lookups, occasional dead ends.
40-70%
Useful for ordinary search, weak on genuinely buried facts.
< 40%
Not a research agent. Expect confident wrong answers.

Why this matters for AI agents

If you are building research agents, this is the closest published proxy for whether the thing will actually find an obscure answer instead of confidently inventing one. Just do not treat small gaps between top models as real. The measurement noise from differing harnesses and context strategies is larger than the differences between the leaders.

Other benchmarks

Premium API: time-series for BrowseComp

The leaderboard above is a snapshot. Want to see how a model's BrowseComp score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

BrowseComp source ·Last refreshed 2026-08-26·Max score 100