Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

OSWorld 2.0 leaderboard

OSWorld 2.0 measures whether a model can actually operate a computer. It presents 108 workflows across seven professional domains inside a real desktop environment, tasks that take a skilled human a median of 1.6 hours and average 318 tool calls to complete. Rather than a single pass or fail, each workflow is graded against many weighted checkpoints, so the benchmark can distinguish an agent that got most of the way from one that never started. It was built by XLANG Lab at the University of Hong Kong and released in 2026 as the successor to the original OSWorld, which frontier models had largely outgrown.

Current leader
Claude Opus 5(Anthropic)70.6%

Last refreshed 2026-08-26. 9 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1Claude Opus 5Anthropic70.6%vendor2026-07
2Claude Fable 5Anthropic66.1%2026-06
3GPT-5.6 SolOpenAI62.6%vendor2026-07
4Kimi K3Moonshot AI58.3%vendor2026-07
5Claude Opus 4.8Anthropic55.7%2026-05
6GPT-5.6 TerraOpenAI50.2%vendor2026-07
7Gemini 3.7 FlashGoogle47.9%vendor2026-08
8GPT-5.6 LunaOpenAI45.6%vendor2026-07
9Gemini 3.6 FlashGoogle33.8%vendor2026-07

Score interpretation

Read the metric before you compare anything. Binary completion requires every scoring checkpoint to pass, while the partial score averages the fraction of checkpoints satisfied, and those two numbers can differ by more than thirty points on the same run. Vendors almost always publish the partial figure while the benchmark authors headline binary completion. Step budget matters too: all current frontier submissions run at 500 steps, and the same model scores lower with a smaller budget. TensorFeed reports the vendor-published figure, so treat this column as one scale and do not mix it with numbers quoted elsewhere without checking which one you are reading.

20%+ binary
Current frontier. Completes a fifth of long professional workflows end to end.
10-20% binary
Capable on shorter workflows, unreliable across a full multi-hour task.
5-10% binary
Usually gets partway, rarely finishes. Needs a human watching.
< 5% binary
Not usable for autonomous desktop work.

Why this matters for AI agents

Every other benchmark on this site hands the model an API. This one hands it a mouse. If you are deploying an agent that clicks through software a human normally drives, spreadsheets, design tools, internal admin panels, this is the only column that tells you whether it will finish the job or stall halfway. The scores are low by design and the ceiling is nowhere close, which makes it the most honest measure of computer-use capability available right now.

Other benchmarks

Premium API: time-series for OSWorld 2.0

The leaderboard above is a snapshot. Want to see how a model's OSWorld 2.0 score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

OSWorld 2.0 source ·Last refreshed 2026-08-26·Max score 100