{"ok":true,"source":"tensorfeed.ai","lastUpdated":"2026-09-14","count":26,"leaderboards":[{"id":"lmsys-arena","name":"Arena (formerly LMArena, originally LMSYS Chatbot Arena)","publisher":"Arena Intelligence","scope":"General chat capability across all major models, plus separate arenas for agents, web dev, vision, documents, search, image, and video","updateCadence":"continuous (live)","scoreType":"Elo-style rating (human pairwise vote)","domain":"general","live":true,"hasAPI":false,"url":"https://arena.ai/leaderboard/text","notes":"The most-cited crowd-sourced model ranking. The lmarena.ai domain now redirects to arena.ai. The Text arena alone passed 8.1M human pairwise votes by September 2026, with separate boards for Agent, WebDev, Vision, Document, Search, Text-to-Image, Image Edit, and video tasks."},{"id":"artificial-analysis","name":"Artificial Analysis","publisher":"Artificial Analysis","scope":"Quality + price + latency across models and providers","updateCadence":"daily","scoreType":"Composite quality index","domain":"general","live":true,"hasAPI":true,"url":"https://artificialanalysis.ai","notes":"Independent benchmark and market-data aggregator that runs evaluations itself. Its Intelligence Index (v4.3 as of September 2026) folds in agentic evals such as GDPval-AA v2, AutomationBench, and Terminal-Bench 4.0 alongside latency and price. Free data API with a key."},{"id":"hf-open-llm","name":"Open LLM Leaderboard","publisher":"Hugging Face","scope":"Open-weights models across 6 academic benchmarks (IFEval, BBH, MATH, GPQA, MUSR, MMLU-Pro)","updateCadence":"archived (no new evaluations)","scoreType":"Average of normalized benchmarks","domain":"open-models","live":false,"hasAPI":true,"url":"https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard","notes":"Archived. The Space still loads with an archived badge and the results datasets remain downloadable, but its last content update was March 2025 and later commits only fix startup crashes. Useful as a historical snapshot of 2024 era open models, not for current rankings."},{"id":"swebench-verified","name":"SWE-bench Verified Leaderboard","publisher":"SWE-bench team","scope":"Coding agents on 500 real GitHub issues","updateCadence":"as submissions land","scoreType":"% issues resolved","domain":"code","live":true,"hasAPI":false,"url":"https://www.swebench.com","notes":"The official board still tops out at 79.2% (Claude 4.5 Opus agent scaffolds, tied) and has taken no new Verified entry since February 2026; in September 2026 the separate bash-only board became a toggle inside Verified. Vendor launches and independent evaluators now report low to mid 90s on their own harnesses, so check the harness before comparing. TensorFeed mirrors top entries at /harnesses."},{"id":"aider-leaderboard","name":"Aider Polyglot Leaderboard","publisher":"Aider","scope":"225 hardest Exercism exercises across 6 languages","updateCadence":"dormant (page last updated November 2025)","scoreType":"% pass@2","domain":"code","live":false,"hasAPI":false,"url":"https://aider.chat/docs/leaderboards/","notes":"Edit-by-diff coding benchmark and a strong cross-language test, maintained by the Aider community. Frozen with gpt-5 (high) on top at 88.0%; no 2026 models have been added."},{"id":"livecodebench","name":"LiveCodeBench","publisher":"UC Berkeley + UW","scope":"Competitive programming problems released after model training cutoff","updateCadence":"no longer updating (last repo commit August 2025)","scoreType":"% pass@1","domain":"code","live":false,"hasAPI":false,"url":"https://livecodebench.github.io/leaderboard.html","notes":"Built to be contamination-free by pulling problems from after each model training cutoff, but the problem window stopped in spring 2025, so that guarantee no longer holds for current models. Vendors still quote it in launch tables; treat those figures as self-reported."},{"id":"bigcodebench","name":"BigCodeBench","publisher":"BigCode","scope":"1140 hard programming tasks with library usage","updateCadence":"not maintained (last results April 2025)","scoreType":"% pass@1","domain":"code","live":false,"hasAPI":false,"url":"https://huggingface.co/spaces/bigcode/bigcodebench-leaderboard","notes":"Multi-library code benchmark that tests realistic code composing external libraries; harder than HumanEval or MBPP for the same models. The board has not added results since April 2025."},{"id":"terminal-bench","name":"Terminal-Bench Leaderboard","publisher":"Terminal-Bench Team (Stanford, Laude Institute; Harbor harness)","scope":"Model plus agent-harness pairs on containerized terminal tasks with deterministic verifiers","updateCadence":"as runs land; versioned releases (4.0 current since August 2026)","scoreType":"% resolution rate with 95% interval","domain":"agent","live":true,"hasAPI":false,"url":"https://www.tbench.ai","notes":"Tests the loop, not just the model. The site now defaults to Terminal-Bench 4.0 (66 tasks) and also hosts 1.0, 2.0, 2.1, 3.0, and Science 0.1; versions are not comparable. On 4.0, Codex with GPT-6 Astra (max) leads at 58.2% as of September 2026, with Claude Code plus Claude Fable 5.1 at 57.9% inside the interval."},{"id":"frontiercode","name":"FrontierCode Leaderboard","publisher":"Cognition","scope":"Coding agents on maintainer-authored open-source tasks, graded for mergeability (Main 100 tasks, Extended 150)","updateCadence":"as models release (changelog on page)","scoreType":"Composite score (tests plus maintainer rubric), pass rate shown separately","domain":"code","live":true,"hasAPI":true,"url":"https://cognition.com/frontiercode","notes":"Grades whether a patch would actually be merged, not just whether tests pass; runs that consult solution-bearing sources are zeroed. Version 1.1 deprecated the Diamond subset. Per-effort results are published as a public JSON file behind the page. Top three on 1.1 Main sit within 0.2 points as of September 2026."},{"id":"deepswe","name":"DeepSWE Leaderboard","publisher":"Datacurve","scope":"113 original long-horizon engineering tasks across 91 repos and 5 languages","updateCadence":"as models release","scoreType":"Pass@1 with error bars, plus cost, output tokens, and steps","domain":"code","live":true,"hasAPI":false,"url":"https://deepswe.datacurve.ai","notes":"Every model runs on mini-swe-agent for consistency, so it compares models rather than harnesses. Tasks are written from scratch to avoid contamination. A three-way tie at 74% (GPT-6 Astra, Gemini 3.8 Flash, Claude Opus 5) leads v1.1 as of September 2026."},{"id":"arc-prize","name":"ARC Prize Leaderboard","publisher":"ARC Prize Foundation","scope":"Abstract reasoning: ARC-AGI-1, ARC-AGI-2, and the interactive ARC-AGI-3 environments","updateCadence":"as verified runs land","scoreType":"% solved, plotted against cost per task","domain":"reasoning","live":true,"hasAPI":true,"url":"https://arcprize.org/leaderboard","notes":"Opens on ARC-AGI-3 and splits Verified from Community results; board data is published as JSON. As of September 2026, GPT-6 Astra (max) scores 95.0% on ARC-AGI-2 against a 100% human panel, and 62.7% on ARC-AGI-3 with the standard harness (99.9% with the OpenAI provider adapter harness). ARC-AGI-2 is now near saturation."},{"id":"mmlu-pro","name":"MMLU-Pro Leaderboard","publisher":"TIGER Lab","scope":"General knowledge + reasoning across 57 subjects","updateCadence":"occasional (results dataset last updated March 2026)","scoreType":"% accuracy","domain":"general","live":true,"hasAPI":false,"url":"https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro","notes":"Successor to MMLU and the standard knowledge plus reasoning check of 2024 to 2026. The maintainer board updates slowly, so September 2026 flagships mostly appear only on independent evaluators such as Vals AI."},{"id":"hle-leaderboard","name":"Humanity's Last Exam","publisher":"CAIS + Scale AI","scope":"2,500 expert-validated questions across 100+ disciplines (finalized April 2025)","updateCadence":"lastexam.ai table is static; the live board is at Scale Labs","scoreType":"% accuracy","domain":"reasoning","live":true,"hasAPI":false,"url":"https://labs.scale.com/leaderboard/humanitys_last_exam","notes":"Every major lab reports against HLE, with and without tools. The original lastexam.ai page still shows a 2025 table; the maintained ranking lives on Scale Labs, where GPT-6 Astra leads at 54.8% and the next entries sit in the mid 40s as of September 2026. Tool setups vary by vendor, so compare like for like."},{"id":"mmmu","name":"MMMU Leaderboard","publisher":"TIGER Lab","scope":"College-level multimodal questions across 30 subjects","updateCadence":"as submissions land (last entry July 2026)","scoreType":"% accuracy","domain":"multimodal","live":true,"hasAPI":false,"url":"https://mmmu-benchmark.github.io","notes":"Standard vision-language benchmark with MMMU (val) and MMMU-Pro tables. The board aggregates submitted results rather than re-running models. Top MMMU-Pro entry is 86.9 and top MMMU val is 85.4 as of September 2026."},{"id":"video-arena","name":"Artificial Analysis Video Arena","publisher":"Artificial Analysis","scope":"Video generation models, ranked separately with and without audio output","updateCadence":"continuous","scoreType":"Elo (human pairwise vote)","domain":"video","live":true,"hasAPI":false,"url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","notes":"Live Elo for video models. The old /text-to-video/arena path now opens the voting page, so link the leaderboard path instead. Wan 3.0 leads the with-audio text-to-video board at 1242 as of September 2026; HappyHorse 1.0 no longer leads."},{"id":"image-arena","name":"Artificial Analysis Image Arena","publisher":"Artificial Analysis","scope":"Image generation and image editing models","updateCadence":"continuous","scoreType":"Elo (human pairwise vote)","domain":"image","live":true,"hasAPI":false,"url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","notes":"Live Elo for image models. The old /text-to-image/arena path now opens the voting page. GPT Image 2.5 Flare (max) leads text-to-image at 1187 as of September 2026."},{"id":"tts-arena","name":"TTS Arena V2","publisher":"TTS-AGI (Hugging Face Space)","scope":"TTS models from commercial and open providers, including stealth entries","updateCadence":"continuous","scoreType":"Elo (human pairwise vote)","domain":"voice","live":true,"hasAPI":true,"url":"https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2","notes":"Live Elo for TTS quality with a public JSON leaderboard endpoint. The original TTS Arena Space is now a read-only legacy board. A stealth model leads V2 as of September 2026 and Eleven v3 sits around 27th. TensorFeed mirrors top entries at /voice-leaderboards."},{"id":"open-asr","name":"Open ASR Leaderboard","publisher":"Hugging Face","scope":"STT models on LibriSpeech + Common Voice + AMI + GigaSpeech","updateCadence":"as submissions land","scoreType":"Word Error Rate","domain":"voice","live":true,"hasAPI":true,"url":"https://huggingface.co/spaces/hf-audio/open_asr_leaderboard","notes":"Aggregated ASR benchmark with English and multilingual tracks, actively updated through September 2026. The English average WER leader has turned over repeatedly since 2025, and AssemblyAI Universal-2 no longer leads."},{"id":"ruler-leaderboard","name":"RULER Leaderboard","publisher":"NVIDIA","scope":"Long-context retrieval and reasoning (effective vs claimed length)","updateCadence":"README table frozen since October 2025 (repo still receives fixes)","scoreType":"% accuracy at varying context lengths","domain":"long-context","live":false,"hasAPI":false,"url":"https://github.com/NVIDIA/RULER","notes":"Reveals the gap between claimed and effective context length. The README leaderboard has not been updated since RULER v2 was added in October 2025, so it predates the 1M-token context flagships of 2026."},{"id":"gaia-leaderboard","name":"GAIA Leaderboard","publisher":"Hugging Face + Meta","scope":"General assistant agents on 466 real-world questions","updateCadence":"as submissions land","scoreType":"% accuracy by difficulty level","domain":"agent","live":true,"hasAPI":false,"url":"https://huggingface.co/spaces/gaia-benchmark/leaderboard","notes":"Tests the full agent loop: web browsing, file ops, multi-step reasoning, at three difficulty levels. Largely saturated: the test-split leader, an agent scaffold, scores 93.36% as of September 2026."},{"id":"webarena","name":"WebArena Leaderboard","publisher":"CMU + UW","scope":"Browser agents on 812 tasks across 5 simulated websites","updateCadence":"as submissions land (maintained as a linked spreadsheet)","scoreType":"% solved","domain":"agent","live":true,"hasAPI":false,"url":"https://webarena.dev/og/","notes":"The browser-automation agent benchmark. The old /leaderboard path now returns 404 and webarena.dev is a WebArena-x hub; the original board is a spreadsheet linked from the /og/ page. Frontier is 74.3% (WebTactix with DeepSeek v3.2) as of September 2026."},{"id":"osworld-leaderboard","name":"OSWorld-Verified Leaderboard (OSWorld 1.0)","publisher":"XLANG Lab, University of Hong Kong","scope":"Computer-use agents on 369 desktop tasks (361 excluding Google Drive tasks)","updateCadence":"as verified submissions land (newest entry August 2026)","scoreType":"% solved","domain":"agent","live":true,"hasAPI":false,"url":"https://osworld-v1.xlang.ai","notes":"The original computer-use benchmark, now near saturation: the Intelligence-Indeed Agent scaffold leads OSWorld-Verified at 90.19% and the best bare general model sits in the mid 80s. os-world.github.io redirects here. For long-horizon computer use, see the separate OSWorld 2.0 board."},{"id":"osworld-2","name":"OSWorld 2.0 Leaderboard","publisher":"XLANG Lab, University of Hong Kong","scope":"Computer-use agents on 108 long-horizon workflows across 7 professional domains","updateCadence":"as official runs land","scoreType":"Binary completion and partial score at a fixed step budget","domain":"agent","live":true,"hasAPI":true,"url":"https://osworld-v2.xlang.ai","notes":"Filter by task release (v2026.06.24 or v2026.08.08), full set or offline no-internet subset, and step budget before comparing anything; vendor launch figures often mix these. Official results are served as a public JSON file. Claude Opus 5 (max) leads the v2026.08.08 full set at 31.43% binary as of September 2026."},{"id":"automationbench","name":"AutomationBench Leaderboard","publisher":"Zapier","scope":"Agents completing cross-app business workflows across 47 simulated SaaS tools in 6 domains","updateCadence":"as models release (versioned, currently 1.0.6)","scoreType":"% tasks completed correctly (strict end-state assertions)","domain":"agent","live":true,"hasAPI":false,"url":"https://zapier.com/benchmarks","notes":"Scored on a held-out private task set with deterministic checks and no LLM judge; the public GitHub task set is easier. Lists cost per task and per-domain leaders. Watch for fallback rows, such as Claude Fable 5.1 with Claude Opus 5 completing refused steps. GPT-6 Astra (max) leads at 41.4% as of September 2026."},{"id":"steel-browsecomp","name":"BrowseComp Leaderboard","publisher":"Steel","scope":"Agentic web research systems on OpenAI's BrowseComp","updateCadence":"as results are published","scoreType":"% accuracy","domain":"agent","live":true,"hasAPI":false,"url":"https://leaderboard.steel.dev/leaderboards/browsecomp/","notes":"BrowseComp has no official leaderboard, so this board collects self-reported figures with a source link and setup notes per row (single agent vs multi-agent, context management). Treat it as a registry of claims, not a controlled comparison."},{"id":"vals-ai","name":"Vals AI Benchmarks","publisher":"Vals AI","scope":"Independent runs of frontier models across academic, legal, finance, healthcare, and coding benchmarks, plus the composite Vals Index","updateCadence":"weekly to monthly per benchmark (dated on each board)","scoreType":"Accuracy per benchmark; Vals Index composite","domain":"general","live":true,"hasAPI":false,"url":"https://www.vals.ai/benchmarks","notes":"Runs every model itself on a fixed harness (SWE-bench Verified uses a bash-only mini-swe-agent), so its numbers often differ from vendor launch claims. Each board shows its own last-updated date and model count."}]}