Fugu Max
Mid-tierby Sakana AI
Fugu Max v1.0 landed on September 11, 2026 as the cost-optimized half of Sakana's orchestration release, sharing an architecture with Fugu Ultra v2 but tuned to route each task to the leanest model that can solve it. It orchestrates Sakana's largest pool yet, weighted toward open-weight and specialized models including NVIDIA's Nemotron family through a partnership announced in August. Pricing is $2 per million input tokens and $6 output, with cached input at a flat $0.25 regardless of context length, which Sakana says puts its output rate 40 to 60 percent under Claude Sonnet 5, GPT-5.6 Terra, and Kimi K3. Context is 1 million tokens. Sakana reports best overall score on six benchmarks including Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and the internal SWEFish, and says it expands the cost-performance frontier on seven of ten tested benchmarks. The billing detail that matters more than the headline rate is tools. Because Fugu Max leans on open models without native web search, web_search and web_fetch run through Sakana's internal tools at $0.007 per call, and Sakana warns a single query may need several calls, a per-request cost that never shows up in the per-token price. The separately reported orchestration token fields that bill at full rates are documented for Fugu Ultra and Fugu Cyber, not Fugu Max, and Sakana says it never stacks model fees when several agents run. Fugu is not yet offered in the EU or EEA. All figures are self-reported and SWEFish is Sakana's own benchmark.
Input Price
$2.00
per 1M tokens
Output Price
$6.00
per 1M tokens
Context Window
1M
tokens
Released
2026-09
API access
Capabilities
Key Strengths
- ✓$2/$6 per 1M tokens with flat $0.25 cached input at any context length
- ✓Output rate 40 to 60 percent under Sonnet 5, GPT-5.6 Terra, and Kimi K3
- ✓Routes to the leanest capable model in a large open-weight pool
- ✓Best overall score on six of Sakana's reported benchmarks
- ✓1M token context window
- ✓Includes NVIDIA Nemotron models in its pool
Best For
- ▸High-volume agentic work on a fixed budget
- ▸Everyday coding and automation at mid-tier cost
- ▸Workloads that favor open-weight models for compliance
- ▸Replacing a $2-tier frontier model without losing capability
Benchmark Scores
| Benchmark | Score | Description |
|---|---|---|
| GPQA Diamond | 95.5 | Graduate-level science questions verified by domain experts |
Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.
Pricing Details
Input tokens
$2.00
per 1M tokens
Output tokens
$6.00
per 1M tokens
Estimated cost per 1K requests
$5.00
~1K input + ~500 output tokens avg
Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.