Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Ember-1

Flagship

by Fireworks AI

Ember-1 is the first model in a new Fireworks Research series, though not the first time Fireworks has post-trained someone else's open weights (FireFunction-v2 was built on Llama 3 70B): a reasoning model post-trained from Moonshot's Kimi K3, released September 23, 2026 as a Research Preview on Fireworks Serverless, where research releases get two weeks of serverless access and become permanent based on community demand. The pitch is efficiency rather than a new capability tier. Fireworks says it learned to cut unnecessary reasoning while keeping the thinking that matters, producing roughly 40 percent fewer tokens than K3 at comparable quality, with live A/B tests showing approximately 35 percent fewer tokens per task. Pricing holds at $3 per million input tokens and $15 output on a 1,048,576 token context window, the same rate card as Kimi K3 itself, so the entire value proposition sits in the token count rather than the sticker price. On Fireworks' own benchmark table Ember-1 scores 92.2 percent on SWE-bench Verified against 93.2 for K3 at max settings and 82.0 on Terminal-Bench 2.1 against 80.9, essentially holding capability steady while using fewer tokens, and the company reports Ember-1 set a new cost-per-task Pareto frontier across open and closed models on Doximity's Bedside Bench, a 500-case clinical benchmark. Every number here is vendor-reported; no independent evaluator has published a score for Ember-1 yet, and Fireworks has not published weights, so this is an API-only offering built on top of Kimi K3's architecture rather than a new open release.

Input Price

$3.00

per 1M tokens

Output Price

$15.00

per 1M tokens

Context Window

1.0M

tokens

Released

2026-09

API access

Capabilities

textvisiontool-usecodereasoning

Key Strengths

  • ✓Roughly 40 percent fewer reasoning tokens than Kimi K3 at comparable quality, vendor-reported
  • ✓92.2 percent on SWE-bench Verified against K3's 93.2 on Fireworks' own run, with a 15.5 percent reduction on that benchmark
  • ✓$3/$15 pricing identical to the Kimi K3 rate card
  • ✓1,048,576 token context window inherited from K3
  • ✓Vendor-reported cost-per-task Pareto frontier on Doximity's Bedside Bench clinical benchmark

Best For

  • ▸Coding and agentic workloads where token spend on reasoning is the cost driver
  • ▸Teams already on Kimi K3 looking to cut inference cost without changing quality
  • ▸High-volume software engineering and search tasks
  • ▸Budget-conscious deployments that still need Kimi K3-class reasoning depth

Benchmark Scores

BenchmarkScoreDescription
SWE-bench92.2Real-world software engineering tasks from GitHub issues (SWE-bench Verified)

Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.

Pricing Details

Input tokens

$3.00

per 1M tokens

Output tokens

$15.00

per 1M tokens

Estimated cost per 1K requests

$10.50

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide