Fireworks Research published Ember-1 on 23 September: a specialized derivative of Kimi K3 that, in the vendor's own measurements, matches K3's quality while emitting 40% fewer tokens. What makes it worth reading is not the headline number but the failure mode it targets.

Reasoning tokens get re-billed every turn

The argument in the post is arithmetic. Reasoning models like Kimi K3 spend "sometimes more than 90%" of generated tokens on internal reasoning rather than the answer. That is expensive once, and much worse across a multi-turn agent loop: each turn replays all prior reasoning back to the model, so context grows roughly quadratically with turn count. Every early-turn reasoning trace is re-read and re-billed on every subsequent call.

The obvious workaround — turning reasoning effort down — was already tried and rejected. Fireworks states plainly that lower effort settings "gave up too much quality." So they trained the model instead, running more than 50 training experiments and over 200 evaluations on Fireworks Serverless Training, using their own data and (they say) no customer data.

The numbers on public benchmarks

Costs below use public Kimi K3 API pricing — uncached input $3/M tokens, cached input $0.30/M, output $15/M — and the delta is per task versus K3 at max (default) effort:

Benchmark N K3 Low K3 High K3 max Ember-1 Ember-1 vs K3 max
Terminal Bench 2.1 89 76.4% 77.6% 80.9% 82.0% -51.9% / -23.1 USD
SWE-bench Verified 500 80.4% 86.0% 93.2% 92.2% -15.5% / -68.1 USD
SWE-Interact 75 6.7% 13.3% 21.3% 20.0% -32.5% / -60.8 USD
DeepSWE 1.1 113 55.8% 62.8% 66.4% 75.2% -23.7% / -126.9 USD
τ-2 Bench Airline 50 64% 64% 64% 66% -5.9% / -0.3 USD

Two rows deserve a second look. DeepSWE 1.1 is the one benchmark where Ember-1 beats K3 max on pass rate (75.2% vs 66.4%) while still spending 23.7% fewer tokens — which is odd enough that I would treat it as the benchmark to verify first on my own workload. And τ-2 Bench Airline is effectively flat on both axes: 2 points of pass rate, 5.9% token reduction on 50 samples. If your traffic looks like that benchmark, Ember-1 buys you almost nothing.

What the live A/B tests actually measured

Two customers ran Ember-1 against Kimi K3 on production coding workloads. Fireworks reports "approximately 35% fewer tokens per task at comparable quality," with task completion, success scores and failure rates all moving the right direction. One customer is now running it in live production with plans to replace the base model entirely.

Their one published A/B table:

Score Steps Output Tokens Reasoning token reduction Total token reduction
Kimi K3 0.751 23.8 49.3K - -
Ember-1 0.753 21.4 29.9K 71.3% 39%

The score is a rounding error apart. Output tokens drop by roughly 19.4K, and the step count falls from 23.8 to 21.4 — fewer wasted loops, not just shorter ones. The token-savings spread across the post (35–50% across seven benchmarks and two customers' production traffic) is the number I would plan against rather than the 40% headline.

Available as a two-week research preview

The distribution model is the part that affects planning. Ember-1 rolls out as a serving option alongside base Kimi K3 on Serverless, labelled a Research Preview. Fireworks says research releases get two-week serverless access, and become permanent based on community demand. That is a hard deadline on any pilot you start — if you build a dependency on Ember-1 and demand does not materialize, it goes away.

There is also training support for Ember-1, so enterprises can fine-tune their own token-efficient variants with their own data. Fireworks frames the direction explicitly: "The future of open models is specialized models trained on your specific workload."

What I would watch for

I would reach for Ember-1 when reasoning tokens dominate the bill — long agentic coding loops, multi-turn tool use — and not for a chatbot, where reasoning is a small fraction of output. The catch is that every number here is vendor-published, on vendor-chosen benchmarks, and the one benchmark it strictly dominates (DeepSWE 1.1) sits oddly next to the rest of the table. Run the τ-2-style short, low-reasoning traffic through it before assuming the savings are universal, and do not build anything load-bearing on a preview that expires in two weeks.

Sources: