Fireworks sells a shorter-thinking build of Kimi K3 — and its own post undercuts its 40% headline
Ember-1 shipped as a research preview on 23 September, trained to cut unnecessary reasoning tokens while holding quality. The headline says 40% fewer tokens; the customer A/B test inside the same post measures about 35%. In reasoning-token pricing, that percentage is the product.
The mechanism is not a new base model. Ember-1 is built on Moonshot's Kimi K3, and what Fireworks changed is the length of the reasoning trace — shorter chains of thought for the same answer. That is a direct attack on the economics of reasoning models, where the customer pays for every intermediate token the model emits while thinking, and where a verbose reasoner can cost several times a terse one for identical output quality.
The number does not hold still, and Fireworks is the one who shows it. The headline claim is Kimi K3's quality with 40% fewer tokens. Further down the same post, a customer A/B test measures approximately 35% fewer tokens per task at comparable quality, and 39% in one specific benchmark. The broadest statement is that across seven benchmarks and two customers' production traffic, Kimi K3's reasoning could be shortened by 35 to 50% without sacrificing accuracy. So 40% is a figure inside a measured range, and the single cleanest customer measurement came in at the bottom of it. Anyone modelling cost savings should use 35%.
The pricing framing needs the same care. The post uses public Kimi K3 API pricing as its comparison baseline — $3 per million uncached input tokens, $0.30 per million cached, $15 per million output — to make the token saving legible in dollars. Ember-1's own per-token price is not stated, so the saving is expressed as a token reduction against someone else's price card rather than as a price. That is a meaningful omission for a product whose whole pitch is cost.
The release posture is provisional. Ember-1 is available as a research preview on Fireworks Serverless, with two-week access to new research models — a trial window, not a committed endpoint. The pattern to watch is the business model: a serving company taking a strong open-weight model from another lab and selling a cheaper-to-run derivative of it, with the margin coming from tokens not emitted.
- Confirmed Ember-1 is built on Kimi K3 and shipped 23 September as a research preview on Fireworks Serverless. Fireworks AI
- Claimed The headline is 40% fewer tokens than Kimi K3, but the customer A/B test in the same post measures about 35% per task, and 39% in one benchmark. Fireworks AI
- Claimed Across seven benchmarks and two customers' production traffic, Fireworks says Kimi K3's reasoning could be shortened 35–50% without sacrificing accuracy. Fireworks AI
Models & releasesMoney & markets
Today in the September 25, 2026 edition · front page