Fireworks ships a Kimi K3 variant that writes 40% fewer output tokens
single source· 1 articles · confidence: medium · first seen 2026-09-28 07:22 UTC
What this means for you
Check what the 40% is measured against: output tokens per task, in one vendor-run A/B test with no named benchmark. If Kimi K3 output length is what drives your bill, Ember-1 is worth testing on your own traffic. Otherwise nothing changes today — it is a Research Preview on unchanged pricing.
Fireworks AI has released Ember-1, a version of Kimi K3 trained further on top of the existing model to write shorter reasoning traces — the working-out a model produces before its answer — rather than to spend less effort solving the problem. Fireworks reports a production A/B test in which output tokens per task fell from 49.3K to 29.9K, about 40% less, at essentially unchanged scores; no benchmark, harness or evaluation date is named. Ember-1 is API-only, labelled a Research Preview, at Kimi K3 pricing. The comparison is the vendor's own.
Models in this story
Key facts
- ·Fireworks AI released Ember-1, a Kimi K3 model trained further on top of the base model, on 28 September 2026. source
- ·Fireworks reports output tokens per task falling from 49.3K to 29.9K, about 40% fewer, in a production A/B test. source
- ·Fireworks says scores were essentially unchanged; no benchmark, harness or evaluation date is given. source
- ·Ember-1 is available as API-only and is described as a Research Preview. source
- ·Ember-1 is offered at Kimi K3 pricing. source
What the sources say
- MarkTechPost — Vendor announcement from Fireworks, with token counts from an internal production test and no named benchmark.
Sources
The original reporting. Follow these — they did the work.
- MarkTechPostFireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens2026-09-28