Fireworks ships a Kimi K3 variant that writes 40% fewer output tokens

single source· 1 articles · confidence: medium · first seen 2026-09-28 07:22 UTC

What this means for you

Check what the 40% is measured against: output tokens per task, in one vendor-run A/B test with no named benchmark. If Kimi K3 output length is what drives your bill, Ember-1 is worth testing on your own traffic. Otherwise nothing changes today — it is a Research Preview on unchanged pricing.

Fireworks AI has released Ember-1, a version of Kimi K3 trained further on top of the existing model to write shorter reasoning traces — the working-out a model produces before its answer — rather than to spend less effort solving the problem. Fireworks reports a production A/B test in which output tokens per task fell from 49.3K to 29.9K, about 40% less, at essentially unchanged scores; no benchmark, harness or evaluation date is named. Ember-1 is API-only, labelled a Research Preview, at Kimi K3 pricing. The comparison is the vendor's own.

Models in this story

Key facts

  • ·Fireworks AI released Ember-1, a Kimi K3 model trained further on top of the base model, on 28 September 2026. source
  • ·Fireworks reports output tokens per task falling from 49.3K to 29.9K, about 40% fewer, in a production A/B test. source
  • ·Fireworks says scores were essentially unchanged; no benchmark, harness or evaluation date is given. source
  • ·Ember-1 is available as API-only and is described as a Research Preview. source
  • ·Ember-1 is offered at Kimi K3 pricing. source

What the sources say

  • MarkTechPost — Vendor announcement from Fireworks, with token counts from an internal production test and no named benchmark.

Sources

The original reporting. Follow these — they did the work.

← the wire