BottleCap AI ships a Qwen3.8-27B fine-tune that generates fewer tokens per answer
single source· 1 articles · confidence: medium · first seen 2026-09-24 18:58 UTC
What this means for you
If you serve Qwen3.8-27B on vLLM or SGLang, this drops in without code changes and cuts generated tokens per answer by 37.2%, for 0.86 percentage points of macro accuracy. Whether that trade pays depends on your workload. No evaluation date or harness is published, so re-measure on your own tasks before switching.
BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B that generates 37.2% fewer tokens before its final answer — what BottleCap calls "thinking" — across 12 benchmarks. Macro accuracy falls from 86.65% to 85.79%, down 0.86 percentage points, while the long-context AA-LCR evaluation improves by 2.25 points. It is a drop-in replacement on the vLLM and SGLang serving stacks, with FP8, NVFP4, GGUF and MLX builds. No evaluation date or harness is given for the reported scores.
Models in this story
Key facts
- ·BottleCap AI released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B, on 24 September 2026. source
- ·The model generates 37.2% fewer tokens before its final answer across 12 benchmarks. source
- ·Macro accuracy moves from 86.65% to 85.79%, a drop of 0.86 percentage points. source
- ·The long-context AA-LCR evaluation improves by 2.25 percentage points. source
- ·The model is a drop-in replacement on vLLM and SGLang, with FP8, NVFP4, GGUF and MLX builds. source
What the sources say
- MarkTechPost (press) — Release note listing the fine-tune's token savings, accuracy change and available build formats.
Sources
The original reporting. Follow these — they did the work.