Ternary weights shrink a 27B model to 5.9 GB

single source· 1 articles · confidence: medium · first seen 2026-09-18 18:06 UTC

What this means for you

If you serve a 27B model and memory is the constraint, a checkpoint at 5.93 GB against 53.80 GB is worth testing. But the 98.2% retention figure is PrismML's own, across 20 benchmarks it has not named — re-run your own evals before swapping anything.

PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B: each weight takes one of three values rather than 16-bit precision, so the checkpoint takes 5.93 GB against 53.80 GB in FP16. PrismML reports it keeps 98.2% of the parent model's average across 20 benchmarks; no evaluation date, harness or benchmark list is given, and the figures are the vendor's own. The release is Apache 2.0 licensed, accepts text and images, and carries a 262K-token context.

Models in this story

Key facts

  • ·Ternary Bonsai 2 27B is a ternary-weight version of Qwen3.8 27B, in which each weight takes one of three values instead of FP16 precision. source
  • ·The model occupies 5.93 GB, against 53.80 GB for the same model in FP16. source
  • ·PrismML reports the model keeps 98.2% of the parent model's average across 20 benchmarks. source
  • ·The release is licensed under Apache 2.0. source
  • ·The model accepts text and images and supports a 262K-token context. source
  • ·PrismML demoed the model driving Cline, a coding tool. source

What the sources say

  • MarkTechPostRelays the vendor's file size, accuracy retention and licence for the ternarised Qwen3.8 model.

Sources

The original reporting. Follow these — they did the work.

← the wire