Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Breakthrough in model compression (near-lossless 9x reduction), highly relevant to AI deployment.
Bonsai 2 27B, based on Qwen3.8 27B, uses ternary {−1,0,+1} weights with FP16 scaling to achieve 1.76 bits per weight and a 5.9GB footprint—9x smaller than full-precision while retaining 98.2% aggregate benchmark performance. Released under Apache 2.0 with a 262K-token context window, it reaches 143 tokens/second on RTX 5090 and 46.8 on M5 Max, making it viable for local coding agents, multimodal workflows, and private document analysis. This closes the retention gap from 95% to over 98% versus the prior Bonsai, pushing low-bit models toward practical losslessness for deployment.
PrismML