Smaller, faster, safer: running Kimi and GLM at scale
7.6 relevance
Score Breakdown
technical depth 8
novelty 8
actionability 6
community 8
strategic 7
personal 9
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Cloudflare blog on running Kimi/GLM at scale, deep technical content on model optimization.
Summary
Cloudflare's Workers AI uses FP8 KV cache quantization to double Kimi K2.6's context capacity from 686K to 1.37M tokens, achieving 2,192 tok/s at 64 concurrent requests (41% higher peak than BF16) with no accuracy loss, and compresses GLM 5.2 weights to INT4 reducing checkpoint from 705GB to 421GB. These optimizations, built on SGLang and disaggregated prefill/decode, lower cost per token while maintaining model fidelity.