Skip to content

Smaller, faster, safer: running Kimi and GLM at scale

7.6 relevance
Score Breakdown
technical depth
8
novelty
8
actionability
6
community
8
strategic
7
personal
9

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

Cloudflare blog on running Kimi/GLM at scale, deep technical content on model optimization.

AI/ML blog.cloudflare.com
Smaller, faster, safer: running Kimi and GLM at scale
Summary

Cloudflare's Workers AI uses FP8 KV cache quantization to double Kimi K2.6's context capacity from 686K to 1.37M tokens, achieving 2,192 tok/s at 64 concurrent requests (41% higher peak than BF16) with no accuracy loss, and compresses GLM 5.2 weights to INT4 reducing checkpoint from 705GB to 421GB. These optimizations, built on SGLang and disaggregated prefill/decode, lower cost per token while maintaining model fidelity.