Cloudflare describes how Workers AI serves Moonshot’s Kimi K-series and Z.ai’s GLM models, which are large, long-context mixture-of-experts systems that are difficult to fit efficiently into GPU memory. The deployment combines three techniques: KV-cache quantization, model-weight compression, and protection for shared caches when more requests are colocated on the same hardware. Cloudflare says its experiments and production traffic use SGLang, and that the optimizations support more customers at lower cost without changing model accuracy. The post focuses on serving engineering rather than introducing new model architectures.
No heat snapshots are available in the last 24 hours.