Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingNewsWatching1 independent reportsincl. 1 official10

Smaller, Faster, Safer: Running Kimi and GLM at Scale

First seen · 8/3/2026, 09:00 PMLatest activity · 8/3/2026, 09:00 PM

Cloudflare describes how Workers AI serves Moonshot’s Kimi K-series and Z.ai’s GLM models, which are large, long-context mixture-of-experts systems that are difficult to fit efficiently into GPU memory. The deployment combines three techniques: KV-cache quantization, model-weight compression, and protection for shared caches when more requests are colocated on the same hardware. Cloudflare says its experiments and production traffic use SGLang, and that the optimizations support more customers at lower cost without changing model accuracy. The post focuses on serving engineering rather than introducing new model architectures.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. OfficialCloudflare Blog8/3, 09:00 PMRepresentative
    Smaller, Faster, Safer: Running Kimi and GLM at Scale