WitCert introduces a runtime meter for KV-cache quantization that provides per-layer, per-head, and per-step upper bounds on the total variation between exact and compressed attention. Its deterministic certificate is designed to remain sound for cache-preserving black-box quantizers and adaptive queries, while a tighter probabilistic certificate targets a controlled subtractively dithered INT8 quantizer under an explicit request-level failure budget. An environment-guarded SGLang integration supports live measurement and risk-based gating. The authors report that gating raises raw-cast FP8 performance from 22.8 to 79.7 on hard RULER tasks, and that certified INT8 serves 1.88 times more KV tokens at equal memory.
No heat snapshots are available in the last 24 hours.