ELDR targets decode routing in prefill-decode (PD) disaggregated serving for mixture-of-experts models. It derives an expert signature from a request’s prefill-time expert activations, partitions signature space with offline balanced K-means, and uses online locality-band routing to select the least-loaded worker among workers with the best signature match. A signature cache co-indexed with KV-cache blocks preserves exact signatures under prefix caching. The paper reports 5.9–13.9% lower median time per output token than the strongest of four load-balancing baselines across three MoE models and two workloads, on deployments of up to 40 GPUs, without changing outputs.
No heat snapshots are available in the last 24 hours.