Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

First seen · 7/2/2026, 12:00 PMLatest activity · 7/2/2026, 12:00 PM

ELDR targets decode routing in prefill-decode (PD) disaggregated serving for mixture-of-experts models. It derives an expert signature from a request’s prefill-time expert activations, partitions signature space with offline balanced K-means, and uses online locality-band routing to select the least-loaded worker among workers with the best signature match. A signature cache co-indexed with KV-cache blocks preserves exact signatures under prefix caching. The paper reports 5.9–13.9% lower median time per output token than the strongest of four load-balancing baselines across three MoE models and two workloads, on deployments of up to 40 GPUs, without changing outputs.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/2, 12:00 PMnot independentRepresentative
    ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving