Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Think Before You Grid-Search: Floor-First Triage for LLM Serving

First seen · 7/7/2026, 02:11 PMLatest activity · 7/7/2026, 02:11 PM

This paper proposes Floor First, a residual-driven workflow for optimizing LLM serving before launching expensive profiling or broad configuration searches. Each decode step is represented as a five-dimensional resource vector: HBM bytes, FLOPs, network bytes, network messages, and KV capacity. Summing each resource and taking the maximum yields optimistic and pessimistic floors; a measured result inside that interval reveals overlap quality. In a case study of a DeepSeek-V3.2-style 671B MoE/MLA model on 16 NVIDIA H20 GPUs, TP16 is favored for single-stream latency, while EP16+DP-attention offers roughly an order-of-magnitude larger capacity wall.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/7, 02:11 PMnot independentRepresentative
    Think Before You Grid-Search: Floor-First Triage for LLM Serving