TurboServe is a serving system designed for streaming video generation, where long-lived sessions produce video progressively in chunks. It jointly optimizes session placement and GPU provisioning through migration-aware rebalancing and load-driven autoscaling. The implementation adds coalesced chunk processing for batching active sessions, GPU-CPU offloading for suspending and resuming sessions, and NCCL-based GPU-to-GPU migration. On production traces from Shengshu Technology, across multiple model sizes and GPU clusters with up to 64 NVIDIA B300 GPUs, the paper reports average reductions of 37.5% in worst-case per-chunk latency and 37.2% in GPU operating cost. The code is publicly available.
No heat snapshots are available in the last 24 hours.