The paper introduces Prefix-Optimal Generative Policies (POGP), which learn a prefix value function at every intermediate denoising step of a diffusion policy. The function provides an auxiliary training signal for producing useful intermediate actions and supports test-time early stopping when further denoising is unlikely to help. Across four MuJoCo environments and 12 baselines, POGP reduces required denoising iterations by approximately 2.7-fold while retaining near-full task performance. Against state-of-the-art dynamic diffusion baselines, prefix training improves final task performance by approximately 3.5%.
No heat snapshots are available in the last 24 hours.