The paper introduces DPBAC, a diffusion-policy algorithm for offline reinforcement learning that combines behavioral advantage corrected policy evaluation (BAC-PE) with diffusion-based policy modeling. BAC-PE uses the behavior policy’s Q-function to correct the learned policy’s Q-function, targeting distribution-shift-related pessimism and overestimation bias. The method also matches the distributions of behavior and learned policies and adds Q-value guidance during training. The authors report stronger results than existing offline methods across multiple D4RL domains, although the abstract does not provide task-level scores, baselines, or statistical details.
No heat snapshots are available in the last 24 hours.