DADiff addresses online dynamics adaptation in reinforcement learning, where a policy is trained with abundant source-domain data but can collect only limited target-domain interactions. The method uses a diffusion-based generative process for next-state prediction and estimates dynamics mismatch from discrepancies between source- and target-domain generative trajectories. It offers two adaptation variants: reward modification and data selection. The paper also derives a theoretical bound relating policy performance differences across domains to generative trajectory deviation, and reports experiments across multiple types of domain shifts.
No heat snapshots are available in the last 24 hours.