This paper introduces BadWAM, a framework for World-Action Drift attacks against world-action models (WAMs), which jointly generate actions and predict future world states. Small visual perturbations can either directly induce task-failing actions or preserve the model’s predicted future while shifting its executed behavior. The latter exposes a WAM-specific failure mode: plausible imagination does not guarantee aligned control. Across WAM variants, the reported action-only attack reduced task success from 96.5% to 43.1%, while moderate future-preserving regularization retained attack effectiveness and reduced imagination drift.
No heat snapshots are available in the last 24 hours.