When cooperative multi-agent environments shift mid-training, past coordination experience can quickly misguide learning agents. This paper introduces Patterns of Past Rewards (PPR), an algorithm-agnostic change-point detector that monitors agents' reward trajectories using statistical drift detection on smoothed returns. Evaluated in a custom Speaker-Listener environment across non-stationary setups, PPR curbs repeated false alarms while reliably flagging environmental shifts.
No heat snapshots are available in the last 24 hours.