arXivFatemeh Saberi Khomami
Online Change-Point Detection for Cooperative MARL: Patterns of Past Rewards (PPR)
Original title:Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning
Papers68
When cooperative multi-agent environments shift mid-training, past coordination experience can quickly misguide learning agents. This paper introduces Patterns of Past Rewards (PPR), an algorithm-agnostic change-point detector that monitors agents' reward trajectories using statistical drift detection on smoothed returns. Evaluated in a custom Speaker-Listener environment across non-stationary setups, PPR curbs repeated false alarms while reliably flagging environmental shifts.
Why it's worth reading
It provides a lightweight, algorithm-agnostic method to detect non-stationarity in cooperative MARL, helping systems recognize exactly when to trigger adaptation.
Tags
MARLReinforcement LearningChange-point DetectionNon-stationarityMulti-AgentOnline Detection