Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Fatemeh Saberi Khomami·Sep 4, 2026, 3:49 PM

Online Change-Point Detection for Cooperative MARL: Patterns of Past Rewards (PPR)

Original title:Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning

Papers68

When cooperative multi-agent environments shift mid-training, past coordination experience can quickly misguide learning agents. This paper introduces Patterns of Past Rewards (PPR), an algorithm-agnostic change-point detector that monitors agents' reward trajectories using statistical drift detection on smoothed returns. Evaluated in a custom Speaker-Listener environment across non-stationary setups, PPR curbs repeated false alarms while reliably flagging environmental shifts.

Why it's worth reading

It provides a lightweight, algorithm-agnostic method to detect non-stationarity in cooperative MARL, helping systems recognize exactly when to trigger adaptation.

Tags

MARLReinforcement LearningChange-point DetectionNon-stationarityMulti-AgentOnline Detection

Score breakdown

  • Novelty68
  • Impact65
  • Practicality72
  • Credibility70
  • Timeliness68