Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Online Change-Point Detection for Cooperative MARL: Patterns of Past Rewards (PPR)

First seen · 9/4/2026, 11:49 PMLatest activity · 9/4/2026, 11:49 PM

When cooperative multi-agent environments shift mid-training, past coordination experience can quickly misguide learning agents. This paper introduces Patterns of Past Rewards (PPR), an algorithm-agnostic change-point detector that monitors agents' reward trajectories using statistical drift detection on smoothed returns. Evaluated in a custom Speaker-Listener environment across non-stationary setups, PPR curbs repeated false alarms while reliably flagging environmental shifts.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv9/4, 11:49 PMnot independentRepresentative
    Online Change-Point Detection for Cooperative MARL: Patterns of Past Rewards (PPR)