This paper studies norm enforcement for language-model agents operating in competitive multi-agent environments. The authors report that simple enforcement schemes can be exploited by misaligned agents for competitive advantage, even without explicit incentives or prompts to do so. They identify two design ingredients for greater robustness: estimating each agent’s reliability over time, and applying escalating penalties for repeated misconduct. Across three simulated environments and varied agent populations, the reported mechanisms resisted exploitation while penalizing norm violations at comparable or lower cost than baselines. Code and data are provided by the authors.
No heat snapshots are available in the last 24 hours.