Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Moral Hazard in Multi-Agent Language Models

First seen · 7/27/2026, 12:13 PMLatest activity · 7/27/2026, 12:13 PM

This paper adapts Holmström’s team moral-hazard model into the Dialogue Moral Hazard Game, a controlled textual environment where agents may pay a query cost to reveal hidden safety information that mainly benefits another agent. It evaluates nine open-weight models and one frontier API model across querying, information transfer, local-reward preservation, unsafe choices, format validity, and team success. The abstract reports that base open-weight models often preserve local reward without achieving team success, or query without transmitting decision-changing information. GPT-5.6 Sol reaches ceiling behavior in the primary setting, with a query threshold closely tracking the theoretical private-share boundary in 3,015 scripted-partner decisions.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/27, 12:13 PMnot independentRepresentative
    Moral Hazard in Multi-Agent Language Models