The paper argues that foundation-model agents performing optimal planning in stylized social dilemmas can consistently converge on stable cooperation, contrary to the mutual-defection prediction of classical game theory. It introduces the embedded Bayesian agent, which models itself as part of the universe and remains uncertain about its own decision-making algorithm. Through similarity inference, an agent treats its own deliberation as evidence about a behaviorally similar partner’s likely choice. The authors formalize this mechanism with embedded equilibrium, a proposed replacement for Nash equilibrium in modeling social behavior by modern AI agents.
No heat snapshots are available in the last 24 hours.