This paper studies training a single-layer self-attention model directly with external-regret and swap-regret objectives over probability-simplex policies. The authors show that particular stationary points have forward passes matching smoothed fictitious play and the classical Blum–Mansour no-pass implementation, respectively. In the swap-regret construction, each attention head performs an external-regret update through smoothed fictitious play. The paper argues that deploying these learned dynamics can yield coarse correlated equilibrium under external regret and correlated equilibrium under swap regret, without supervised traces of the underlying algorithms. The claims are theoretical and centered on a specified minimal architecture.
No heat snapshots are available in the last 24 hours.