This paper proposes a two-stage EEC explanatory framework and an Agent Response Map (ARM) for interpreting multi-agent reinforcement learning policies. ARM exposes spatial decision patterns, including aggregation and avoidance regions, and suggests that robots implicitly learn geometric fields in their environments as navigation targets. In cooperative shape assembly, the unoccupied target interior becomes a preferred destination and shifts toward the boundary as the center fills. In competitive predator-prey pursuit-evasion, prey agents converge toward the boundary of the predators’ Voronoi diagram. The framework is evaluated on these two task types.
No heat snapshots are available in the last 24 hours.