This paper studies how episodic exploration bonuses interact with neural memory in partially observable reinforcement learning. Across three environments, the same bonus produces three outcomes: it magnifies memory-architecture differences when useful information must be discovered and retained without direct supervision; equalizes architectures when the sought cue is directly reward-supervised; or has no effect when observations are scheduled. Controlled reward changes suggest that reward structure, rather than reward density alone, determines the interaction. The authors formalize this distinction with observation-anchored reward machines and argue that exploration and memory are complementary: exploration exposes information, while memory converts it into return.
No heat snapshots are available in the last 24 hours.