Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning
This paper studies how episodic exploration bonuses interact with neural memory in partially observable reinforcement learning. Across three environments, the same bonus produces three outcomes: it magnifies memory-architecture differences when useful information must be discovered and retained without direct supervision; equalizes architectures when the sought cue is directly reward-supervised; or has no effect when observations are scheduled. Controlled reward changes suggest that reward structure, rather than reward density alone, determines the interaction. The authors formalize this distinction with observation-anchored reward machines and argue that exploration and memory are complementary: exploration exposes information, while memory converts it into return.
Why it's worth reading
It reframes sparse-reward design by showing that bonus effectiveness depends on what the reward supervises and what memory must retain, not simply on how often rewards occur.