This paper studies generalization in offline reinforcement learning for contextual Markov decision processes. It argues that excessive pessimism does not inherently prevent optimal generalization; the crucial factor is whether the pessimistic value function respects the symmetries of the optimal solution. The paper shows theoretically that a mildly pessimistic but asymmetric value function can generalize worse than a highly pessimistic symmetric one. It connects pessimistic structure to dataset coverage and proposes applying data augmentation through a consistency loss during policy extraction, rather than ordinary offline training on augmented data. The approach is evaluated with IQL and CQL in a rotationally symmetric reacher environment.
No heat snapshots are available in the last 24 hours.