The paper introduces PAMD, or Pairwise Adaptive Mahalanobis Distance, for learning latent state similarity in visual reinforcement learning. It addresses a design tension in bisimulation-based representation learning: fixed global norms may be too restrictive, while unconstrained pairwise distances can produce degenerate solutions. PAMD parameterizes a positive-definite, pair-conditioned Mahalanobis metric and is intended as a plug-in component for existing bisimulation methods. According to the abstract, experiments on visual MuJoCo continuous-control tasks show substantial improvements in the final performance of several recent bisimulation-based RL algorithms.
No heat snapshots are available in the last 24 hours.