This paper introduces a self-supervised masked-prediction auxiliary task for vision-based reinforcement learning. Instead of reconstructing raw observations, the method uses sequences collected by an agent and their contextual information to predict masked content in latent space. Combined with Transformers, it learns compressed representations that can be provided to reinforcement-learning agents. The abstract reports improved sample efficiency and performance surpassing state-of-the-art sample-efficient methods across multiple continuous- and discrete-control benchmarks. Detailed benchmark settings, ablations, compute costs, and statistical significance require inspection of the full paper.
No heat snapshots are available in the last 24 hours.