The paper introduces GenDa (Generalizable Data-efficient Agent), a unified framework for unsupervised reinforcement learning. It targets two claimed bottlenecks in current off-policy URL methods: non-stationary skill semantics and brittle generalization. GenDa uses skill relabeling to reduce semantic non-stationarity and improve pre-training data efficiency, while a Complementary Information Bottleneck (CIB) encourages skill policies to rely on ego-centric features and resist distribution shifts in downstream control tasks. The authors report improvements in scalability, generalization, and data efficiency across multiple experiments, although the abstract does not provide quantitative results.
No heat snapshots are available in the last 24 hours.