The paper introduces LaMem-VLA, a latent-memory-native framework for long-horizon robotic manipulation. Unlike approaches that append observation windows or retrieve history as auxiliary policy context, LaMem-VLA reconstructs retrieved experience into latent memory tokens within the VLA model’s native embedding space. A curator maintains short-term and long-term memory vaults, a seeker retrieves relevant evidence, a condenser compresses it into latent tokens, and a weaver interleaves those tokens with the current observation and instruction. The authors report experiments on SimplerEnv and LIBERO, but the provided abstract does not include quantitative results.
No heat snapshots are available in the last 24 hours.