RoboTTT introduces Test-Time Training into robot foundation models, using fast weights updated by gradient descent during both training and inference to compress long visuomotor histories into recurrent state. The approach scales context to 8K timesteps without increasing inference latency. On challenging real-robot manipulation tasks, the authors report an 87% overall improvement over a single-step context baseline. An 8K-context model outperforms the same model pretrained with 1K timesteps by 62%, and completes a five-minute, ten-stage assembly task that no baseline completes. The method also enables one-shot imitation from human videos, online policy improvement, and greater perturbation robustness.
No heat snapshots are available in the last 24 hours.