GORDON learns dense reinforcement-learning rewards from action-free video demonstrations by representing scenes as graphs of detected objects and spatial relations. A self-supervised graph neural network maps these graphs into a task-aligned latent space, while activity-aware pooling emphasizes relevant objects and masks robot-dominated motion. Distances to demonstrated goal configurations define progress rewards. Temporal reward profiles expose stage-wise object transitions, enabling automatic subtask discovery and sequential policy composition. On seven tasks from MAGICAL and ManiSkill3, the authors report a 74.4% average success rate on long-horizon tasks, about 35 percentage points above the best learned baseline and 25 points above an oracle-based comparison.
No heat snapshots are available in the last 24 hours.