UniVR studies whether complex reasoning, fine-grained physical dynamics, and long-horizon planning can be learned jointly from pure visual demonstrations, without image-text pairs or task-specific heuristics. Its central method, VR-GRPO, combines global and step-level rewards to encourage both logical coherence and physical consistency during visual reasoning. The authors introduce VR-X, a benchmark assembled from 16 sources covering long-horizon manipulation, spatial puzzles, and physical reasoning. The abstract reports improvements of up to 25% on VR-X and gains on other multimodal understanding benchmarks, while stating that code, data, and models are open-sourced.
No heat snapshots are available in the last 24 hours.