APIVOT is a vision-language-model planner for long-horizon robot tasks that adaptively interleaves language thoughts with visual thoughts. Language reasoning handles semantic decomposition, object selection, and action sequencing, while visual thoughts represent imagined future states used to verify geometric feasibility, including free-space and collision constraints. The authors report that APIVOT outperforms general-purpose VLMs and prior planning frameworks on long-horizon kitchen tasks, with the largest improvements in spatially constrained settings. The abstract also claims that the model learns meaningful modality-selection behavior, improving both planning success and reasoning efficiency.
No heat snapshots are available in the last 24 hours.