SeededGrasp separates semantic reasoning from geometric grasp execution. A vision-language model predicts a task-relevant seed point from language and visual scene context, while a lightweight downstream model generates the final grasp. This design avoids expensive end-to-end VLM-grasp training and is intended to support multiple robot embodiments. The authors release a multi-embodiment tabletop grasping dataset containing more than 2.5 million grasps in cluttered scenes. They report 72% success in simulation and 78% in real-world grasping experiments, alongside code and data through the project website.
No heat snapshots are available in the last 24 hours.