Play2Perfect proposes task-agnostic reinforcement-learning pretraining through play on diverse objects and goals before fine-tuning for precise assembly. The learned manipulation priors include grasping, in-hand reorientation, and pose reaching. The paper studies object diversity, training objectives, trajectory diversity, and goal precision. According to the abstract, the resulting prior is 33x more sample-efficient than training from scratch, even with dense multi-stage rewards. It reports zero-shot sim-to-real transfer with 60% success on insertions with only 0.5 mm contact clearance, plus more than 50% success on long-horizon multi-part assembly and screwing.
No heat snapshots are available in the last 24 hours.