This paper treats demonstration organization as a first-class factor in Vision-Language-Action imitation learning. Using a dual-arm robotic platform, it proposes three principles: decompose complex manipulation into progressively learnable sub-skills, standardize the interaction environment, and order demonstrations by increasing task complexity. The strategy is evaluated on block grasping and sorting, and towel folding. According to the abstract, it consistently improves task success rate and training stability over directly collecting complete end-to-end trajectories. The work focuses on dataset construction and curriculum structure rather than proposing a new VLA architecture.
No heat snapshots are available in the last 24 hours.