Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

Xiaomi introduces Xiaomi-Robotics-1, a vision-language-action model pretrained on more than 100,000 hours of real-world manipulation trajectories collected with UMI devices. Its two-stage recipe combines broad action pretraining with post-training for robot embodiments and imperative human instructions. An auto-labeling pipeline adds natural-language descriptions of scene-state transitions to trajectory clips. The paper reports scaling gains with both data and model size, transfer of those gains to unseen real-robot tasks, and efficient fine-tuning for dexterous tasks. On simulation benchmarks, it reports 57.6% success on RoboCasa365 versus 46.6% for the previous best, and a 20.07 average score on RoboDojo versus 13.07.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories