Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

First seen · 7/9/2026, 12:00 PMLatest activity · 7/9/2026, 12:00 PM

This paper introduces LingBot-Video, a DiT-based video pretraining paradigm designed for embodied intelligence rather than primarily for visual content creation. It uses a Mixture-of-Experts architecture to increase modeling capacity while improving inference efficiency, and scales the model from scratch. A data profiling engine augments internet video with robot-oriented footage covering manipulation, navigation, and egocentric perspectives. Training adds a multidimensional reward system targeting physical rationality and task completion, beyond aesthetics, prompt following, and motion consistency. The authors describe it as the first large-scale open-source MoE video foundation model aimed at connecting visual generation with physical actuation.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/9, 12:00 PMnot independentRepresentative
    Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence