Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

RxBrain: An Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

First seen · 7/17/2026, 12:00 PMLatest activity · 7/17/2026, 12:00 PM

The paper introduces Hy-Embodied-RxBrain, an embodied cognition foundation model that combines language reasoning and visual imagination in a single planning sequence. Language encodes task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination predicts intermediate and final physical states. The model uses a unified multimodal Mixture-of-Transformers architecture for language, image, and video understanding and generation. The authors also present an automatic pipeline that turns embodied videos into joint text-visual planning supervision, and introduce RxBrain-Bench. The abstract reports results in embodied understanding, generation, and continuous robot action generation, including promising real-robot performance without large-scale action-data pretraining.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/17, 12:00 PMnot independentRepresentative
    RxBrain: An Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination