ABot-M0.5 is a world action model for mobile manipulation that targets three problems: coarse temporal modeling, entangled navigation and manipulation actions, and mismatch between inverse-dynamics training and autoregressive inference. It introduces intermediate latent actions, a dual-level Mixture-of-Transformers architecture that separates modality representations and action subspaces, and a dream-forcing training strategy based on model-predicted videos. The abstract reports state-of-the-art results on challenging mobile and fine-grained manipulation benchmarks, but does not provide benchmark names, numerical gains, or detailed evaluation settings.
No heat snapshots are available in the last 24 hours.