WorldDiT introduces a unified diffusion Transformer for robot learning that jointly generates continuous action chunks and predicts normalized RGB patches from future camera frames. The architecture is designed to couple action generation with visual world modeling without relying on a large pretrained vision-language model as its action backbone. On all four LIBERO simulation suites, the paper reports that WorldDiT lies on the Pareto frontier of total parameter count versus mean success among methods reporting the complete set of suites. The authors position it as a strong sub-billion-parameter baseline for future scaling studies.
No heat snapshots are available in the last 24 hours.