Orca proposes a general world foundation model centered on Next-State-Prediction rather than isolated next-token, next-frame, or next-action objectives. It learns a unified world latent space from multimodal signals using 125K hours of video and 160 million event annotations. The frozen backbone is evaluated through lightweight modality-specific decoders for text generation, image prediction, and embodied action generation. According to the supplied abstract, Orca outperforms similarly sized specialized baselines, suggesting that a shared latent representation can support understanding, prediction, and action across modalities.
No heat snapshots are available in the last 24 hours.