The paper introduces the World Critic Model (WCM), a lightweight LeJEPA-based critic for vision-language-action reinforcement learning. Instead of estimating value from single-frame observations or backbone latents alone, WCM jointly predicts future latent states and values, providing an explicit world-modeling objective for temporal representation learning. It is designed to work with both on-policy and off-policy pipelines and is compatible with Pi0, Pi0.5, and OpenVLA-OFT. The authors report experiments across 149 tasks on four benchmarks, plus seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5 with off-policy RL.
No heat snapshots are available in the last 24 hours.