PhiZero is a physical world model organized around “physical language,” a compact discrete representation of world-state transitions. Rather than predicting future frames directly in pixel space, it learns transition sequences from in-the-wild videos through self-supervision, reasons over those sequences, and then renders the predicted evolution into video. The paper reports experiments on generation and understanding benchmarks, as well as applications to interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer. The supplied abstract does not provide model size, dataset scale, benchmark names, or numerical results.
No heat snapshots are available in the last 24 hours.