Apple-PI introduces a benchmark that evaluates video models against explicit physical laws rather than judging only the plausibility of generated outputs. Its Orchard dataset contains 400 videos spanning ten classical-mechanics tasks, separating single-law diagnosis from multi-law generalization. The protocol evaluates three stages: Perception, Formulation, and Deduction, using chain-of-frames prompting and infographic-annotated initial frames. A hybrid suite combines MLLM-based subjective scoring with law-grounded objective metrics. Across 11 models, the best video model scores 0.473. The analysis identifies bottlenecks between perception, formulation, and deduction, weak multi-law state transfer, and a persistent Sim-to-Real gap.
No heat snapshots are available in the last 24 hours.