PhysMRV is a training-free framework for physical plausibility reasoning in video-language models. It converts training videos into a Hierarchical Memory Bank with three structured levels: scene descriptions, physical-event graphs representing object interactions and causal structure, and summaries of reusable physics rules. At inference time, the framework retrieves physically relevant evidence and uses it to guide a frozen VLM in verifying whether a video is physically plausible. The authors evaluate it across ImplausiBench, IntPhys2, and GRASP Level 2 with multiple state-of-the-art VLMs, reporting consistent improvements over direct prompting without fine-tuning or parameter updates.
No heat snapshots are available in the last 24 hours.