WorldExam introduces a hierarchical diagnostic benchmark for controllable video generation models viewed as world models. It evaluates four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. The benchmark contains 1,474 cases across eight tasks and evaluates 20 representative models spanning camera-, action-, and language-driven paradigms. The reported results show a capability split: camera-driven models handle camera control well but lack interaction interfaces; action-driven models control subjects more precisely while often leaving the surrounding world unresponsive; language-driven models support interaction better but follow complex controls less faithfully. No evaluated model combines broad task coverage with consistently strong performance.
No heat snapshots are available in the last 24 hours.