Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

First seen · 8/4/2026, 12:00 PMLatest activity · 8/4/2026, 12:00 PM

WorldExam introduces a hierarchical diagnostic benchmark for controllable video generation models viewed as world models. It evaluates four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. The benchmark contains 1,474 cases across eight tasks and evaluates 20 representative models spanning camera-, action-, and language-driven paradigms. The reported results show a capability split: camera-driven models handle camera control well but lack interaction interfaces; action-driven models control subjects more precisely while often leaving the surrounding world unresponsive; language-driven models support interaction better but follow complex controls less faithfully. No evaluated model combines broad task coverage with consistently strong performance.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/4, 12:00 PMnot independentRepresentative
    WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity