This paper argues that generative world models used to judge robotic or autonomous-driving policies must themselves be accredited before their verdicts can count as assurance evidence. Drawing on Verification, Validation & Accreditation (VV&A), Safety of the Intended Functionality (SOTIF), and scenario-based testing, it proposes an embodiment-agnostic L0-L4 admissibility ladder. In an autonomous-driving instantiation involving two driving world models, the model with better visual-generation quality at L0 ranked worse on action-following at L1-L2, showing that visual realism is not a reliable proxy for closed-loop decision validity.
No heat snapshots are available in the last 24 hours.