O-VAD introduces a training-free, domain-knowledge-free agentic framework for industrial video anomaly detection. Instead of judging clips only at the scene level, it tracks detected objects across space and time, models their state evolution and transformations, and reasons over object-wise temporal trajectories to ground anomalous objects in specific frames. The abstract reports experiments on three IVAD datasets, claiming improvements over frontier vision-language models, agentic frameworks, and traditional VAD methods fine-tuned on each dataset. It also produces interpretable reports describing anomaly processes and types, although detailed metrics and implementation evidence are not available in the supplied abstract.
No heat snapshots are available in the last 24 hours.