V-DEAL studies a counterintuitive safety failure in Video Large Language Models: harmful videos paired with benign queries produce higher attack success rates than the same videos paired with explicitly harmful queries. Across six Video LLMs and three public benchmarks, the models recognized harmful video content with over 81% accuracy, yet the average attack success rate remained 48.33% for harmful-video/benign-query pairs. Hidden-state analysis indicated that visual understanding activates a weaker refusal tendency than textual understanding. The proposed prompt-injection intervention reduced attack success rates by an average of 48.24 percentage points.
No heat snapshots are available in the last 24 hours.