I2VShield is a proactive privacy-defense method for image-to-video generation models based on Diffusion Transformers (DiTs). It combines a text-adaptive perturbation generation framework with adversarial learning to reduce the computational burden of creating protective perturbations while preserving visual imperceptibility. Its untargeted Multimodal Attention Disruption (MAD) attack pushes internal attention features away from their clean states, aiming to impair generated-video quality and spatiotemporal coherence. The abstract reports competitive protection across multiple datasets and mainstream DiT-based I2V models, but does not provide quantitative results, model names, or resource measurements.
No heat snapshots are available in the last 24 hours.