This paper studies adversarial vulnerability in transformer-based vision-language models through the spectral structure of their intermediate linear transformations. It introduces SSGRA, a white-box spectral-subspace-guided attack that aligns intermediate representations with the subspace spanned by bottom right singular vectors. According to the abstract, experiments show improved attack effectiveness over existing baselines. The authors frame the method as both an attack technique and a spectral interpretation of why VLMs are vulnerable, with potential implications for robustness improvements. Specific models, datasets, perturbation budgets, baselines, and numerical results are not included in the supplied abstract.
No heat snapshots are available in the last 24 hours.