The paper introduces QR-Structured Thermal Triggers (QR-STT), a training-free, black-box attack framework for infrared vision-language models. It preserves the functional regions of a QR pattern while assigning each module a cold, neutral, or hot thermal state. A gradient-free search jointly optimizes module topology and rendering parameters, including position, scale, rotation, intensity, blur, and roundness. According to the abstract, QR-STT redirects image-text alignment toward attacker-selected concepts and transfers from classification to image captioning and visual question answering, producing target-consistent semantic drift.
No heat snapshots are available in the last 24 hours.