The paper presents a token-centric dual-view framework for mammography classification using a frozen vision transformer. Dedicated fusion tokens mediate bidirectional communication between craniocaudal (CC) and mediolateral oblique (MLO) views through cross-attention, while fusion modules are inserted at multiple transformer depths. This design aims to preserve view-specific information while progressively propagating complementary evidence. Experiments on VinDr-Mammo and CMMD reportedly outperform linear probing, prompt-only adaptation, and conventional fusion baselines. On VinDr-Mammo BI-RADS classification, the method reaches 50.40% F1 and 0.8090 AUC, with a 0.10 AUC gain over a dual-view fusion baseline in the binary setting.
No heat snapshots are available in the last 24 hours.