The paper proposes SFT+RL, a two-stage robust unsupervised domain adaptation framework built on CLIP’s pretrained visual encoder. Supervised fine-tuning first adversarially trains a linear classifier on labeled source data with PGD perturbations while partially unfreezing the projection layer. A reinforcement-learning stage progressively selects target-domain pseudo-labels using a decaying confidence threshold, then trains on mixed clean and adversarial batches. On OfficeHome, PACS, and VisDA, the authors report average gains of 10.2% in clean accuracy and 15.8% in adversarial robustness. These claims are currently supported only by the supplied abstract.
No heat snapshots are available in the last 24 hours.