This paper presents a parameter-efficient CLIP adaptation framework for long-term animal re-identification. Its main contribution is continuous metadata conditioning, which injects numerical attributes into prompt representations without discretizing them into textual categories. The framework combines low-rank visual adaptation, prompt-based supervision, and cross-modal alignment. According to the abstract, experiments on a seven-year longitudinal fish dataset and multiple wildlife benchmarks cover closed-set, open-set, and time-aware evaluation protocols, showing improved robustness to morphological and seasonal shifts. Metadata is used during training but is not required at inference, preserving a purely visual pipeline.
No heat snapshots are available in the last 24 hours.