This paper presents a fine-grained multimodal framework for binary depression detection from audio-visual data. It combines a temporal encoder with a mutual transformer for cross-modal fusion. Its main contribution is Binary Advantage-weighting Ranking Loss, which has two components: Advantage-weighted Separation mines difficult pairs using pairwise prediction differences, while Advantage-weighted Compactness reduces within-class variance around class centers. The authors report that the method reconstructs latent ordinal structure and achieves state-of-the-art performance on the D-vlog and LMVD datasets, although the abstract provides no numerical results.
No heat snapshots are available in the last 24 hours.