The paper introduces SVF-CR, a multimodal framework for recognizing ambivalence and hesitancy. It extracts whole-video and cropped-face segment tokens using the same temporal partition, then applies intra-modal self-attention and bidirectional visual-facial cross-attention before constructing consistency- and discrepancy-based evidence. Temporal modeling and attention pooling are applied to the visual-facial stream, while textual and acoustic features receive lighter refinement and are combined through pairwise evidence fusion. On the public evaluation split of the BAH challenge, SVF-CR reports a public macro-F1 of 0.7156, outperforming the paper’s referenced global visual-face fusion and synchronized evidence baselines.
No heat snapshots are available in the last 24 hours.