This paper introduces Multimodal Voice Activity Projection (MM-VAP), extending audio-only future voice-activity prediction to synchronized audio-visual inputs for turn-taking in social robots. The method adapts pretrained audio-visual speech encoders with Low-Rank Adaptation, independently encodes speakers, and uses inter-speaker attention to model conversational relations. A semantic consistency loss regularizes the 256-state output space around higher-level dialogue activity patterns. Experiments on NoXi, NoXi+J, and the Haru EDR corpus report improvements over current baselines for some turn-taking events, especially in mediation-oriented human-robot interaction. The abstract does not provide exact metric values or statistical significance.
No heat snapshots are available in the last 24 hours.