Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Multimodal Voice Activity Projection for Turn-Taking in Social Robots

First seen · 7/8/2026, 07:34 PMLatest activity · 7/8/2026, 07:34 PM

This paper introduces Multimodal Voice Activity Projection (MM-VAP), extending audio-only future voice-activity prediction to synchronized audio-visual inputs for turn-taking in social robots. The method adapts pretrained audio-visual speech encoders with Low-Rank Adaptation, independently encodes speakers, and uses inter-speaker attention to model conversational relations. A semantic consistency loss regularizes the 256-state output space around higher-level dialogue activity patterns. Experiments on NoXi, NoXi+J, and the Haru EDR corpus report improvements over current baselines for some turn-taking events, especially in mediation-oriented human-robot interaction. The abstract does not provide exact metric values or statistical significance.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/8, 07:34 PMnot independentRepresentative
    Multimodal Voice Activity Projection for Turn-Taking in Social Robots