Objects as Audio-Visual Modal Sound Fields
The paper introduces Audio-Visual Modal Sound Field (AV-MSF), an object-level acoustic representation reconstructed from multi-view images and only a few impact recordings. It combines 3D Gaussian Splatting with dense 3D visual features to obtain a geometry-aware prior, while modeling the impact sound field with compact, physically meaningful modal parameters. According to the abstract, experiments on two real-world datasets achieve state-of-the-art impact-sound rendering against physics-based and data-driven baselines. The representation also supports contact localization and object sound editing.
Why it's worth reading
It connects few-shot sound reconstruction with an editable 3D object representation, with potential relevance to robotics, immersive interaction, and audiovisual generation; the paper’s dataset and quantitative details still require verification.