AVE-Compass evaluates audio-video editing as a coordinated capability rather than treating visual and audio edits separately. The benchmark contains 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It measures instruction following, fidelity preservation, realism, and editing intent through checklist-based MLLM judging, a dedicated realism rubric, and automated cross-modal, video, and audio metrics. The authors report that current state-of-the-art models still struggle with cross-modal edits and preserving non-target content. They also introduce AVE-Agent, a modular framework using dependent task decomposition, self-reflection, and evaluator feedback.
No heat snapshots are available in the last 24 hours.