The paper introduces V2N (Video to Notes), a complete visual piano transcription system that predicts note onset, offset, key hold, and velocity from video. A shared temporal backbone feeds task-specific heads, while training uses per-frame supervision instead of supervising only the center frame of each window. According to the supplied abstract, multi-task supervision enables offset and velocity prediction and also improves onset accuracy. Longer temporal context provides additional gains, and V2N reports state-of-the-art results on the PianoVAM and R3 benchmarks.
No heat snapshots are available in the last 24 hours.