Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Multi-Task Multi-Frame Visual Piano Transcription

First seen · 8/4/2026, 04:00 AMLatest activity · 8/4/2026, 04:00 AM

The paper introduces V2N (Video to Notes), a complete visual piano transcription system that predicts note onset, offset, key hold, and velocity from video. A shared temporal backbone feeds task-specific heads, while training uses per-frame supervision instead of supervising only the center frame of each window. According to the supplied abstract, multi-task supervision enables offset and velocity prediction and also improves onset accuracy. Longer temporal context provides additional gains, and V2N reports state-of-the-art results on the PianoVAM and R3 benchmarks.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/4, 04:00 AMnot independentRepresentative
    Multi-Task Multi-Frame Visual Piano Transcription