This paper describes DS@GT ARC’s BirdCLEF+ 2026 system for multi-label animal-vocalization detection in Pantanal soundscapes. Its supervised baseline ensembles a frozen Perch v2 backbone, a trained HGNetV2-B0 sound-event-detection model, and a non-bird prototypical head. The system achieved a private leaderboard score of 0.936 and rank 1894 within a 90-minute CPU budget. The authors then evaluate whether token-based representations can compete, comparing neural audio-codec representations and semantic foundation-model embeddings with two bioacoustic specialists and four AudioSet-trained token encoders.
No heat snapshots are available in the last 24 hours.