This paper tests whether pretrained audio embeddings encode phylogenetic structure that was not part of their training objectives. Across 32 marine mammal species and 20 bird species, general-purpose models recovered evolutionary relationships from vocalizations. For cetaceans, CLAP and BEATs-bio reached Mantel correlations of 0.82, while AST reached 0.74; handcrafted 105-dimensional MFCC features showed no significant signal. The result remained after dimensionality matching and partial control for dominant frequency. On birds, AST and CLAP again outperformed the domain-focused BirdNET and BEATs-bio, suggesting that domain-specific pretraining alone does not guarantee stronger biological representations.
No heat snapshots are available in the last 24 hours.