Fast, Large-Scale Linguistic Phylogenetics via Self-Supervised Lexical Representations
Original title:Self-Supervised Lexical Representation Learning for Fast, Large-Scale Phylogenetic Inference
Reconstructing global language family trees has long been constrained by labor-intensive manual cognacy judgments and prohibitive computational costs. This paper presents a fully self-supervised dual-contrastive learning framework that derives lexical representations directly from unaligned, IPA-transcribed wordlists. Across 3,399 language varieties, the method infers a global phylogenetic tree matching Glottolog benchmark quality within minutes on a standard notebook GPU, while also capturing diachronic concept stability without supervision.
Why it's worth reading
It demonstrates that global-scale linguistic phylogenetics across thousands of languages can be accurately and automatically reconstructed on a single consumer GPU without cognate annotations.