This paper studies whether large language models encode the hierarchical similarity structure among languages. The authors report that model representations largely recover the Indo-European language family tree, clustering languages from the same subfamilies in latent space. They also report a correlation between the degree of this structural alignment and performance on XNLI, a multilingual natural language inference benchmark. The central claim is that models representing similar languages similarly are better able to generalize from one language to another, extending similarity-based accounts of generalization to multilingual LLMs.
No heat snapshots are available in the last 24 hours.