This paper studies hypernetworks as a train-time mechanism for injecting factual knowledge into large language models. Given a corpus of facts, the hypernetwork generates a fixed LoRA adapter for a target model. The authors introduce MegaWikiQA, a dataset containing tens of millions of multi-hop question-answer examples across 39 domains, derived from Wikidata5M. They analyze how hypernetwork depth, width, and target-model size affect loss, reasoning accuracy, and out-of-distribution generalization. The reported results show broadly predictive power-law scaling across architecture axes, with increasingly large hypernetworks achieving reliable OOD generalization and steeper scaling exponents than LoRA and full fine-tuning in the reported OOD evaluations.
No heat snapshots are available in the last 24 hours.