This paper asks whether an individual neural-network weight can be understood globally across the training distribution, rather than only within a behavior-specific circuit. The authors introduce an automated LLM pipeline that describes when ablating a weight changes predictions and validates the description on held-out text. Across two sparse and two dense Transformers, sparse models contain a higher fraction of interpretable weights. After unreliable descriptions are removed, the gap widens. The reported fraction of sparse-model weights receiving one short, generalizing description ranges from 12% to 31%.
No heat snapshots are available in the last 24 hours.