Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Individual Parameters in Weight-Sparse Transformers Appear Interpretable

First seen · 7/3/2026, 01:15 PMLatest activity · 7/3/2026, 01:15 PM

This paper asks whether an individual neural-network weight can be understood globally across the training distribution, rather than only within a behavior-specific circuit. The authors introduce an automated LLM pipeline that describes when ablating a weight changes predictions and validates the description on held-out text. Across two sparse and two dense Transformers, sparse models contain a higher fraction of interpretable weights. After unreliable descriptions are removed, the gap widens. The reported fraction of sparse-model weights receiving one short, generalizing description ranges from 12% to 31%.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/3, 01:15 PMnot independentRepresentative
    Individual Parameters in Weight-Sparse Transformers Appear Interpretable