The paper introduces NeuroCogMap, a cognitive-neuroscience-inspired framework for organizing internal LLM features into functional parcels linked to interpretable functions, capabilities, and a cognitive hierarchy. According to the abstract, the resulting organization is stable and semantically coherent, with some cross-model conservation. It associates hallucination, bias, refusal failure, and sycophancy with distinct representational or behavioral-control disruptions. The framework also reportedly improves prediction of human cortical responses during naturalistic language comprehension, especially in higher-order association cortex, and reveals latent strategies relevant to models of human decision-making.
No heat snapshots are available in the last 24 hours.