This paper adapts Funder’s person-situation-behavior framework to analyze personality-like structure in large language models. It treats the person component as internal personality-related representations, situations as contexts that elicit trait-relevant responses, and behavior as response patterns in broader social tasks. Using contrastive behavior pairs and sparse autoencoder decomposition, the authors identify features associated with opposing trait poles. Feature-level interventions produce bidirectional shifts across diverse situations while preserving response validity. Social-intelligence evaluations show benefit-tradeoff patterns that the authors describe as consistent with human personality research. The abstract does not specify the evaluated models, datasets, traits, or effect sizes.
No heat snapshots are available in the last 24 hours.