This pilot study uses Natural Language Autoencoders to verbalize layer-20 residual-stream activations in Qwen2.5-7B-Instruct. It examines 30 prompts organized as 15 matched Colombian-Spanish and English pairs, covering explicit Colombian cues, implicit cues, and neutral controls. Activations are sampled across four positional quartiles to investigate whether nationality, socioeconomic status, or stereotype-related information appears internally before reaching the model output. The authors report descriptive rates and qualitative observations rather than statistically powered effects, so the work should be read as an exploratory interpretability and bias-evaluation study.
No heat snapshots are available in the last 24 hours.