The paper proposes a Structured Sparse Autoencoder (S²AE) for improving concept consistency in vision-language models. It groups image patches using Transformer attention similarity and spatial proximity, then combines intra-group group sparsity with inter-group exclusive sparsity. On Qwen2.5-VL-7B-Instruct, the method reports a 6.06% average gain in semantic alignment measured by mIoU, a representational-efficiency score of 60.81 based on lower l0 norm, and explained variance above 99%. Cross-modal features also show gains in semantic consistency and monosemanticity.
No heat snapshots are available in the last 24 hours.