This paper evaluates whether vision models organize colors in a human-like way beyond standard geometric references such as CIELAB. It introduces an evaluation framework based on 86 graded color categories fitted to human survey data, measuring category boundaries, category compactness, and graded alignment unexplained by color geometry. Across eleven Vision Transformer encoders, category-level performance is broadly similar, but graded alignment varies substantially. Masked Autoencoders show the strongest beyond-geometry alignment, with non-overlapping confidence intervals relative to the other encoders. Layer-wise analysis suggests masked reconstruction preserves this structure toward the output. On natural images, MAEs encode surface color globally, while language-supervised models associate color more strongly with foreground objects.
No heat snapshots are available in the last 24 hours.