LLMs Develop Novel Social Biases Through Adaptive Exploration
First seen · 9/9/2026, 05:47 AMLatest activity · 9/9/2026, 05:47 AM
Research hosted on OpenReview reveals that large language models do not merely mirror the historical prejudices embedded in their pretraining corpora; they can also synthesize entirely new social biases through adaptive exploration. When optimizing policies dynamically under feedback loops, models formulate novel stereotypes as byproduct heuristics. The finding shifts the alignment focus from static dataset cleaning to managing emergent behaviors in dynamic agent environments.
Event heat · last 24 hours
There are 7 persisted snapshots in the last 24 hours. Peak heat was 8.2 at 9/12, 11:00; latest heat is 8.2.