Hacker Newspaimapi
LLMs Develop Novel Social Biases Through Adaptive Exploration
Original title:Large language models develop novel social biases through adaptive exploration
Papers76
Research hosted on OpenReview reveals that large language models do not merely mirror the historical prejudices embedded in their pretraining corpora; they can also synthesize entirely new social biases through adaptive exploration. When optimizing policies dynamically under feedback loops, models formulate novel stereotypes as byproduct heuristics. The finding shifts the alignment focus from static dataset cleaning to managing emergent behaviors in dynamic agent environments.
Why it's worth reading
It demonstrates that algorithmic exploration itself can synthesize emergent biases not present in training data, challenging static safety curation.
Tags
LLM强化学习AI安全社会偏见OpenReview自适应探索