StabilityBench converts single-turn benchmark queries into multi-turn interaction histories using realistic user simulations, including demographic proxies and sycophantic baits, while preserving the original task intent. The authors apply it to four benchmarks covering mathematical reasoning, health question-answering, and safety, evaluating nine large language models. They report consistent instability under these injections, with considerable degradation on three of the four benchmarks. The paper also introduces StabilityBench-Mini, a size-preserving variant that samples across diversification axes without increasing evaluation costs.
No heat snapshots are available in the last 24 hours.