Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

StabilityBench: Benchmarking Instability in LLMs

First seen · 7/18/2026, 12:09 AMLatest activity · 7/18/2026, 12:09 AM

StabilityBench converts single-turn benchmark queries into multi-turn interaction histories using realistic user simulations, including demographic proxies and sycophantic baits, while preserving the original task intent. The authors apply it to four benchmarks covering mathematical reasoning, health question-answering, and safety, evaluating nine large language models. They report consistent instability under these injections, with considerable degradation on three of the four benchmarks. The paper also introduces StabilityBench-Mini, a size-preserving variant that samples across diversification axes without increasing evaluation costs.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/18, 12:09 AMnot independentRepresentative
    StabilityBench: Benchmarking Instability in LLMs