Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

First seen · 7/29/2026, 08:40 PMLatest activity · 7/29/2026, 08:40 PM

This paper tests whether 4-bit post-training weight quantization is effectively lossless for multi-turn, tool-calling agents. Across eight τ²-bench cells with 456 episodes each, standard scores show no change surviving multiple-comparison correction. Process-level logs tell a different story: quantization reportedly amplifies existing failures by up to 2.5×, adding 17.6 percentage points of errors per task while creating almost no novel failures. The benchmark’s ten-error allowance masks this damage; reducing it to two errors reveals a 17-point score gap in the affected cell.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/29, 08:40 PMnot independentRepresentative
    Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents