This paper tests whether 4-bit post-training weight quantization is effectively lossless for multi-turn, tool-calling agents. Across eight τ²-bench cells with 456 episodes each, standard scores show no change surviving multiple-comparison correction. Process-level logs tell a different story: quantization reportedly amplifies existing failures by up to 2.5×, adding 17.6 percentage points of errors per task while creating almost no novel failures. The benchmark’s ten-error allowance masks this damage; reducing it to two errors reveals a 17-point score gap in the affected cell.
No heat snapshots are available in the last 24 hours.