KronQ proposes a post-training quantization framework that incorporates gradient covariance into the Kronecker-factored Hessian approximation. Unlike GPTQ-style methods that rely primarily on input activation statistics, it models both input and output-side importance. The method applies bidirectional incoherence processing, extending random rotations to the output dimension, and derives a mixed-precision sensitivity metric from activation and gradient Hessian traces. The paper reports that for 2-bit weight-only quantization of LLaMA-3-70B, KronQ reaches 7.93 WikiText-2 perplexity, while GPTQ and GPTAQ diverge or produce degenerate quantizations above 2,000 perplexity.
No heat snapshots are available in the last 24 hours.