NANQ is a mixed-precision, non-uniform quantization framework designed around the hardware noise floor of analog compute-in-memory systems. It derives adaptive quantization density from magnitude-dependent noise measured on an eFlash CIM array and selects layer-wise bit widths using each layer’s precision saturation point. According to the supplied abstract, on-chip eFlash CIM SoC experiments show an 8.05 percentage-point average vision accuracy improvement and a 54.7% average language-model perplexity reduction over PowerQuant under 2-bit weight-magnitude quantization. Mixed-precision configurations reportedly retain most additional-precision gains using 3.2–3.8 equivalent bits.
No heat snapshots are available in the last 24 hours.