Why compensated rounding beats naive quantization at low bit-widths, and how this project's own integration goes further than Apple's shipped mlx_lm GPTQ: reading a real per-tensor mixed-precision plan instead of one flat bit-width, and using this project's own calibration instead of a fixed public dataset.
Every weight in a neural network has to land on one of a small number of grid points once quantized. The question is how you choose which point — and at 3-4 bits, that choice is the difference between a usable model and a broken one.
Round-to-nearest (RTN) rounds every weight independently, blind to the weights around it. GPTQ instead measures each weight's rounding error and folds it into the weights it hasn't rounded yet — a local, Hessian-informed compensation, so errors don't compound unchecked across a layer.
Pattern confirmed directly from the real paper (Frantar et al., "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers," ICLR 2023, arXiv:2210.17323, Figure 1): 4-bit RTN spikes sharply and unpredictably on the OPT model family while 4-bit GPTQ tracks close to the FP16 baseline across every model size tested. Same shape shown for 3-bit on BLOOM.
Apple's own mlx_lm.quant.gptq is real, working GPTQ — the
Cholesky-based compensation math is genuine and this project reuses it directly, unmodified. What
differs is everything around that math: what bit-width gets applied, where the calibration
data comes from, which bit-widths can even be packed, and whether the model's other real components
survive the process at all.
Every row above except the shaded "shared core" box was verified directly against the real installed source this session — not assumed. The two implementations differ in everything surrounding the actual rounding math, not the rounding math itself.