← Back to index
VegaLaboratories·HybridOptiQ Research
Quantization methods · Stage 2, rounding

GPTQ for MLX — multi-precision, plan-aware rounding

Why compensated rounding beats naive quantization at low bit-widths, and how this project's own integration goes further than Apple's shipped mlx_lm GPTQ: reading a real per-tensor mixed-precision plan instead of one flat bit-width, and using this project's own calibration instead of a fixed public dataset.

Why rounding quality matters

Naive rounding breaks at low bit-widths. GPTQ doesn't.

Every weight in a neural network has to land on one of a small number of grid points once quantized. The question is how you choose which point — and at 3-4 bits, that choice is the difference between a usable model and a broken one.

Round-to-nearest (RTN) rounds every weight independently, blind to the weights around it. GPTQ instead measures each weight's rounding error and folds it into the weights it hasn't rounded yet — a local, Hessian-informed compensation, so errors don't compound unchecked across a layer.

Perplexity vs. model size, RTN vs. GPTQ vs. FP16 — illustrative pattern

Redrawn to match the real published shape, not the source's exact pixel values
worse better PPL model size → 4-bit RTN — sharp failures 4-bit GPTQ — near FP16 FP16 baseline

Pattern confirmed directly from the real paper (Frantar et al., "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers," ICLR 2023, arXiv:2210.17323, Figure 1): 4-bit RTN spikes sharply and unpredictably on the OPT model family while 4-bit GPTQ tracks close to the FP16 baseline across every model size tested. Same shape shown for 3-bit on BLOOM.

“GPTQ can quantize GPT models with 175 billion parameters in approximately four GPU hours, reducing the bitwidth down to 3 or 4 bits per weight, with negligible accuracy degradation relative to the uncompressed baseline.” — Frantar, Ashkboos, Hoefler, Alistarh. GPTQ, ICLR 2023, abstract (verbatim).
Apple's native GPTQ vs. this project's integration

Same trusted core math. A much more capable pipeline around it.

Apple's own mlx_lm.quant.gptq is real, working GPTQ — the Cholesky-based compensation math is genuine and this project reuses it directly, unmodified. What differs is everything around that math: what bit-width gets applied, where the calibration data comes from, which bit-widths can even be packed, and whether the model's other real components survive the process at all.

Architecture comparison — native mlx_lm GPTQ vs. this project's integration

APPLE'S NATIVE mlx_lm GPTQ THIS PROJECT'S INTEGRATION Calibration source Fixed public gist URL — same text for every model No relationship to what the model will actually see Calibration source This project's own domain-stratified calibration (6 domains, balanced — or flat, for explicit comparison) Bit-width granularity ONE global bits value for the entire model Ignores any per-tensor sensitivity finding Bit-width granularity PER-TENSOR — read live from plan.json, 497 real independent decisions from the Pareto/MILP allocator Supported bit-widths {2, 4, 8} only — word-aligned packing scheme Crashes on 5/6-bit tensors, a real plan gap Supported bit-widths {2, 3, 4, 5, 6, 8} — general packer, verified bit-exact against MLX's own real quantizer, 700+ configurations MTP / vision sidecars Not handled — a separate concern entirely Left to whatever wraps this function, if anything MTP / vision sidecars Reattached automatically via this project's own already -trusted OptiQ functions — not silently dropped Interruption / resume None — a killed run loses all progress save() only runs once, at the very end Interruption / resume Per-tensor checkpoint cache — a stopped run resumes exactly where it left off, tensor by tensor SHARED, UNCHANGED CORE Damped-Cholesky Hessian inversion + column-by-column compensated rounding — Apple's real, trusted math, reused directly and never reimplemented

Every row above except the shaded "shared core" box was verified directly against the real installed source this session — not assumed. The two implementations differ in everything surrounding the actual rounding math, not the rounding math itself.

497tensors in a real plan
371now GPTQ-rounded (~75%)
136GPTQ-rounded before this fix (~27%)
700+packing test configs, bit-exact
© Hakim Ghelab, VegaLaboratories LTD. This page documents original research and methods developed for this project — including domain-stratified calibration, cascaded sensitivity measurement, the Pareto/MILP allocator lineage, and the per-tensor-plan-aware GPTQ/MLX integration described here. All rights reserved.