GPTQ compensates each tensor's own rounding error, locally. YAQA compensates for how that tensor's error propagates all the way to the model's final output — a genuine second-order estimate of the whole model's KL divergence, not just one layer's own reconstruction error.
Both methods round weights using a Hessian-informed compensation instead of naive round-to-nearest. The difference is what that Hessian is actually measuring.
GPTQ's Hessian is built from one layer's own input activations — it has no information about anything downstream. YAQA's Hessian sketch is built with respect to the full model's output KL divergence, so the same rounding decision is informed by how it actually affects the final prediction, not just this layer's local reconstruction error.
The paper offers two ways to build the Hessian sketch. Sketch A needs an iterative power-iteration bootstrap; Sketch B is a single round from a clean start — confirmed directly from the paper's own text as both cheaper and empirically better than Sketch A. That's the one this project's porting plan targets.
Equations verified directly from the paper's own pages 4-6, not summarized: Sketch B (Eq. 12-13) and the fixed-point rounding update (Eq. 5-6). Real cost, from the paper's own text: as cheap as ~1 GPU-hour-equivalent at ~2K sequences while still reaching comparable results — the real cost is tunable, not fixed.
The first pass on this question was too pessimistic — a deeper read of both YAQA's real PyTorch/CUDA implementation and MLX's actual documented API corrected it.
mx.custom_function confirmed to give simultaneous (input, grad_output) access on a real layer shape. Passing.
H_I/H_O verified two independent ways, exact symmetry, exact round-trip storage. Passing.
Fixed-point update (Eq. 5-6) using the validated synthetic Hessian. Next up.
End-to-end on one real 27B-model layer, compared against naive rounding.
Only after 1-4 are each independently confirmed correct.
Full command-by-command log, real captured output, and exact resume
instructions: scripts/yaqa_port/PORT_LEDGER.md.
Authorship, independently confirmed directly from the real downloaded paper's own PDF metadata (not just the announcement): the paper's listed authors are Albert Tseng, Zhaofeng Sun, and Christopher De Sa (Cornell). A public announcement from Together AI accompanied the release — that attribution comes from Together AI's own blog post, not from something this project independently verified in the paper itself.
Timing: the arXiv submission is 2505.22988 — independently corroborated here by the paper's own embedded figures, whose file metadata timestamps read May 14–16, 2025, matching the "late May 2025" release window reported publicly. A wider conference showing (ICML) was reported afterward — that claim is attributed to its source below, not re-checked against icml.cc directly this session.
Is it MLX-compatible today? No — and this one is this project's own
directly verified finding, not a secondhand claim: the real, official implementation lives at
Cornell-RelaxML/yaqa
(cloned locally this session), built on PyTorch/CUDA-specific mechanisms
(torch.autograd.Function, torch.distributed).
MLX has no native YAQA support, and there is no existing MLX port anywhere — confirmed by a dedicated
search this session, not inferred from silence. That is exactly the gap the staged port above exists
to close, methodically, not by "saving as MLX."
| Dimension | GPTQ | YAQA |
|---|---|---|
| Optimization target | Minimize error inside one isolated layer | Minimize the whole model's final-output KL divergence |
| Error compounding across layers | Not modeled — each layer assumes the others are fine | Modeled directly — this tensor's effect on the final prediction is the objective |
| Real published error reduction | Baseline | ≈30% lower than GPTQ/LDLQ (verbatim, paper's own abstract) |
| Hessian approximation | Local, per-layer empirical Hessian | Kronecker-factored whole-model Hessian sketch (Sketch A/B) |
| Bit-width coupling | Tied to standard affine INT formats | Designed to be quantizer-agnostic |
The error-reduction and Kronecker-sketch rows are this project's own verified claims (real quote and real equations, pulled directly from the paper PDF, cited earlier on this page). The bit-width/quantizer-agnostic framing is reported from Together AI's announcement and is consistent with, but not separately re-derived from, this project's own reading of the paper.
Sources: arXiv:2505.22988 · Cornell-RelaxML/yaqa · Together AI announcement · OpenReview