← Back to index
VegaLaboratories·HybridOptiQ Research
Research complete · Not yet implemented
Quantization methods · Stage 2, rounding · beyond GPTQ

YAQA for MLX — whole-model-aware rounding

GPTQ compensates each tensor's own rounding error, locally. YAQA compensates for how that tensor's error propagates all the way to the model's final output — a genuine second-order estimate of the whole model's KL divergence, not just one layer's own reconstruction error.

Why whole-model awareness matters

GPTQ sees one tensor. YAQA sees the whole model.

Both methods round weights using a Hessian-informed compensation instead of naive round-to-nearest. The difference is what that Hessian is actually measuring.

What each method's compensation actually sees

GPTQ YAQA layer N-1 (outside view) this tensor local Hessian (this layer's own input activations) layer N+1... (outside view) Compensation scope: within this weight matrix only — no visibility past this one tensor layer 1 this tensor layer N final output KL Compensation scope: Kronecker-factored sketch of how THIS tensor's error reaches the final prediction

GPTQ's Hessian is built from one layer's own input activations — it has no information about anything downstream. YAQA's Hessian sketch is built with respect to the full model's output KL divergence, so the same rounding decision is informed by how it actually affects the final prediction, not just this layer's local reconstruction error.

“YAQA is provably better than GPTQ/LDLQ and empirically reduces the error by ≈30% over these methods.” — Cornell-RelaxML. YAQA: Model-Preserving Adaptive Rounding, arXiv:2505.22988, abstract (verbatim).
The actual mechanism

How Sketch B works — the version this port targets

The paper offers two ways to build the Hessian sketch. Sketch A needs an iterative power-iteration bootstrap; Sketch B is a single round from a clean start — confirmed directly from the paper's own text as both cheaper and empirically better than Sketch A. That's the one this project's porting plan targets.

Sketch B — end to end

Calibration batch real text windows through the model Gradient collection ∇W*ℓ custom backward pass mx.custom_function / mx.grad Sketch B (one round) H_I ← E[∇ᵗ∇] / m H_O ← E[∇∇ᵗ] / n no LDLQ bootstrap needed — single round from identity confirmed better than Sketch A Rounding update W = Q(vec(W*) + vec(Δ)(L_O⊗L_I−I)) fixed-point correction, whole-model-aware ✓ quantized weights

Equations verified directly from the paper's own pages 4-6, not summarized: Sketch B (Eq. 12-13) and the fixed-point rounding update (Eq. 5-6). Real cost, from the paper's own text: as cheap as ~1 GPU-hour-equivalent at ~2K sequences while still reaching comparable results — the real cost is tunable, not fixed.

Feasibility, verified not assumed

MLX has the primitives. The port is real work, not a wall.

The first pass on this question was too pessimistic — a deeper read of both YAQA's real PyTorch/CUDA implementation and MLX's actual documented API corrected it.

mx.custom_functiondirect match for torch.autograd.Function
mx.grad / mx.vjpdirect match for gradient hooks
~30%real published error reduction vs. GPTQ/LDLQ
weeksnot months — realistic estimate, both research passes agree

Staged port — each step gates the next, no chaining on faith

STEP 1 ✓

Isolated gradients

mx.custom_function confirmed to give simultaneous (input, grad_output) access on a real layer shape. Passing.

STEP 2 ✓

Synthetic Sketch B

H_I/H_O verified two independent ways, exact symmetry, exact round-trip storage. Passing.

STEP 3

Rounding step

Fixed-point update (Eq. 5-6) using the validated synthetic Hessian. Next up.

STEP 4

One real layer

End-to-end on one real 27B-model layer, compared against naive rounding.

STEP 5

Full model

Only after 1-4 are each independently confirmed correct.

Full command-by-command log, real captured output, and exact resume instructions: scripts/yaqa_port/PORT_LEDGER.md.

Origin, timeline, and today's real compatibility

Where YAQA came from, and whether it runs on MLX today

Authorship, independently confirmed directly from the real downloaded paper's own PDF metadata (not just the announcement): the paper's listed authors are Albert Tseng, Zhaofeng Sun, and Christopher De Sa (Cornell). A public announcement from Together AI accompanied the release — that attribution comes from Together AI's own blog post, not from something this project independently verified in the paper itself.

Timing: the arXiv submission is 2505.22988 — independently corroborated here by the paper's own embedded figures, whose file metadata timestamps read May 14–16, 2025, matching the "late May 2025" release window reported publicly. A wider conference showing (ICML) was reported afterward — that claim is attributed to its source below, not re-checked against icml.cc directly this session.

Is it MLX-compatible today? No — and this one is this project's own directly verified finding, not a secondhand claim: the real, official implementation lives at Cornell-RelaxML/yaqa (cloned locally this session), built on PyTorch/CUDA-specific mechanisms (torch.autograd.Function, torch.distributed). MLX has no native YAQA support, and there is no existing MLX port anywhere — confirmed by a dedicated search this session, not inferred from silence. That is exactly the gap the staged port above exists to close, methodically, not by "saving as MLX."

GPTQ vs. YAQA — five real differences

DimensionGPTQYAQA
Optimization targetMinimize error inside one isolated layerMinimize the whole model's final-output KL divergence
Error compounding across layersNot modeled — each layer assumes the others are fineModeled directly — this tensor's effect on the final prediction is the objective
Real published error reductionBaseline≈30% lower than GPTQ/LDLQ (verbatim, paper's own abstract)
Hessian approximationLocal, per-layer empirical HessianKronecker-factored whole-model Hessian sketch (Sketch A/B)
Bit-width couplingTied to standard affine INT formatsDesigned to be quantizer-agnostic

The error-reduction and Kronecker-sketch rows are this project's own verified claims (real quote and real equations, pulled directly from the paper PDF, cited earlier on this page). The bit-width/quantizer-agnostic framing is reported from Together AI's announcement and is consistent with, but not separately re-derived from, this project's own reading of the paper.

Sources: arXiv:2505.22988 · Cornell-RelaxML/yaqa · Together AI announcement · OpenReview

© Hakim Ghelab, VegaLaboratories LTD. This page documents original research and methods developed for this project — including domain-stratified calibration, cascaded sensitivity measurement, the Pareto/MILP allocator lineage, and the YAQA/MLX feasibility and porting plan described here. All rights reserved.