Port begins
No MLX port of YAQA existed anywhere — confirmed by a direct GitHub search. Steps 1–4 built and verified on a real layer of the real 27B model: YAQA’s correction cuts the Hessian-weighted error 75% versus naive rounding.
2026-09-01
A real, staged port of YAQA (Yet Another Quantization Algorithm — model-preserving adaptive rounding, Cornell-RelaxML) from its original PyTorch/CUDA implementation to Apple’s MLX framework.
Author: Hakim Ghelab, VegaLaboratories LTD
Scroll or drag sideways — every stop below traces back to a real log, commit, or doc in this repo. Nothing here is aspirational.
No MLX port of YAQA existed anywhere — confirmed by a direct GitHub search. Steps 1–4 built and verified on a real layer of the real 27B model: YAQA’s correction cuts the Hessian-weighted error 75% versus naive rounding.
Generalizing to the whole model surfaced three real, structural findings: 52% of tensors have no valid Hadamard factor, the real deployment format is trellis-VQ not affine, and holding every tensor’s Hessian at once needs 557GB. Fixed with a hand-derived rotation-aware kernel and memory-bounded batching.
One tensor’s correction blew up 302x instead of improving. Root-caused twice — first to ill-conditioning, then all the way down to BF16 precision loss inside Hessian construction itself. Fixed and independently reverified.
363/363 tensors corrected (362 YAQA + 1 GPTQ), 6.015 real bits/weight, vision and MTP sidecars reattached. Three post-build regressions found by direct comparison against the source model — each root-caused and fixed.
All 7 multi-token-prediction tensors given real YAQA correction — up to 9,675x better than naive on the worst tensor. Full production build relaunched with the fix applied.
A new floor-plus-lexicographic solver mechanism lets the bit-budget optimizer actively protect the tensors the Hessian signal flags as most at risk — not just react to them. Independently code-reviewed; two real bugs found and fixed before shipping.
MMLU, GSM8K, IFEval, BFCL, and HumanEval measured end to end, temperature 0, across 6 real built models — the first quantitative read on the whole pipeline, not just internal error metrics.
YAQA is a whole-model-aware quantization rounding method: instead of
correcting each weight using only local, per-layer information (as GPTQ
does), it collects a real Hessian sensitivity sketch — both an
input-side map (H_I) and an output-side map
(H_O) — from the entire model’s real backward
pass, so the correction accounts for how each weight actually affects
the model’s real output. GPTQ only ever has the equivalent of
H_I; YAQA’s LDLQ_2hess correction algorithm is
two-sided by construction and needs both. No MLX port of YAQA existed
before this work (checked directly via GitHub repo/code search,
2026-09-03 — see PORT_LEDGER.md).
Staged 5-step plan — all 5 steps complete,
independently verified end to end on real layers of a real 27B model
(Qwen3.5-27B hybrid architecture — standard attention + GatedDeltaNet
linear-attention layers), and shipped as a real, loadable production
build (363/363 tensors corrected — 362 YAQA + 1 scoped GPTQ for
lm_head — 6.015 real bits/weight), including a rebuild with
the MTP sidecar’s own real YAQA correction applied. Since then: a new
MILP-aware Hessian mechanism that actively protects the tensors the
Hessian signal flags as most at risk, and this project’s first real
end-to-end benchmark numbers (MMLU / GSM8K / IFEval / BFCL / HumanEval,
temperature 0) across 6 real built models. See the timeline above for
the full, dated story — every stop traces back to a real log, commit, or
doc in this repo.
mx.custom_function) works as required.H_O at all; see
PORT_LEDGER.md Step 2.)LDLQ_2hess
block-sequential rounding algorithm, cross-verified against an
independent NumPy reference.lm_head) have no valid Hadamard factor — a genuine
limitation of the real upstream algorithm, not a porting bug. Second
real finding (2026-09-04): the reference repo’s actual deployment format
is a trellis-VQ codebook with runtime SU/SV/Hadamard rotation via custom
CUDA kernels, not affine quantization. Third real finding (same day):
the rotation-aware forward pass IS derivable and buildable on real MLX
primitives (mx.hadamard_transform,
mx.quantized_matmul) — derived by hand, verified
numerically, built as yaqa_core.YaqaQuantizedLinear.
Production decision: --incoherence none
(unrotated, real mlx_lm-loadable today) is the default;
--incoherence hadamard stays as a real, verified research
branch pending a loader that doesn’t exist yet. Fourth real finding: the
original “wrap all 363 tensors, one pass” architecture needs 557GB
simultaneously — does not fit in 128GB. Fixed via memory-bounded
batching (--hessian-batch-budget-gb) plus running each
batch as a separate process (run_full_yaqa.sh) after a real
Metal resource-handle limit was found to survive
mx.clear_cache(). The real full run has since
completed (2026-09-07): 363/363 tensors corrected, real
production build saved and loadable, three post-build regressions found
by direct comparison against the source model and fixed. The MTP
sidecar’s own real YAQA correction (2026-09-08) followed the
same path — all 7 MTP tensors corrected, up to 9,675x better than naive
on the worst tensor — and the production build was relaunched with it
applied. See PORT_LEDGER.md for the full story and
00_DOCS/HTML/yaqa_pipeline_comparison.html for a visual
comparison.See PORT_LEDGER.md for the full, honest engineering log
— every real bug found and fixed, every real command and its real
output, and a plain-English explanation of what any of this means,
written for a non-technical reader.
3.11.15) with a real MLX install. Verified working versions
(2026-09-03):
mlx 0.32.2mlx-lm 0.31.3numpy 2.4.604_real_layer_test.py,
05_full_model_quantize.py): a local copy of the real BF16
source model this project targets, plus enough RAM to hold the model
and, for Step 5, every target tensor’s Hessian simultaneously during
collection (this project’s own machine: 128GB RAM, confirmed comfortable
for the largest real Hessian in the current plan, ~1.2GB for a
17408×17408 H_I).01_gradient_collection_test.py through
03_rounding_update_test.py,
test_regression.py) need no real model and run in seconds
to a couple of minutes.This project uses a uv-managed Python tool environment.
If you already have mlx/mlx-lm installed some
other way, any Python 3.11 environment with the versions above will work
— just substitute your own interpreter path below.
# This project's own real interpreter (adjust if you're not on the same machine):
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
# Or set up your own, from the real pinned versions in requirements.txt:
uv venv --python 3.11
uv pip install -r requirements.txtcd scripts/yaqa_port
# Fast, synthetic-only regression suite (no real model needed, ~1-2 min):
"$PYBIN" test_regression.py
# Steps 1-3 individually (all synthetic, no real model needed):
"$PYBIN" 01_gradient_collection_test.py
"$PYBIN" 02_sketch_b_synthetic_test.py
"$PYBIN" 03_rounding_update_test.py
# Step 4: one real layer, end to end (needs the real source model on disk):
"$PYBIN" 04_real_layer_test.py
"$PYBIN" 04_real_layer_test.py --from-cache # skip the model load/backward pass on repeat runs
"$PYBIN" 04_real_layer_test.py --sigma-reg 0.5 # override the damping constant
# Step 5: many real tensors at once (needs the real source model + a real per-tensor plan on disk):
"$PYBIN" 05_full_model_quantize.py --subset default # 3 representative real tensors, comparison only, no save
"$PYBIN" 05_full_model_quantize.py --subset "language_model.model.layers.63.mlp.up_proj" # one tensor by name--incoherence controls whether Hadamard rotation is
carried into the DEPLOYED tensor (not the correction step, which uses it
either way where a valid factor exists):
none (production default, 2026-09-04):
deploys unrotated for every tensor. Reuses
08_gptq_apply_plan.py’s real save/sidecar/MTP mechanism
completely unchanged, real mlx_lm-loadable today, no
runtime overhead. Keeps YAQA’s actual advantage over GPTQ (two-sided
H_I+H_O correction) on every tensor.hadamard (research branch): deploys rotation-aware via
yaqa_core.YaqaQuantizedLinear (real, derived + verified end
to end — test_regression.py test [8]) for the 173/363
tensors that support it. Requires a real loader that does not
exist yet — the saved model is NOT loadable with plain
mlx_lm.load() until that’s built.Real memory constraint, found 2026-09-04: all 363
real plan targets’ Hessians held simultaneously need 557GB (310GB even
excluding lm_head, which is excluded from YAQA’s two-sided
correction separately — its 248,320-wide output dimension alone needs
246.7GB, and even the real reference algorithm doesn’t run its per-layer
correction loop on this tensor). lm_head is NOT left
uncorrected, though: 06_lm_head_gptq.py gives it a real,
scoped one-sided GPTQ correction instead (its input side is small and
safe; only its output side was ever the problem) — see the Files table
below and PORT_LEDGER.md for the full reasoning. The
original “wrap everything, one pass” design for the other 362 tensors
does not fit in 128GB either. Fixed by splitting targets into
memory-bounded batches (--hessian-batch-budget-gb, default
8GB). A SECOND real constraint found immediately after: running multiple
batches inside one long-lived process hits a hard Metal resource-handle
limit ([metal::malloc] Resource limit (499000) exceeded)
that mx.clear_cache() does NOT reset (tested directly,
twice). Fixed by running each batch as its own separate process — use
run_full_yaqa.sh, not
05_full_model_quantize.py directly, for any run with more
than a couple of tensors:
# The real full run (production, mlx_lm-loadable):
./run_full_yaqa.sh /path/to/output --n-calibration 4 --incoherence none
# Watch it live in another terminal (auto-refreshing, color-coded):
./monitor.sh <output_dir>.yaqa_resume/logs/batch_<N>.logrun_full_yaqa.sh queries the real batch count, runs each
batch as a fresh process
(--batch-only N --resume-dir <output>.yaqa_resume),
then does one final --assemble-from-resume pass that loads
every batch’s cached result and performs the real save. STATUS
(2026-09-04): the batching + separate-process fix is being validated now
on the 3-tensor subset — do not treat it as confirmed working until
PORT_LEDGER.md says so.
Manual/advanced flags on 05_full_model_quantize.py
(normally only used by run_full_yaqa.sh itself, not
directly): --resume-dir <dir>,
--batch-only N, --print-num-batches,
--assemble-from-resume.
04_real_layer_test.py and
05_full_model_quantize.py both have real, hardcoded paths
near the top of the file (SOURCE_MODEL,
PLAN_PATH, TARGET_LAYER) pointing at this
project’s own real model and plan — edit those constants to point at
your own real model/plan if running this elsewhere.
| File | What it is |
|---|---|
01_gradient_collection_test.py |
Step 1 |
02_sketch_b_synthetic_test.py |
Step 2 |
03_rounding_update_test.py |
Step 3 |
04_real_layer_test.py |
Step 4 — one real layer, end to end |
05_full_model_quantize.py |
Step 5 — many real tensors at once. Do NOT call directly for a real
multi-batch run – use run_full_yaqa.sh |
run_full_yaqa.sh |
Real orchestrator for a full run – runs each memory-bounded batch as
its own process (works around the real Metal resource-limit crash), runs
06_lm_head_gptq.py automatically, assembles the final
save |
06_lm_head_gptq.py |
Real, scoped one-sided GPTQ correction for
language_model.lm_head – the one tensor YAQA’s two-sided
correction structurally cannot reach (its H_O alone needs
246.7GB). Wraps only this tensor (safe – wrapping every real
Linear/SwitchLinear the way 08_gptq_apply_plan.py’s own
function does needs 125.93GB). Runs automatically via
run_full_yaqa.sh, not meant to be called standalone except
for testing |
monitor.sh |
Live terminal monitor for a running batch – color-coded status/health/progress, real ps + log data, auto-refreshing |
yaqa_core.py |
Shared, validated core math (Hadamard rotation orchestration,
block-Cholesky damping, the LDLQ_2hess correction loop, the
real trace-based error metric) — Steps 4 and 5 both import this, so
there is exactly one implementation of each validated piece |
hadamard_transform.py |
The real Hadamard incoherence-processing transform, ported from
lib/utils/matmul_had.py and independently verified against
the actual running real algorithm (not just its source) |
hadamard_tables.npz |
The 12 real Hadamard factor tables, extracted mechanically
(AST-parsed, not retyped) from the real source and independently
verified (H @ H.T == n·I, all entries ±1) |
test_fixtures/ |
Real reference input/output arrays from actually running the real PyTorch Hadamard transform, used by the regression suite so it never needs a torch install to re-verify against |
test_regression.py |
Full regression suite — every real bug found this session gets a real, independently-cross-checked guard |
interpret_manifest.py |
Live interpretation of a real run’s manifest.json –
explains Metric A (Hessian-weighted error) vs. Metric B (Frobenius) in
plain terms, breaks real progress down by tensor kind and bit-width |
requirements.txt |
Pinned real dependency versions, for a clean install elsewhere |
CHANGELOG.md |
Chronological log of every real change made to this codebase, one entry per real fix/addition |
PORT_LEDGER.md |
The complete, honest engineering ledger – the detailed why
behind each change; CHANGELOG.md is the short chronological
what |
YAQA_COMMAND_REFERENCE.md |
Short, exact command reference (also exported as PDF/HTML) |
RUNNING_GUIDE.md |
Complete, untruncated, step-by-step guide to running everything in this repo, start to finish (also exported as PDF/HTML) |
exports/PORT_LEDGER.html,
exports/PORT_LEDGER.pdf |
Themed HTML/PDF rendering of the ledger |
exports/README.html,
exports/README.pdf |
Themed HTML/PDF rendering of this file |
exports/YAQA_COMMAND_REFERENCE.html,
exports/YAQA_COMMAND_REFERENCE.pdf |
Themed HTML/PDF rendering of the command reference |
../../00_DOCS/HTML/yaqa_pipeline_comparison.html |
Visual side-by-side of the real reference pipeline (trellis-VQ,
runtime rotation) vs. this port’s pipeline (affine
mx.quantize, unrotated deployment) – lives in the project’s
real HTML docs folder, not here |
../../00_DOCS/HTML/yaqa_batching_architecture.html |
Real, spatial diagram of the batching + resume architecture, a verified-claims table, and the proposed name for the method: YAQA-UMA |
graphify-out/ |
Knowledge-graph map of this codebase |
© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.