← Back to index

YAQA_MLX_Vegalaboratories_LTD

2026-09-01

A real, staged port of YAQA (Yet Another Quantization Algorithm — model-preserving adaptive rounding, Cornell-RelaxML) from its original PyTorch/CUDA implementation to Apple’s MLX framework.

Author: Hakim Ghelab, VegaLaboratories LTD

Real, dated, independently-verified progress

The story so far

Scroll or drag sideways — every stop below traces back to a real log, commit, or doc in this repo. Nothing here is aspirational.

01
2026-09-03
Origin

Port begins

No MLX port of YAQA existed anywhere — confirmed by a direct GitHub search. Steps 1–4 built and verified on a real layer of the real 27B model: YAQA’s correction cuts the Hessian-weighted error 75% versus naive rounding.

02
2026-09-04
Discovery

Full-model reality check

Generalizing to the whole model surfaced three real, structural findings: 52% of tensors have no valid Hadamard factor, the real deployment format is trellis-VQ not affine, and holding every tensor’s Hessian at once needs 557GB. Fixed with a hand-derived rotation-aware kernel and memory-bounded batching.

03
2026-09-05 → 06
Root-cause

The 302x bug, hunted down

One tensor’s correction blew up 302x instead of improving. Root-caused twice — first to ill-conditioning, then all the way down to BF16 precision loss inside Hessian construction itself. Fixed and independently reverified.

04
2026-09-07
Ship

First real production build

363/363 tensors corrected (362 YAQA + 1 GPTQ), 6.015 real bits/weight, vision and MTP sidecars reattached. Three post-build regressions found by direct comparison against the source model — each root-caused and fixed.

05
2026-09-08
Fix

MTP sidecar gets its own correction

All 7 multi-token-prediction tensors given real YAQA correction — up to 9,675x better than naive on the worst tensor. Full production build relaunched with the fix applied.

06
2026-09-14
Mechanism

MILP-aware Hessian

A new floor-plus-lexicographic solver mechanism lets the bit-budget optimizer actively protect the tensors the Hessian signal flags as most at risk — not just react to them. Independently code-reviewed; two real bugs found and fixed before shipping.

07
2026-09-15
Proof

First real benchmark numbers

MMLU, GSM8K, IFEval, BFCL, and HumanEval measured end to end, temperature 0, across 6 real built models — the first quantitative read on the whole pipeline, not just internal error metrics.

What this is

YAQA is a whole-model-aware quantization rounding method: instead of correcting each weight using only local, per-layer information (as GPTQ does), it collects a real Hessian sensitivity sketch — both an input-side map (H_I) and an output-side map (H_O) — from the entire model’s real backward pass, so the correction accounts for how each weight actually affects the model’s real output. GPTQ only ever has the equivalent of H_I; YAQA’s LDLQ_2hess correction algorithm is two-sided by construction and needs both. No MLX port of YAQA existed before this work (checked directly via GitHub repo/code search, 2026-09-03 — see PORT_LEDGER.md).

Status

Staged 5-step plan — all 5 steps complete, independently verified end to end on real layers of a real 27B model (Qwen3.5-27B hybrid architecture — standard attention + GatedDeltaNet linear-attention layers), and shipped as a real, loadable production build (363/363 tensors corrected — 362 YAQA + 1 scoped GPTQ for lm_head — 6.015 real bits/weight), including a rebuild with the MTP sidecar’s own real YAQA correction applied. Since then: a new MILP-aware Hessian mechanism that actively protects the tensors the Hessian signal flags as most at risk, and this project’s first real end-to-end benchmark numbers (MMLU / GSM8K / IFEval / BFCL / HumanEval, temperature 0) across 6 real built models. See the timeline above for the full, dated story — every stop traces back to a real log, commit, or doc in this repo.

See PORT_LEDGER.md for the full, honest engineering log — every real bug found and fixed, every real command and its real output, and a plain-English explanation of what any of this means, written for a non-technical reader.

Requirements

Install

This project uses a uv-managed Python tool environment. If you already have mlx/mlx-lm installed some other way, any Python 3.11 environment with the versions above will work — just substitute your own interpreter path below.

# This project's own real interpreter (adjust if you're not on the same machine):
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"

# Or set up your own, from the real pinned versions in requirements.txt:
uv venv --python 3.11
uv pip install -r requirements.txt

Running

cd scripts/yaqa_port

# Fast, synthetic-only regression suite (no real model needed, ~1-2 min):
"$PYBIN" test_regression.py

# Steps 1-3 individually (all synthetic, no real model needed):
"$PYBIN" 01_gradient_collection_test.py
"$PYBIN" 02_sketch_b_synthetic_test.py
"$PYBIN" 03_rounding_update_test.py

# Step 4: one real layer, end to end (needs the real source model on disk):
"$PYBIN" 04_real_layer_test.py
"$PYBIN" 04_real_layer_test.py --from-cache          # skip the model load/backward pass on repeat runs
"$PYBIN" 04_real_layer_test.py --sigma-reg 0.5        # override the damping constant

# Step 5: many real tensors at once (needs the real source model + a real per-tensor plan on disk):
"$PYBIN" 05_full_model_quantize.py --subset default   # 3 representative real tensors, comparison only, no save
"$PYBIN" 05_full_model_quantize.py --subset "language_model.model.layers.63.mlp.up_proj"  # one tensor by name

Step 5c: real output model

--incoherence controls whether Hadamard rotation is carried into the DEPLOYED tensor (not the correction step, which uses it either way where a valid factor exists):

Real memory constraint, found 2026-09-04: all 363 real plan targets’ Hessians held simultaneously need 557GB (310GB even excluding lm_head, which is excluded from YAQA’s two-sided correction separately — its 248,320-wide output dimension alone needs 246.7GB, and even the real reference algorithm doesn’t run its per-layer correction loop on this tensor). lm_head is NOT left uncorrected, though: 06_lm_head_gptq.py gives it a real, scoped one-sided GPTQ correction instead (its input side is small and safe; only its output side was ever the problem) — see the Files table below and PORT_LEDGER.md for the full reasoning. The original “wrap everything, one pass” design for the other 362 tensors does not fit in 128GB either. Fixed by splitting targets into memory-bounded batches (--hessian-batch-budget-gb, default 8GB). A SECOND real constraint found immediately after: running multiple batches inside one long-lived process hits a hard Metal resource-handle limit ([metal::malloc] Resource limit (499000) exceeded) that mx.clear_cache() does NOT reset (tested directly, twice). Fixed by running each batch as its own separate process — use run_full_yaqa.sh, not 05_full_model_quantize.py directly, for any run with more than a couple of tensors:

# The real full run (production, mlx_lm-loadable):
./run_full_yaqa.sh /path/to/output --n-calibration 4 --incoherence none

# Watch it live in another terminal (auto-refreshing, color-coded):
./monitor.sh <output_dir>.yaqa_resume/logs/batch_<N>.log

run_full_yaqa.sh queries the real batch count, runs each batch as a fresh process (--batch-only N --resume-dir <output>.yaqa_resume), then does one final --assemble-from-resume pass that loads every batch’s cached result and performs the real save. STATUS (2026-09-04): the batching + separate-process fix is being validated now on the 3-tensor subset — do not treat it as confirmed working until PORT_LEDGER.md says so.

Manual/advanced flags on 05_full_model_quantize.py (normally only used by run_full_yaqa.sh itself, not directly): --resume-dir <dir>, --batch-only N, --print-num-batches, --assemble-from-resume.

04_real_layer_test.py and 05_full_model_quantize.py both have real, hardcoded paths near the top of the file (SOURCE_MODEL, PLAN_PATH, TARGET_LAYER) pointing at this project’s own real model and plan — edit those constants to point at your own real model/plan if running this elsewhere.

Files

File What it is
01_gradient_collection_test.py Step 1
02_sketch_b_synthetic_test.py Step 2
03_rounding_update_test.py Step 3
04_real_layer_test.py Step 4 — one real layer, end to end
05_full_model_quantize.py Step 5 — many real tensors at once. Do NOT call directly for a real multi-batch run – use run_full_yaqa.sh
run_full_yaqa.sh Real orchestrator for a full run – runs each memory-bounded batch as its own process (works around the real Metal resource-limit crash), runs 06_lm_head_gptq.py automatically, assembles the final save
06_lm_head_gptq.py Real, scoped one-sided GPTQ correction for language_model.lm_head – the one tensor YAQA’s two-sided correction structurally cannot reach (its H_O alone needs 246.7GB). Wraps only this tensor (safe – wrapping every real Linear/SwitchLinear the way 08_gptq_apply_plan.py’s own function does needs 125.93GB). Runs automatically via run_full_yaqa.sh, not meant to be called standalone except for testing
monitor.sh Live terminal monitor for a running batch – color-coded status/health/progress, real ps + log data, auto-refreshing
yaqa_core.py Shared, validated core math (Hadamard rotation orchestration, block-Cholesky damping, the LDLQ_2hess correction loop, the real trace-based error metric) — Steps 4 and 5 both import this, so there is exactly one implementation of each validated piece
hadamard_transform.py The real Hadamard incoherence-processing transform, ported from lib/utils/matmul_had.py and independently verified against the actual running real algorithm (not just its source)
hadamard_tables.npz The 12 real Hadamard factor tables, extracted mechanically (AST-parsed, not retyped) from the real source and independently verified (H @ H.T == n·I, all entries ±1)
test_fixtures/ Real reference input/output arrays from actually running the real PyTorch Hadamard transform, used by the regression suite so it never needs a torch install to re-verify against
test_regression.py Full regression suite — every real bug found this session gets a real, independently-cross-checked guard
interpret_manifest.py Live interpretation of a real run’s manifest.json – explains Metric A (Hessian-weighted error) vs. Metric B (Frobenius) in plain terms, breaks real progress down by tensor kind and bit-width
requirements.txt Pinned real dependency versions, for a clean install elsewhere
CHANGELOG.md Chronological log of every real change made to this codebase, one entry per real fix/addition
PORT_LEDGER.md The complete, honest engineering ledger – the detailed why behind each change; CHANGELOG.md is the short chronological what
YAQA_COMMAND_REFERENCE.md Short, exact command reference (also exported as PDF/HTML)
RUNNING_GUIDE.md Complete, untruncated, step-by-step guide to running everything in this repo, start to finish (also exported as PDF/HTML)
exports/PORT_LEDGER.html, exports/PORT_LEDGER.pdf Themed HTML/PDF rendering of the ledger
exports/README.html, exports/README.pdf Themed HTML/PDF rendering of this file
exports/YAQA_COMMAND_REFERENCE.html, exports/YAQA_COMMAND_REFERENCE.pdf Themed HTML/PDF rendering of the command reference
../../00_DOCS/HTML/yaqa_pipeline_comparison.html Visual side-by-side of the real reference pipeline (trellis-VQ, runtime rotation) vs. this port’s pipeline (affine mx.quantize, unrotated deployment) – lives in the project’s real HTML docs folder, not here
../../00_DOCS/HTML/yaqa_batching_architecture.html Real, spatial diagram of the batching + resume architecture, a verified-claims table, and the proposed name for the method: YAQA-UMA
graphify-out/ Knowledge-graph map of this codebase

© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.