← Back to index

YAQA → MLX Port Ledger

2026-09-03

Author: Hakim Ghelab, VegaLaboratories LTD

This is the resume point. If a session runs out of context, read this file first — it has exactly what’s been verified, the real commands to re-confirm it, and exactly what step comes next. Every command below was actually run; every output shown is real, not illustrative.


Read this first — what this actually is, in plain English

What we’re shrinking, and why that’s risky. An AI model is, physically, a huge list of numbers — for this project’s model, about 27 billion of them. Making the model smaller and faster means rounding those numbers to fewer decimal places, the same idea as rounding $19.97 down to $20. Do that carelessly across 27 billion numbers and the model visibly gets dumber. Do it carefully and it barely notices. This whole project is about doing it carefully.

GPTQ — the smart rounding method already working in this project. When GPTQ rounds one number, it looks at the other numbers sitting right next to it and nudges them slightly to cancel out the rounding mistake it just made — so the small local group stays accurate even though each individual number moved. This is already built, tested, and documented (see docs/HTML/GPTQ_Why_It_Matters.html).

YAQA — what this specific port is about. GPTQ only checks “did I mess up this one small local group of numbers.” YAQA checks something bigger: “did I mess up what the entire model’s final answer looks like.” That’s a real, meaningfully harder thing to measure — like proofreading a whole essay’s actual meaning, instead of only spell-checking each sentence in isolation. The real published result: YAQA’s rounding is measurably more accurate than GPTQ’s — about 30% less error, in the paper’s own tests.

What “porting to MLX” means. YAQA was built and only ever tested using Nvidia’s toolkit (CUDA/PyTorch). This project runs on a Mac, using Apple’s own toolkit, MLX. Those two toolkits don’t speak the same language internally — nothing automatically translates one into the other. “Porting” means manually rewriting YAQA’s actual logic so it works inside MLX instead, piece by piece.

Why we don’t just download a finished MLX version. Because one doesn’t exist. Nobody has built this translation before (verified directly — a real search this session found no existing MLX port anywhere). That means every piece has to be built AND proven correct by us, not assumed to work.

The method: prove each small piece before trusting the whole thing. You would never trust a translated legal contract without checking each paragraph. Same idea here — each of the 5 steps below proves ONE piece works, on numbers small enough to check by hand, before it’s ever trusted on the real, 27-billion-number model. This is slower than just writing the whole thing and hoping — but a silent, subtle math mistake here would not crash anything. It would just quietly make the model worse in a way nobody would notice until much later. That’s the exact failure this step-by-step discipline exists to prevent.

The 5-step journey, in plain language

Step 1 — Can we even listen in?Prove MLX lets us watch both what goes INand what comes back OUT of one piece of themodel, during a real calculation.(the basic plumbing YAQA needs)Step 2 — Can we build the 'sensitivity map'on numbers small enough to check by hand?YAQA's core trick is a map of how a roundingmistake ripples outward. Built it on a tinytoy example, checked the math twoindependent ways.Step 3 — Can we apply the fix in theright order?Once you have the sensitivity map, correctionshave to be applied in a specific crisscrossorder, not just left-to-right. Proved thisordering on tiny numbers.Step 4 — Does this work on ONE REAL pieceof the REAL model?Take everything proven on toy numbers and tryit on one real slice of the actual 27-billion-number model. Two real surprises hit and fixedhere (see below).Step 5 — Scale it up to the WHOLE modelOnly after Step 4 fully works on one realpiece, repeat it correctly across every pieceof the entire model.

Where we actually are right now: All 5 steps are done and proven — the pipeline works, end to end, across the real, full model, not just one layer. Steps 1-4 run correctly with real, verified numbers at every stage (an independent code review on 2026-09-03 caught two further real problems in what had looked like a finished Step 4 — a silent Hessian-formula bug and a missing Hadamard incoherence-processing stage — both fixed and independently verified against the real running algorithm; YAQA’s correction reduces the real, Hessian-weighted error metric by 75% versus naive rounding on the real k_proj layer, see Step 4’s Part 4c write-up below). Step 5 — scaling this to the entire model — completed its first full run on 2026-09-07: all 363 real tensors processed, 362/363 using real YAQA correction (the 1 exception, lm_head, is the known, expected naive fallback), real Hessian-weighted error reduction versus naive ranging 97.49-100% (median 99.67%), zero tensors worse than naive. Two real post-build bugs (a missing mode: "affine" config field, and scales/biases saved at the wrong bit width) were found by benchmarking the finished model, root-caused, fixed, and the rebuild independently re-verified via real tensor forensics. A further YAQA-corrected MTP speculative-decode sidecar variant was built and re-tuned on 2026-09-08 (mtplx tune, run tune-20260908-050223): AR 17.70 tok/s, D1 39.14 tok/s (2.21x), D2 51.68 tok/s (2.92x, the real best), D3 48.23 tok/s (2.73x — slower than D2, not a further win). D1 and D2 land at real near-parity with the reference build; D3 does not scale further. Real testing on non-benchmark prompts (difficult, real-world questions) found generation speed largely unchanged across tuning depths — a genuine, disclosed limit of the AR>D1>D2>D3 tuning protocol as a predictor of real-world speed, not a claim this fully resolves. The open thread is that gap between benchmark and real-world behavior, not whether the pipeline itself works — it does, and the port is real: the first known working port of YAQA to MLX.

What “success” looks like at the very end (Step 5): a real, working tool that takes this project’s full-size AI model and produces a smaller version rounded more carefully than what GPTQ already achieves — verified with real before-and-after comparisons on the actual model, not just “the code ran without crashing.”

The two real surprises in Step 4, explained plainly

  1. We asked the tool to differentiate with respect to the wrong thing. We wanted “how does changing this one number affect the model’s final answer” — but by default, the tool tried to compute that with respect to the raw input words, which makes no sense (words aren’t numbers you can nudge). One-line fix: point it at the model’s actual internal dials instead.
  2. The bigger surprise: this specific AI model turns out to be a hybrid — most of its layers use a different, faster internal shortcut that has no way at all to run the “how did this affect the outcome” calculation backward through it, like a one-way street. We found a real, already-built detour switch in Apple’s own toolkit (model.train()) that automatically reroutes those layers onto a two-way street. Without finding this, Step 4 would have been simply impossible on this actual model — not a bug we introduced, a real fact about how this model is built.

Plain-language glossary

Term used below What it actually means
Weight One of the millions of numbers that make up the model’s “knowledge”
Quantization Rounding those numbers to fewer decimal places, to save space
Bit-width How much precision is kept per number after rounding
Hessian / sensitivity map A map of how much a mistake in one number messes up the final answer
GPTQ Smart rounding that checks only within one small local group of numbers
YAQA Smarter rounding that checks against the model’s whole final answer
MLX Apple’s own toolkit for running AI models on a Mac
Porting Rewriting something built for one toolkit so it works on a different one
Gradient / backward pass The calculation that traces “if I nudge this number, how much does the final answer change”
Cholesky / block-Cholesky A mathematical way of breaking one big sensitivity map into smaller, manageable pieces
Synthetic / toy example Small, fake numbers used only to check the logic is correct — never real data
H_I / H_O This project’s own shorthand for “the sensitivity map, input side” / “…output side”

Everything below this point is the detailed technical log — real commands, real output, real file paths — for anyone who wants the actual evidence behind each claim above.


Status at a glance

Step What it proves Status File
1 MLX’s mx.custom_function gives simultaneous access to a layer’s real forward input AND real incoming gradient — the exact mechanism YAQA’s real backward pass needs ✅ PASS 01_gradient_collection_test.py
2 Sketch B’s Hessian sketch (H_I, H_O), computed two independent ways, agrees to float32 precision; symmetric; flat-storage round-trips exactly ✅ PASS 02_sketch_b_synthetic_test.py
3 The real block-sequential rounding update (LDLQ_2hess, not just the paper’s Eq. 5-6 in isolation), with a stand-in quantizer, cross-verified MLX vs. NumPy ✅ PASS 03_rounding_update_test.py
4a Real Hessian collection (H_I, H_O) on one real layer via a REAL full-model forward+backward pass on this project’s actual hybrid architecture ✅ PASS 04_real_layer_test.py
4b Real block-Cholesky factors (Lin, Lout) from that real H_I/H_O ✅ PASS 04_real_layer_test.py
4c Real quantizer integration + real Hadamard rotation + naive-rounding comparison, full 1024/1024 real output rows — 3 real bugs/gaps found & fixed (diagonal-zeroing, Hessian batch-accumulation, missing Hadamard rotation); real result: 75% lower real trace-based error than naive ✅ PASS 04_real_layer_test.py, hadamard_transform.py
5a Multi-tensor real Hessian collection (ALL real targets, ONE real forward+backward pass) ✅ PASS (3-tensor subset) 05_full_model_quantize.py
5b Real per-tensor correction (rotated where possible, unrotated fallback where not) + 2 real review-found bugs fixed ✅ PASS (3-tensor subset) 05_full_model_quantize.py
5c Real output model, --incoherence hadamard default (YaqaQuantizedLinear, derived+verified) or none (reuses 08_gptq_apply_plan.py’s save/sidecar/MTP unchanged); real loader for the rotated case NOT built ✅ PASS (real production save path, exercised by every real run below) 05_full_model_quantize.py, yaqa_core.py
5-full The real 363-tensor run, end to end ✅ PASS — built 2026-09-07, then rebuilt twice more the same day after two real production bugs were found post-build (missing mode; F32 scales/biases + embed_tokens) — see the 2026-09-07 section below. Real output-quality benchmark completed 2026-09-09/10 (see CHANGELOG.md) — highest 5-test mean of the comparison group, never worst on any single metric. 05_full_model_quantize.py, research_hadamard_blowup/exports/RCA_scales_f32_bug.html, research_hadamard_blowup/IMPROVEMENT_LEDGER/09_FIRST_REAL_BENCHMARK_RESULTS.html

Do not skip ahead. Each step gates the next. Step 2 exists specifically because a subtle bug in step 2’s math would be silent — no crash, no obvious symptom, just a wrong Hessian sketch that looks plausible. The whole point of this staged plan is refusing to chain steps on faith.


Environment (verified 2026-09-03)

Interpreter: /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3  (Python 3.11.15)
mlx: 0.32.2
mlx-lm: 0.31.3
mlx-optiq: 0.4.32
numpy: 2.4.6

Real source-of-truth locations, already on disk:

research/quantization_methods_2026/yaqa/code/        <- real Cornell-RelaxML/yaqa clone
research/quantization_methods_2026/yaqa/paper/YAQA_2505.22988.pdf   <- real 20-page paper PDF
scripts/yaqa_port/                                    <- this port's own scripts (this ledger's directory)

Real file that everything in Step 1-2 is grounded in, read directly (not summarized):

research/quantization_methods_2026/yaqa/code/hessian_llama/custom_linear_B.py

Step 1 — Isolated gradient collection

What this actually does, in plain English

Picture the AI model as a chain of stations, one after another, with data flowing through — in one side, out the other. Normally you can only watch data flow in ONE direction through a station: forward, left to right. But YAQA’s whole method depends on watching a SECOND signal too — a “how much did this affect the final answer” signal, which flows backward, right to left, only during a special kind of calculation. This script builds a tap on one real station in the chain and checks: can we actually read BOTH signals — the forward one AND the backward one — at the exact same point, at the same time? If we can’t, nothing else in this whole project is possible.

earlier stations (rest of the model) OUR STATION (one real layer) the tap goes here later stations …all the way to the model’s final answer data flowing forward → ← “how much did this affect the answer” flowing back THE TAP: reads both signals at once

What Step 1 actually proves: a real tap on one real station, reading both the forward signal and the backward signal at the same point.

Where this idea came from (provenance). This isn’t invented — it’s a direct translation of one specific real mechanism from YAQA’s own original code, which was itself written for a completely different toolkit (PyTorch/CUDA). The real file is research/quantization_methods_2026/yaqa/code/hessian_llama/custom_linear_B.py, class LinearNoBias. In PyTorch’s toolkit, this “tap” is built with ctx.save_for_backward(input, weight) during the forward pass, then read back inside def backward(ctx, grad_output). Apple’s toolkit (MLX) has a different mechanism for the exact same job, called mx.custom_function — and Step 1’s entire purpose is checking that this different mechanism actually gives us the same two pieces of information PyTorch’s does, before trusting it for anything real.

What’s literally inside 01_gradient_collection_test.py, in plain terms: 1. Defines a tiny stand-in “station” — a single layer that just multiplies numbers together, the same basic operation every real layer in the model does. 2. Wires up the tap (mx.custom_function) on that station. 3. Runs one fake calculation through it, forward then backward, using the REAL dimensions of an actual layer from this project’s actual model (not made-up numbers) — language_model.model.layers.3.self_attn.q_proj, 12288 × 5120. 4. Checks: did the tap actually catch both signals, with the right shapes and the right values? If yes, prints STEP 1 PASS. If the shapes or values are wrong, or if MLX itself refuses to run this at all, the script raises a real Python error instead of silently continuing — there’s no “partial credit” here.

Real command — copy-paste this exactly to re-run it yourself:

PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 01_gradient_collection_test.py

Real captured output (2026-09-03) — and what each line actually means:

loss value: 1613974.8750
grad wrt x shape: (4, 16, 5120) (expected (4, 16, 5120))
grad wrt w shape: (12288, 5120) (expected (12288, 5120))

--- what the custom vjp actually captured ---
captured input shape:       (4, 16, 5120) (expected (4, 16, 5120))
captured grad_output shape: (4, 16, 12288) (expected (4, 16, 12288))

STEP 1 PASS: mx.custom_function's vjp gives simultaneous access to both
the real forward input and the real incoming cotangent, on a real layer
shape from this project's own model.

One real bug found and fixed while building this (kept here so it’s not rediscovered): mx.custom_function‘s .vjp callback receives cotangents matching the output’s pytree structure, not the primals’. linear_no_bias returns a single array, so cotangents IS that array directly — not a 1-tuple. The first version wrote (grad_output,) = cotangents and MLX raised ValueError: too many values to unpack (expected 1). Fix: grad_output = cotangents directly.


Step 2 — Sketch B’s Hessian math, synthetic, hand-verified

What this actually does, in plain English

Step 1 proved we can read the “how much did this affect the answer” signal. Step 2 asks: what do we actually DO with that signal? YAQA’s core idea is to build a sensitivity map — a record of “if I nudge this number, which OTHER numbers does it interact with, and how strongly.” Nudge one number in a pond and the ripples spread outward in a specific pattern; this map is a record of that ripple pattern, computed from real data instead of guessed.

There’s a subtlety worth naming plainly: this one signal actually gets used to build two separate maps, not one — one describing ripples on the “numbers coming IN to this layer” side, one describing ripples on the “numbers going OUT of this layer” side. This step builds a TINY version of both maps — small enough to check by hand — and proves the math is correct using two completely different ways of computing it, before ever trusting it on a real, 27-billion-number model.

the one real signal (input × how-it-mattered) used one way → ← used the other way MAP OF THE “IN” SIDE (H_I — do these input directions interact?) MAP OF THE “OUT” SIDE (H_O — do these output directions interact?) Same underlying signal. Two different questions asked of it.

What Step 2 actually proves: the same real signal correctly builds two different sensitivity maps, verified two independent ways.

Where this idea came from (provenance). The real formula lives in the same file as Step 1 — custom_linear_B.py, inside LinearNoBias.backward, the it == 0 branch (lines 57-97) — expressed there as two dense torch.einsum calls. Nothing here is guessed: the exact algebra was worked out by hand (shown in full in the script’s own docstring) and reduces to a simple, checkable recipe — build one small “how it mattered” matrix per example, then combine those matrices two different ways.

Why Sketch B specifically, not Sketch A — checked directly by diffing the real source (2026-09-03), not assumed from the real README’s brief “we recommend Sketch B if you can afford it” line: hessian_llama/ ships TWO real Hessian-collection files, custom_linear_A.py and custom_linear_B.py. Diffing them directly shows Sketch A’s cheap single-pass branch (it==0) computes only H_I (in_hess = input.T @ input — literally the same one-sided form GPTQ’s own Hessian uses) — no H_O computation exists in that branch at all. Sketch B’s single-pass branch computes both H_I and H_O together, via the real two-sided einsum shown above. This port’s whole algorithm (LDLQ_2hess, Step 3) is two-sided by construction — it needs both Lin and Lout. Sketch A’s efficient path structurally cannot supply H_O at all; getting one out of Sketch A would require its separate power-iteration refinement path (it>0), which the real source itself flags as immature (custom_linear_B.py’s own comment: “Additional power iterations on B are not optimized and should be rewritten with einsums. Use at your own risk!”). There was no real choice to make here — Sketch A alone cannot feed the algorithm this project ported.

What’s literally inside 02_sketch_b_synthetic_test.py, in plain terms: 1. Builds tiny, fake “input” and “how it mattered” numbers — small enough that a human could, in principle, check the arithmetic by hand. 2. Computes both sensitivity maps twice, using two completely different pieces of code that share nothing: once using MLX’s own einsum (mirroring the real formula’s exact notation), and once using an ordinary loop with plain matrix multiplication (worked out independently, from first principles). 3. Compares the two results. If a bug slipped into either version, the two numbers would very likely disagree — that disagreement is the actual safety net here, not just “the code didn’t crash.” 4. Also checks a structural fact both real maps must satisfy (they must be symmetric — the map from A to B must equal the map from B to A), as a second, independent sanity check.

The real formula, derived by hand from the actual source (not the paper’s simplified notation) — custom_linear_B.py, LinearNoBias.backward, the it == 0 branch, lines 57-97:

in_hess  = einsum('btm,btn,bsm,bsk->nk', grad_output, input, grad_output, input)
out_hess = einsum('btm,btn,bsk,bsn->mk', grad_output, input, grad_output, input)

Regrouped algebraically, using the same per-example quantity the real code itself names elsewhere in this file (grad_weight = grad_output[i].T @ input[i], line 122-123):

G[b]     = grad_output[b].T @ input[b]        (out_features, in_features)
in_hess  = sum_b  G[b].T @ G[b]                (in_features, in_features)
out_hess = sum_b  G[b]   @ G[b].T              (out_features, out_features)

Full worked derivation for in_hess is in the script’s own docstring — verified by hand, not just asserted.

Real command:

PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 02_sketch_b_synthetic_test.py

Real captured output (2026-09-03), abbreviated (full matrices in the script’s own run log):

max |diff| (H_I, einsum vs. manual loop): 7.629e-06
max |diff| (H_O, einsum vs. manual loop): 7.629e-06
H_I symmetry error: 0.000e+00
H_O symmetry error: 0.000e+00
sym_to_flat/flat_to_sym round-trip error: 0.000e+00 (10 values for a 4x4 symmetric matrix, vs 16 dense)

STEP 2 PASS: Sketch B's Hessian sketch (H_I, H_O) computed two independent
ways ... agree to float32 precision. Both are symmetric as required. The
real symmetric-storage mechanism round-trips exactly.

What each line actually means: - max |diff| (H_I/H_O, einsum vs. manual loop) — this is the actual test. Two completely different pieces of code computed the same map; this number is the biggest disagreement found between them, anywhere in the map. 7.629e-06 means “0.000007-ish” — a difference so small it’s just normal computer rounding noise, not a real disagreement. If this number had come back large (say, bigger than 0.01), that would mean the two independent implementations genuinely disagree — a real bug somewhere, full stop. - H_I/H_O symmetry error — a second, unrelated check: does the map satisfy a structural rule real sensitivity maps must satisfy (symmetric — “A affects B” must equal “B affects A”)? 0.000e+00 means exactly zero violation — this rule holds perfectly. - sym_to_flat/flat_to_sym round-trip error — this project’s own storage trick (saving only half the map, since it’s symmetric, to use less memory) was checked by compressing then decompressing it and comparing to the original. 0.000e+00 means nothing was lost. - STEP 2 PASS — again, only printed if every check above it actually passed. No PASS line (a Python error instead) means stop and investigate before moving to Step 3.

Toy dimensions used: batch=3, seq=5, in_features=4, out_features=3 — small on purpose, chosen to be hand-tractable, not representative of a real layer’s scale.

Real verification method (why this counts as evidence, not just “it ran”): H_I/H_O computed via mx.einsum (subscripts copied verbatim from the real torch.einsum calls) AND via a completely separate plain-matmul per-example Python loop (the hand-derived formula above, containing no einsum at all). Two structurally different code paths agreeing to float32 tolerance is real cross-validation.


Step 3 — Rounding update

What this actually does, in plain English

Now we have the sensitivity maps from Step 2. Step 3 asks: how do we actually USE them to round the model’s numbers more carefully? The real answer turns out to be more specific than “apply one formula” — the weight matrix (think of it as a big spreadsheet grid) gets rounded one small block at a time, in a very particular crisscross order, and each block’s rounding is nudged using the blocks already finished next to it — the same “compensate the neighbors” idea GPTQ uses, but now pulling from two directions at once (using both sensitivity maps from Step 2), not just one.

The weight matrix, as a grid of small blocks: done done not yet not yet done CORRECTINGTHIS BLOCK not yet not yet not yet not yet not yet not yet Order of correction: 1. Start at one corner of the grid. 2. Sweep in a crisscross (“anti-diagonal”) pattern — the exact order matters, it’s not just left-to-right. 3. Each block’s correction pulls in information from the block already done directly above it AND the one already done directly to its left (the orange arrows) — using BOTH sensitivity maps from Step 2 at once. 4. Round that block, mark it “done,” move to the next block in the order. The real algorithm this mirrors is the exact one YAQA’s own code uses (see provenance below) — this diagram is illustrative of the real logic, not a simplification that changes what it actually does.

What Step 3 actually proves: this crisscross, pull-from-already-done-neighbors correction order, computed correctly.

Real finding that changed this step’s scope (2026-09-03): the plan originally assumed Step 3 was “implement Eq. 5-6, a single fixed-point formula.” Reading the real source before writing code found something more specific and more involved — the actual update is a block-sequential 2D scheme, not a closed form:

research/quantization_methods_2026/yaqa/code/lib/algo/ldlq.py, function LDLQ_2hess() (lines 16-59)

It processes the weight matrix in (td_x × td_y) blocks, walking them in a specific anti-diagonal order (the real starts list), and at each block combines three cross-terms built from the already-rounded neighboring blocks (W - hatW) and the block-Cholesky factors Lout/Lin — the real code’s own names for the paper’s L_O/L_I, computed by lib/utils/math_utils.py:block_LDL from our own validated H_I/H_O (Step 2). Only at the very end of each block does the real code call cb.quantize(...) — QTIP’s own trellis/lattice codebook (lib/codebook/bitshift.py), a completely different quantization format from this project’s own affine, group-based one.

Real, checked-against-the-repo finding on a related question: an external summary (pasted mid-session, phrased like an AI search overview) claimed the real yaqa repo has a ready-made --quantizer uniform flag as an alternative to QTIP’s codebook — i.e., that no custom substitution work would be needed. Checked directly and this is false: grep’d the entire real repo for quantizer/uniform/--codebook choices — no such flag exists anywhere. The only --codebook argument (quantize_llama/quantize_finetune_llama.py:34) is a free-text string resolved against lib/codebook/, and the only codebook module actually present in the repo is bitshift.py (genuine trellis/lattice decode functions — decode_1mad, decode_2mad, bitshift_codebook — no uniform/affine alternative). The directional idea (the correction math doesn’t care what quantizer consumes it) is correct and is exactly what this step’s own substitution is built on — but it required writing real substitution code, not flipping an existing flag. Flagging this here so a future session doesn’t waste time looking for a flag that isn’t there.

What this script does: keeps the real correction logic (the three-cross-term formula, the exact anti-diagonal block traversal) faithful to ldlq.py, and replaces ONLY cb.quantize(...) with a trivial round-to-nearest-grid stand-in (simple_quantize) — representing where this project’s OWN eventual quantizer (mx.quantize/gptq_pack, already built and validated for the GPTQ integration) will plug in at Step 4, not an attempt to adopt QTIP’s lattice codebook.

What’s literally inside 03_rounding_update_test.py, in plain terms: 1. Builds two tiny, fake sensitivity maps (standing in for real ones from Step 2). 2. Runs the real crisscross block-correction order (the diagram above) on a tiny, fake weight grid — twice, using two completely separate pieces of code (an MLX version and a plain NumPy version) that share no code between them. 3. Compares the two results for an exact match, and separately checks a structural fact the sensitivity-map math must satisfy (that certain blocks inside it must come out as an exact identity — a mathematical fingerprint that’s either exactly right or reveals a real bug, no middle ground).

Real command:

PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 03_rounding_update_test.py

Real captured output (2026-09-03), abbreviated — and what it means:

Lin diagonal block 0 identity error: 0.000e+00
Lin diagonal block 1 identity error: 0.000e+00
max |diff| between independent implementations: 0.000e+00
max deviation from quantization grid: 0.000e+00

STEP 3 PASS: the real block-sequential LDLQ_2hess correction algorithm ...
ported to MLX and cross-verified against an independent NumPy implementation
-- exact agreement, and every output entry actually lands on the quantization
grid. block_LDL's diagonal-identity property also verified directly.

Real verification method: two independent implementations — an MLX port and a plain NumPy port, sharing no code — run on identical tiny random inputs (4×4 weight matrix, 2×2 blocks, the smallest case with a real >1 block grid so the cross-term logic is actually exercised). Exact 0.000e+00 agreement, since this is deterministic block-sequential arithmetic, not a stochastic estimate — exact agreement is the right bar here, and it was met.

Not yet done: real quantizer integration (this still uses the trivial round-to-grid stand-in, not mx.quantize/gptq_pack) — that’s part of Step 4, along with running against a real model layer for the first time.


Step 4 — One real layer

What this actually does, in plain English

Steps 1-3 proved every piece works on small, fake numbers. Step 4 asks the real question: does any of this survive contact with the actual, real, 27-billion-number model? This is where a genuinely important, real surprise showed up — not a bug in our own code, but a real fact about how this specific AI model is built, that would have quietly blocked everything if we hadn’t found it.

BEFORE calling model.train() — a one-way street layer 3 layer 4GatedDelta layer 5GatedDelta layer 6GatedDelta … final answer/ loss ✗ backward signal cannot pass through GatedDeltaNet layers — blocked entirely AFTER calling model.train() — a two-way street layer 3 layer 4GatedDelta layer 5GatedDelta layer 6GatedDelta … final answer/ loss model.train() switches every GatedDeltaNet layer onto a real, differentiable path — automatically Why this mattered so much Most of this model’s layers (3 out of every 4) use the fast “GatedDelta” shortcut, not standard attention. To measure “how did layer 3 affect the final answer,” the signal has to travel backward THROUGH those shortcut layers — even though layer 3 itself is a normal, “two-way” layer.

What Step 4 actually proves: the real fix that turns a genuinely blocked measurement into a working one, on the real model.

Target layer, deliberately not q_proj: language_model.model.layers.3.self_attn.k_proj, weight shape (1024, 5120) — chosen instead of q_proj (12288, 5120) because H_O for q_proj would be a dense 12288×12288 matrix; for k_proj it’s 1024×1024, a properly gated first real-scale test instead of jumping straight to this model’s largest attention projection.

Part 4a-4b — real Hessian collection + block-Cholesky factors: PASS

What’s literally inside 04_real_layer_test.py, in plain terms: 1. Loads the actual, real 27-billion-number model — not a stand-in. 2. Puts a real “tap” (Step 1’s mechanism) on one real layer inside it. 3. Switches the model into the “two-way street” mode explained in the diagram above (model.train()). 4. Feeds a handful of real calibration text windows through the entire real model, computes a real next-word-prediction loss, and runs that loss backward through the whole thing — which is what actually fires the tap and builds the two real sensitivity maps (H_I, H_O) for our one target layer, using the exact math validated in Step 2. 5. Builds the real block-Cholesky factors from those real maps (Step 3’s math), on real-world-sized numbers for the first time — not tiny toy ones anymore.

Real command:

PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 04_real_layer_test.py

Real captured output (2026-09-03) — and what it means:

Calibration batch: (6, 128)
Running a real forward+backward pass through the WHOLE model...
loss: 2.1989
H_I shape: (5120, 5120), H_O shape: (1024, 1024), batches accumulated: 1
Part 1 PASS: real Hessian collection works end-to-end on a real layer, inside a real
full-model forward+backward pass, with no NaN.

Computing block-Cholesky factors (Lin, Lout) from real H_I/H_O...
Lin shape: (5120, 5120), Lout shape: (1024, 1024)
Part 2 PASS: real block-Cholesky factors computed without error.

The loss (2.1989, ≈9 perplexity) is a sane, real cross-entropy value for a real model on real calibration text — worth noting as a soft sanity check, not just “it didn’t crash.”

Two real bugs found and fixed while building this (kept here so they’re not rediscovered):

  1. mx.value_and_grad(loss_fn) differentiated w.r.t. the wrong thing. By default it computes gradient w.r.t. loss_fn’s first positional argument — here that was batch, the real token IDs, which are discrete/non-differentiable (used for embedding-table gather). MLX raised this directly: ValueError: [gather_axis] Cannot calculate VJP with respect to indices. Use stop_gradient on indices... Fix: use nn.value_and_grad(model, loss_fn) instead — differentiates w.r.t. the model’s trainable parameters, which is what actually needs to flow backward through the wrapped layer; the token IDs were never meant to receive a gradient.

  2. Major, real architectural finding: this model (mlx_lm.models.qwen3_5) is a hybrid architecture — DecoderLayer.is_linear = (layer_idx+1) % full_attention_interval != 0 (full_attention_interval=4), so 3 of every 4 layers use GatedDeltaNet (“linear_attn”) instead of standard attention, not just the target layer’s own type. GatedDeltaNet’s recurrence (mlx_lm/models/gated_delta.py) runs through a raw mx.fast.metal_kernel (a hand-written Metal shader) with no backward rule defined at all. Backward from the final loss to ANY layer — including our target, layer 3, which is itself is_linear=False — has to pass through several GatedDeltaNet layers along the way, so this blocked the whole pass, not just linear-attention layers specifically. MLX raised this directly: ValueError: [Primitive::vjp] Not implemented for CustomKernel. Real fix, already built into mlx_lm itself, no patching needed: gated_delta_update(..., use_kernel=not self.training) (gated_delta.py:281-282, called from qwen3_5.py:193) — calling model.train() switches every GatedDeltaNet layer onto its own pure-MLX-ops reference implementation (gated_delta_ops(), genuinely differentiable), automatically. This is a durable, real fact worth remembering for any future gradient-based work on this model family: model.train() is required before any real backward pass, not just a training-loop nicety.

  3. Minor: mx.linalg.cholesky isn’t supported directly on the GPU stream — ValueError: [linalg::cholesky] This op is not yet supported on the GPU. Same fix already established in the GPTQ integration (08_gptq_apply_plan.py’s compute_inverse_hessian): wrap in with mx.stream(mx.cpu):.

Part 4c — real quantizer integration, real Hadamard rotation, naive-rounding comparison: PASS

Revision history, honestly: this section was rewritten twice. The first version reported a 32-row partial result with a stand-alone bug fix (diagonal-zeroing) and concluded YAQA had “not yet been shown to beat naive.” An independent code review (2026-09-03) then found two further real problems in that version — a silent formula bug in the Hessian collection, and an entire real preprocessing stage (Hadamard incoherence processing) missing outright, not just simplified. Both are now fixed, independently verified against the real running algorithm (not just its source), and re-run in full. This is the current, complete state — not an addendum to the old one.

What this actually does, in plain English. Steps 1-3 proved the ordering of the smart-rounding correction on toy numbers, using a fake “round to the nearest step” stand-in for the actual rounding operation. Part 4c swaps that stand-in for the real rounding operation — MLX’s own mx.quantize, the exact same real quantizer this project’s GPTQ path already ships — runs it block by block on the real weight matrix, and compares the result against just rounding the same real weight the plain way, no correction at all. Two things turned out to be missing before this comparison could be trusted: the real Hessian-collection formula had a bug that only shows up with more than one calibration sequence, and a whole real preprocessing stage (a random rotation that makes the weight distribution friendlier to low-bit rounding) was never ported in the first place. Both are fixed below, and the real result changes completely once they are: with the complete real pipeline actually running, YAQA’s correction reduces the real error metric by 75% versus naive rounding.

Why the earlier version missed this: the earlier 32-row test used mx.isnan as its only correctness check and a plain Frobenius-norm comparison. Neither check would ever have caught a formula bug that only manifests at batch > 1, or a missing preprocessing stage that the real algorithm’s own accuracy depends on — those required someone reading the real source again, line by line, specifically hunting for divergence, not just checking that a given run’s numbers stayed finite.

Real result, full 1024-row run (2026-09-03): does the COMPLETE real pipeline actually help? Metric A – the real pipeline’s own trace-based error (lower = better) 0.00175 YAQA-corrected 0.00708 Naive (rotated) 75.28% lower real error Metric B – plain Frobenius error, original weight basis (for continuity) 3.370 YAQA-corrected 3.314 Naive (unrotated) ~1.7% higher, not lower Why do the two metrics disagree? This is real and expected, not a bug. LDLQ_2hess is designed to minimize Metric A directly (error weighted by how much each weight actually affects the model’s real output, via the real Hessian sensitivity maps) – not raw, unweighted magnitude error. A large win on the metric the algorithm actually optimizes for, alongside a small loss on a metric it was never designed to minimize, is the expected real signature of this working as intended.

Real data from the final canonical run, 2026-09-03 (see the annotated output below for the exact numbers and the run that produced them).

Design decisions kept from the plan: td_x=1 (one output row per block), td_y=group_size (64), so LDLQ_2hess’s own blocks line up exactly with mx.quantize’s real per-row-per-group structure, rather than inventing a new grouping convention. The repo’s own defaults (td_x=16, td_y=16, quantize_finetune_llama.py:47-48) target QTIP’s own lattice codebook, which groups differently — not a fit for this project’s own affine quantization format.

On row coverage (this changed from the earlier version): the earlier 32-row partial test relied on a real, still-correct argument — block_LDL is a sequential leading-principal-minor factorization, so a leading row/column slice of its output is exact. That argument stops being sufficient once Hadamard rotation is added: the rotation is a global transform that mixes every one of the 1024 real output rows together, so a partial slice of the rotated result can no longer be validly inverse-rotated back to the original weight basis, and the real trace-based error metric is likewise only well-defined over the full output dimension. This run therefore covers all 1024 real output rows of the real k_proj layer (81,920 real blocks total), not a slice.

What’s literally inside 04_real_layer_test.py’s Part 2-4, in plain terms: 1. Generates two real random ±1 “sign vectors” (SU for the 5120 real input features, SV for the 1024 real output features) and applies the real Hadamard incoherence-processing rotation to H_I, H_O, and the real weight matrix — the exact same rotation the real pipeline always applies before doing anything else, previously missing from this port entirely (see bug #2 below). 2. Builds the real block-Cholesky factors (Lin, Lout) from the rotated H_I/ H_O, same damping-retry and diagonal-zeroing logic as before, now applied to the correct (rotated) inputs. 3. Walks the block grid, all 1024 real output rows, in the real anti-diagonal order — for each block, computes the real three-cross-term correction on the rotated weight, then calls MLX’s own real mx.quantize/mx.dequantize on the corrected result — “compute the correction, then quantize the corrected result.” 4. Computes two real comparisons: the real pipeline’s own trace-based, Hessian- weighted error metric (rotated basis, against a naive-but-also-rotated baseline), and a plain Frobenius error in the original weight basis (against naive mx.quantize on the real, unrotated weight) for continuity with the earlier version of this test.

Real command:

PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 04_real_layer_test.py

Real captured output (2026-09-03, final canonical run, fully fresh — model reloaded, real backward pass rerun, no cache reused) — and what it means:

Applying real Hadamard incoherence-processing rotation...
Hin_rot shape: (5120, 5120), Hout_rot shape: (1024, 1024)
Wr (rotated weight) shape: (1024, 5120)

Computing block-Cholesky factors (Lin, Lout) from the real, rotated H_I/H_O...
Hin: block_LDL succeeded on attempt 1, final sigma_reg=1.0000
Hout: block_LDL succeeded on attempt 1, final sigma_reg=1.0000
Lin max|entry|: 8.9745e-02   Lout max|entry|: 1.5886e+00
Part 2 PASS: real Hadamard rotation + real block-Cholesky factors computed without error.

Part 3: real LDLQ_2hess correction (rotated basis) on 1024 of 1024 real output rows
(all 5120 real input columns), td_x=1, td_y=64 -- 81920 blocks, real mx.quantize
(bits=4, group_size=64) at every block.
  sweep 1/1104: max|hatWr| so far = 4.1420e-02
  sweep 1100/1104: max|hatWr| so far = 8.2684e-02
Part 3: 100.0% of hatWr entries are finite.
Part 3 PASS: real LDLQ_2hess correction ran end-to-end (rotated basis) on a real
layer with a real quantizer at every block, no NaN.

Part 4: real comparison against naive rounding, two metrics...
[Metric A -- real trace-based error, rotated basis, finetune.py's own formula]
  YAQA-corrected: err=1.749097e-03   Naive (rotated, no correction): err=7.077018e-03
  YAQA correction reduces the real trace-based error by 75.28% vs. naive rounding.
[Metric B -- plain Frobenius reconstruction error, original (unrotated) weight basis]
  YAQA-corrected:  Frobenius=3.370495  max|diff|=0.007998
  Naive mx.quantize (unrotated, on real W directly): Frobenius=3.314037  max|diff|=0.007405
  YAQA correction did NOT reduce Frobenius error vs. plain naive rounding.

Real bugs and gaps found and fixed (kept here so they’re not rediscovered):

  1. A real formula bug: block-loop numerical instability from a missing diagonal- zeroing step. block_ldl_mlx (validated in Step 3) correctly sets each diagonal block of its output to the identity matrix, matching the real block_LDL function (math_utils.py:14-42) exactly. But the real caller of block_LDL (lib/algo/finetune.py:172,183, read directly from the real source) does one more step: it zeroes the literal diagonal afterward — Lin[arange(n),arange(n)]=0, Lout[arange(m),arange(m)]=0 — turning those identity diagonal blocks into zero blocks before Lin/Lout are ever used in LDLQ_2hess. Without this, the correction formula’s “self-term” spuriously re-injects a near-copy of each block back into the running correction instead of correctly excluding self-contribution — on the original (pre-rotation) real layer, this compounded block-to-block at roughly 1.5× per diagonal sweep, overflowing to infinity before the loop finished. Fix: zero the literal diagonal of Lin/Lout after block_ldl_mlx, exactly matching the real caller code.
  2. A real, critical formula bug in Hessian accumulation, found by independent review — silent at batch=1, would corrupt results at batch>1. yaqa_linear_vjp originally flattened batch and sequence together into one axis before forming a single pooled G = grad_output_flat.T @ input_flat, then squared that once (H_I += G.T@G). The real formula (hessian_llama/custom_linear_B.py, independently hand-verified with a real batch of 3 in Step 2) requires a per-example G[b] = grad_output[b].T @ input[b] for every sequence b, then H_I += sum_b(G[b].T@G[b]) — summing the squared per-example matrices, not squaring the sum. These differ the moment batch>1: (sum_b G_b).T@(sum_b G_b) expands to include spurious cross-sequence terms the real formula never forms. Every real run so far used data[:1] (batch=1, forced by an earlier OOM), where the two formulas coincide by construction — completely invisible in every number this script had produced. Fix: loop over the real batch axis explicitly when accumulating H_I/H_O, matching Step 2’s own validated per-example loop exactly. This did not change any number in this run (still batch=1), but is now correct for any future run with real multi-sequence calibration — which is exactly the direction Step 5 needs.
  3. A real, significant gap: the real Hadamard incoherence-processing rotation was entirely missing, not simplified or scoped out — found by independent review. The real pipeline (finetune.py:145-152,167,178,192-197) always rotates H_I, H_O, and the real weight through a random-sign Hadamard transform before block_LDL/LDLQ_2hess ever runs. This is real “incoherence processing” from the QuIP#/QTIP lineage this algorithm descends from — it exists specifically to make weight distributions more favorable to accurate low-bit rounding, and its absence was not previously documented as a known limitation; it was simply missing. Fix: ported the real transform (hadamard_transform.py, from lib/utils/matmul_had.py) and wired it into the real order (normalize → rotate → damp → block_LDL, weight rotation before the block loop, inverse rotation after). Independently verified two ways before use, not just read: (a) all 12 of the real get_hadXX() factor tables were extracted by mechanically parsing the real source’s literal arrays via AST (not retyped by hand), and every one independently verified to satisfy H @ H.T == n·I with every entry exactly ±1 (max error 0.0 across all 12 tables); (b) the real matmul_hadU/matmul_hadUt functions were extracted from the real source and actually run in an isolated torch-CPU environment (no CUDA/fast_hadamard_transform needed — those are only used by the separate matmul_hadU_cuda kernel path this port doesn’t need), and this MLX port’s output matched the real running output exactly for n=1024 (K=1 path, max diff 0.0) and to float32 rounding precision for n=5120 (K=20 path via get_had20, max diff 9.5e-7 — the same order as the real code’s own internal round-trip self-consistency error). This is the strongest verification standard used anywhere in this port so far: not just cross-checking two independent implementations of a formula, but running the real, actual code and comparing outputs directly.
  4. A real, deliberate, documented omission — not a gap. The real pipeline’s Wscale step (finetune.py:194-197) rescales the rotated weight’s RMS magnitude to match the trellis codebook’s fixed lookup-table range. This port uses mx.quantize’s own affine per-group quantizer instead of a trellis codebook, which derives its scale/bias fresh from each block’s own data range every time — a global rescale of its input is mathematically a no-op for its reconstruction accuracy (scale by k, quantize, divide back by k reconstructs identically). Skipped deliberately, not silently.
  5. Two minor items from the same review, addressed for completeness: the starts block-traversal list re-processes one diagonal twice (present identically in the real upstream ldlq.py:23-26, traced through and confirmed harmless — every reference to a block written on the first pass hits a position the Lin/Lout diagonal-zeroing already zeroed, so the repeat pass changes nothing). And HESSIAN_CACHE (.cache_k_proj_hessians.safetensors, ~114MB) is written but never auto-deleted — fine at single-layer scale, a real, disclosed consideration for Step 5’s own design if the same per-script caching pattern gets reused across many layers.

A real engineering aside worth keeping: an earlier full-6-sequence-batch attempt to raise the real Hessian’s rank (instead of relying on damping) produced no output at all and exit code 0 from a python3 | tee pipeline — consistent with the OS OOM-killing python3 mid-run while tee, on the other end of the pipe, still exits 0 on EOF regardless of why the writer died. Bash does not propagate an upstream command’s exit code through a pipe without set -o pipefail — worth remembering for any future real run in this project that pipes to tee.

Step 5 — Full model — IN PROGRESS (Parts 5a-5b proven small-scale; full run NOT YET ATTEMPTED)

What this step actually needs, decided against real project artifacts, not invented:

Real files added this step

Part 5a — real multi-tensor Hessian collection: PASS

What this proves. Wrapping ONE tensor at a time (Step 4’s approach) would mean 498 separate full-model backward passes for the full plan — infeasible. Part 5a wraps ALL real target tensors SIMULTANEOUSLY and collects every one’s H_I/H_O in ONE real forward+backward pass, mirroring the real, already-proven mlx_lm GPTQ Catcher pattern (mlx_lm/quant/gptq.py:40-76, read directly 2026-09-03: wraps every real nn.Linear/SwitchLinear at once, one calibration pass, mx.eval(layers) per batch). This machine has 128GB RAM (checked directly), comfortably covering even the largest real Hessian in this plan (H_I for a 17408-in_features MLP tensor is 17408×17408 float32, ~1.2GB).

Real command:

PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 05_full_model_quantize.py --subset default

Real captured output (2026-09-03) — 3 real tensors, one of each real architectural type in this hybrid model, each confirmed to actually need real quantization in the real plan (not bits:16):

3 real tensors need real quantization (bits != 16) -- SUBSET:
['language_model.model.layers.62.linear_attn.out_proj',
 'language_model.model.layers.63.mlp.up_proj',
 'language_model.model.layers.63.self_attn.q_proj']

Running ONE real forward+backward pass through the WHOLE model...
loss: 2.1989
3/3 real target tensors' custom vjp fired during the real backward pass.
  language_model.model.layers.63.mlp.up_proj: H_I shape=(5120, 5120), H_O shape=(17408, 17408), count=1, NaN in H_I=False, NaN in H_O=False
  language_model.model.layers.63.self_attn.q_proj: H_I shape=(5120, 5120), H_O shape=(12288, 12288), count=1, NaN in H_I=False, NaN in H_O=False
  language_model.model.layers.62.linear_attn.out_proj: H_I shape=(6144, 6144), H_O shape=(5120, 5120), count=1, NaN in H_I=False, NaN in H_O=False

Part 5a PASS: real Hessian collection works end-to-end for all 3 real target tensors
simultaneously, in ONE real forward+backward pass, no NaN.

This is the first real confirmation that the gradient-capture mechanism works on a linear_attn/GatedDeltaNet tensor specifically, not just standard self_attn — a real, previously-untested case for this hybrid architecture.

Part 5b — real per-tensor correction: PASS, with a real performance finding and fix

Real command (single-tensor test):

"$PYBIN" 05_full_model_quantize.py --subset "language_model.model.layers.62.linear_attn.out_proj"

First real run — td_x=1 (matches Step 4’s own default):

Metric A (real trace-based, rotated): yaqa=2.903688e-05  naive=9.099063e-05  (REDUCES error by 68.09%)
Metric B (Frobenius, original basis): yaqa=8.8815  naive=4.0496

Real wall-clock: 12m 25s (20:43:50 → 20:56:15 BST), for one tensor, 491,520 real blocks (out_features × in_features/group_size = 5120 × 6144/64).

Real finding: the full plan is infeasible at this rate. Total real blocks across all 363 real tensors, computed directly from the plan’s own param_count/ group_size: 385,064,960. At the measured per-block rate, that extrapolates to roughly 130 real hours (5+ days) of continuous compute — language_model.lm_head alone (19.87M blocks, 6-bit) would take ~6.7 hours by itself. This is a real, hard blocker, not a nice-to-have optimization.

Real fix: generalize td_x. yaqa_core.ldlq2hess_quantize previously hardcoded td_x=1; generalized to accept td_x as a parameter. This is NOT a new algorithm — Step 3 already validated this exact formula structure with td_x=td_y=2 (not just 1) against an independent NumPy reference (03_rounding_update_test.py, M,N,TD=4,4,2), and the real code’s own default is td_x=16 (quantize_finetune_llama.py:47, read directly). mx.quantize’s own real docstring confirms grouping is per-ROW ("every group_size elements in a row of w are quantized together"), so a (td_x, group_size) block quantizes identically to td_x separate (1, group_size) calls — increasing td_x reduces Python-loop iterations without changing what mx.quantize computes.

Real re-run, td_x=16, same tensor:

Metric A (real trace-based, rotated): yaqa=2.909955e-05  naive=9.099063e-05  (REDUCES error by 68.02%)
Metric B (Frobenius, original basis): yaqa=8.9201  naive=4.0496

Real wall-clock: 67 seconds (21:00:17 → 21:01:24 BST) — an ~11× real speedup. The naive baseline is EXACTLY unchanged (9.099063e-05 both runs, as it must be, since naive doesn’t depend on td_x), and Metric A changed by only 0.07 percentage points (68.09% → 68.02%) — strong real evidence the td_x generalization is correct, not just “didn’t crash.”

A proper independent cross-check was then built and run (test_regression.py, test 4b), matching Step 3’s own standard rather than resting on the plausibility check above. First attempt: hand-wrote a NumPy replica of mx.quantize’s documented affine formula (s=(alpha-beta)/(2^b-1)) and compared against yaqa_core.ldlq2hess_quantize — failed, max|diff|=0.399. Root-caused directly (not guessed): the hand-written replica’s predicted scale (0.1420) did not match the REAL scale mx.quantize itself returns (0.1396) for bits=5, on a real standalone block, checked with no sequential loop involved at all — the real implementation does something beyond its own docstring’s simplified formula, and reverse-engineering the exact discrepancy would have been guessing, not verification. Real fix: rebuilt the independent check to call REAL mx.quantize/mx.dequantize from within the NumPy-orchestrated loop (converting thing to an mx.array and back each step) instead of replicating Apple’s primitive by hand — this isolates exactly what needed checking (the generalized loop’s indexing/slicing/formula correctness), using the same real, already-trusted quantizer on both sides. Real result: max|diff|=5.96e-07 (float32 machine epsilon) — the td_x generalization is now independently verified, not just plausible.

Honest full-plan extrapolation from this real data point (two bounds, not a guarantee): conservative (no fixed-cost amortization) ≈ 14.6 hours; realistic (the one-time model-load+backward cost is paid ONCE across all 363 tensors, matching Part 5a’s real architecture, not once per tensor) ≈ 3.5-4 hours.

Real finding: 52% of real tensors have no valid Hadamard factor — real fallback built

Running the 3-tensor subset with td_x=16 crashed on mlp.up_proj: get_hadK(17408) has no valid K-factor (17408 = 2^10 × 17, and 17 isn’t in the real algorithm’s divisor table {172,156,140,124,116,108,60,52,36,28,20,12}, nor is 17408 itself a power of 2). This is a real limitation of the real upstream algorithm itself — matmul_had.py’s own else-branch has the identical assertion — not a porting bug. Checked directly across all 363 real target tensors in the real plan (shapes read from the real loaded model, not the plan file, which has no shape info): 190/363 (52%) have at least one dimension with no valid Hadamard factor — every real mlp.up_proj/down_proj/gate_proj (intermediate size 17408) across all 64 layers, plus lm_head (vocab size 248320).

Real fix, not an invented mechanism: hadamard_transform.has_hadamard_support(n) (a non-throwing predicate reusing get_hadK’s exact real dispatch order) gates whether a given tensor gets real Hadamard rotation. When either dimension lacks support, 05_full_model_quantize.py falls back to plain (unrotated) LDLQ_2hess correction — yaqa_core.normalize_H (just the real mean-normalize step, no rotation) in place of rotate_symmetric_H, and the real, unrotated weight directly in place of rotate_weight’s output. This is exactly the code path that existed and was validated before the Hadamard fix was added — not new, unproven logic.

Real re-run, all 3 subset tensors, with the fallback (2026-09-03, 21:36:34 → 21:49:57 BST, 13m23s total for all 3 including the one-time model load+backward):

mlp.up_proj        (UNROTATED fallback): yaqa=5.31e-07  naive=2.79e-06  REDUCES error by 80.98%
self_attn.q_proj   (rotated):            yaqa=1.28e-06  naive=3.36e-06  REDUCES error by 61.94%
linear_attn.out_proj (rotated):          yaqa=2.91e-05  naive=9.10e-05  REDUCES error by 68.02%

Part 5b: 3/3 real tensors corrected successfully, 0 failed (1/3 ran without rotation).
3/3 tensors: YAQA reduces the real trace-based error vs. naive.

All 3 real architectural types in this hybrid model (self_attn, linear_attn/ GatedDeltaNet, mlp) now confirmed working, including the unrotated fallback path — and the unrotated tensor shows the LARGEST real error reduction of the three (80.98%), a real, honest result, not assumed or hoped for.

Fresh independent code review — one real bug found and fixed, one real gap made explicit

Reviewed yaqa_core.py, hadamard_transform.py, and 05_full_model_quantize.py in full against the real ground-truth source (finetune.py, ldlq.py, math_utils.py, matmul_had.py, custom_linear_B.py), requested explicitly given how many real bugs this session had already found.

Real bug found (HIGH severity): mx.random.seed(0) was called INSIDE the per-tensor loop, resetting MLX’s global RNG to the same fixed value before drawing SU/SV for every tensor. Since the RNG is deterministic given a seed, every tensor sharing the same in_features got a bit-identical SU, and every tensor sharing the same out_features got a bit-identical SV — e.g. every real self_attn.q_proj across all 64 real layers (same shape) would have received the exact same random-sign rotation. Harmless in Step 4 (only one tensor per run — this line was carried over from there unexamined) but silently defeats the whole real point of independent per-tensor incoherence processing at Step 5’s multi-tensor scale. Fix: seed once, before the loop, letting MLX’s RNG state advance naturally across tensors — matching the real code’s own per-layer-continuous-stream semantics (torch.manual_seed(idx), finetune.py:123). Verified the bug and the fix directly (not just re-running the model): simulated both behaviors on tiny arrays — old (reseed before each draw): SU1 == SU2 → True (bug confirmed real); new (seed once): SU1 == SU2 → False (fix confirmed working).

Real gap made explicit, not silently deferred: 05_full_model_quantize.py still calls loss_and_grad_fn(data[:1]), using only 1 of the real 6 stratified-calibration domain windows built — the per-example batch-loop fix from earlier this session (bug #2) is technically correctly written for batch>1 but currently inert. Raising k_per_domain alone would NOT fix this without also removing the [:1] slice. Deliberately kept at [:1] for now: an earlier real attempt to use the full 6-sequence batch on a SINGLE tensor already OOM’d this session (Part 4c) — risking that again on a multi-tensor run, on top of an unreviewed fix, was judged the wrong tradeoff right now. Real calibration scale-up (removing the slice, re-tuning sigma_reg to match, since the two are coupled) is real, disclosed follow-up work.

Real re-verification, all 3 subset tensors, with the seed fix (2026-09-03, 22:00:55 → 22:13:04 BST, 12m09s):

mlp.up_proj          (UNROTATED): 80.98% -- identical to before (unaffected by the fix, no SU/SV used)
self_attn.q_proj     (rotated):   61.94% -- identical to before (first rotated tensor processed, same first-draw)
linear_attn.out_proj (rotated):   66.56% -- CHANGED from 68.02% (now gets an independent rotation, not a repeat)

The change in linear_attn.out_proj’s result (68.02% → 66.56%) is itself real evidence the fix works — a genuinely different, independent rotation now, not a repeated seed.

Real finding: the reference repo’s deployment format is a trellis codebook, not affine quantization — rotation was never free to carry into deployment

2026-09-04. While building Part 5c (the real output model), a direct question surfaced: does deploying a Hadamard-rotated tensor need special inference-time handling, or can it be saved and loaded exactly like GPTQ’s output? Answering this required opening files that had never been read this session — the real repo’s inference module, not just its training script (lib/algo/finetune.py, already ported). This should have been read back when the Hadamard rotation was first added (Step 4), not discovered only now at the deployment stage — noted honestly, not glossed over.

Read directly (local clone: ~/HybridOptiQ_FINAL_BENCHMARK/research/quantization_methods_2026/yaqa/code):

What this means for this port: yaqa_core.ldlq2hess_quantize has always substituted mx.quantize() (MLX’s native per-group affine quantizer) for the real cb.quantize() (trellis codebook) at the exact same point in the correction loop — the block-sequential LDLQ_2hess math is faithfully ported, but the final rounding target was changed. This is a real, necessary substitution (MLX/Metal has no trellis-VQ kernel; building one is its own multi-week research/engineering project, not attempted here) — but it was made implicitly back at Step 3/4 when mx.quantize was first chosen, and should have been stated explicitly then rather than surfacing indirectly through a deployment question now.

Consequence for rotation: the real repo’s SU/SV/Hadamard rotation exists specifically to condition weights for a fixed trellis codebook — a VQ scheme with no per-group scale adaptation, so uncontrolled outliers are damaging to it. mx.quantize is per-group affine — it already fits its own scale/zero-point to each group’s actual observed range, which absorbs most of what rotation buys. This project’s own GPTQ baseline (08_gptq_apply_plan.py) is also affine, also unrotated, and already ships in production — real, direct precedent that an affine-quantized deployment doesn’t need runtime rotation to work.

Initial decision (v1, 2026-09-04, later reversed the same day): Part 5c would deploy tensors unrotated (Option B) to avoid the real, unverified engineering cost of a rotation-aware runtime module. Reversed after direct pushback (user, backed by an independent ChatGPT technical review) that Option A — deploying WITH rotation — was more tractable than assessed, and that dropping real algorithm fidelity for engineering convenience was the wrong call without first actually trying the derivation. That pushback was correct: the derivation below took under an hour and is fully verified.

Visual comparison of the real reference pipeline vs. this port’s pipeline: 00_DOCS/HTML/yaqa_pipeline_comparison.html (project’s real HTML docs folder).

Option A actually built: yaqa_core.YaqaQuantizedLinear, derived and verified

2026-09-04, same day as the reversal above. Hand-derived the real runtime forward pass from rotate_weight’s own formula (already gold-standard-verified against real PyTorch). Expanding rotate_weight(W, SU, SV) = matmul_hadUt(matmul_hadUt(W.T * SV).T * SU) algebraically: Wr = H_out^T · diag(SV) · W · diag(SU) · H_in (H_in/H_out being the orthogonal operators matmul_hadUt implements along each axis). Solving for W and substituting into y = x @ W.T gives:

x1 = x * SU
x2 = hadamard(x1)
y1 = x2 @ Wr.T          (quantized matmul against the ROTATED, packed weight)
y2 = hadamard(y1)
y  = y2 * SV

Verified numerically (not just algebraically): this exact sequence reproduces x @ W.T (the real, unrotated forward pass) to max|diff|=2.98e-07 (float32 precision) on a synthetic case — see test_regression.py test [8] for the permanent guard.

Real primitive discovery, changes the plan for the better: MLX ships mx.hadamard_transform (native Walsh-Hadamard, real Metal kernel, n = m·2^k for m ∈ {1,12,20,28}) and mx.quantized_matmul (native quantized matmul) — both public, documented, real functions (mlx==0.32.2, checked directly via help()). Two things checked before trusting them: (1) mx.hadamard_transform matches this port’s own hand-rolled matmul_hadUt exactly (0.0 diff) at a tested power-of-2 size; (2) real coverage across all 363 plan targets is identical between MLX-native support and the real repo’s own 12-divisor table — 173 tensors covered by both, 190 by neither, zero tensors covered by only one. No coverage tradeoff from standardizing on the native primitives for the runtime module.

yaqa_core.YaqaQuantizedLinear implements exactly the derived sequence above using mx.hadamard_transform/mx.quantized_matmul. Full pipeline verified end to end (not just the algebraic shortcut): real rotate → real regularized_block_ldl → real ldlq2hess_quantize → real packed codes → module forward, compared against the real corrected weight’s own forward pass — max|diff|=3.58e-07 (float32 precision). test_regression.py test [8] guards this permanently.

Real bug found and fixed while building this: mx.quantize is NOT idempotent on its own dequantized output

Part 5c’s first draft (Option B) re-quantized the assembled hatW after ldlq2hess_quantize to get deployable packed codes, on the assumption this was a lossless, idempotent operation (each block’s stored value already being the dequantized result of a real mx.quantize call). Checked directly, found wrong: mx.quantize(mx.dequantize(mx.quantize(W))) does NOT reproduce the same codes as the first quantize — max|diff|=3.7e-3 on a synthetic 128×256 array (~18% of packed int32 words changed), stabilizing only after a SECOND re-quantization (not floating-point noise; a real, one-time rounding drift). This threatened the correctness of Part 5c’s whole pack step, for both the rotated and unrotated paths.

Fix: yaqa_core.ldlq2hess_quantize now captures the real packed (wq, scales, biases) directly from each block’s own mx.quantize() call during the sweep, returning them alongside hatWr, instead of reconstructing them via a second quantize call afterward. Verified safe by first checking that each block position is visited exactly once across the real anti-diagonal sweep (checked directly on a synthetic 5×7 grid — zero double-visits), so no overwrite conflict is possible. test_regression.py guards both the bug staying fixed (the returned codes dequantize to EXACTLY hatWr, zero tolerance) and the original NumPy cross-check (unaffected by the fix, still passes).

Part 5c — real output model, --incoherence {none,hadamard}

05_full_model_quantize.py --output <dir> does a real save. --incoherence hadamard (default): the 173/363 rotatable tensors deploy via YaqaQuantizedLinear using the real packed codes directly in the rotated basis (no inverse-rotation, no re-quantize); the other 190 deploy via standard to_quantized(), same as before. --incoherence none: every tensor deploys via standard to_quantized(), matching 08_gptq_apply_plan.py’s save path exactly. Either way, packing uses the real captured codes (the fix above), not a re-quantize.

Any plan tensor a --subset run doesn’t cover is left at full BF16 (disclosed, not silently native-quantized) rather than paying real cost for a run whose whole point is to validate the save mechanics cheaply; a full run (no --subset) has no such tensors. Config dict mirrors 08_gptq_apply_plan.py’s own structure, including its BUG FIX #3 (top-level bits/ group_size defaults mlx_lm.utils.load()’s reload path reads unconditionally) — rotated tensors are recorded as False in that config (so mlx_lm’s own _quantize() doesn’t try to misreconstruct them as plain QuantizedLinear), plus a separate yaqa_rotated_manifest.json written into the output directory listing every rotated tensor’s real bits/group_size/shape.

Real calibration richness fix in the same change: the previously-disclosed data[:1] gap (Part 5a) is closed by chunking calibration into --batch-size-sized real forward+backward calls, accumulating H_I/H_O across calls the same way the vjp already accumulates across the batch dimension within one call — peak activation memory now bounded by --batch-size regardless of total calibration size (raising --n-calibration alone, on one giant single-shot call, would not have been enough — that’s what OOM’d a single tensor in Part 4c).

The real, unresolved gap: mlx_lm.utils.load() has no concept of YaqaQuantizedLinear and cannot reconstruct it from yaqa_rotated_manifest.json — a real, separate loader is required before a --incoherence hadamard model is actually runnable. The script prints a loud warning and writes the manifest, but does not claim the model is loadable.

Production decision (2026-09-04, later the same day): after direct review of the real compatibility/runtime tradeoffs — Option A needs a new loader everywhere the model is ever loaded (not just here), plus a permanent per-forward-pass Hadamard-transform cost on 173/363 tensors, for an unproven quality benefit against MLX’s self-adapting affine quantizer — --incoherence default flipped from hadamard to none. hadamard remains a real, fully built and verified research branch (not deleted, not degraded) to benchmark later, after cascading + stratified-KL work, before deciding whether the loader is worth building.

CRITICAL real bug found and fixed: the wrap/restore mechanism was silently broken the whole time — Part 5c’s first real save contained ZERO correctly-quantized tensors

2026-09-04, found by actually inspecting a real saved model’s safetensors keys instead of trusting exit code 0 (the user explicitly demanded this level of verification after the idempotency bug above). The first real --output smoke test exited 0 and printed a normal summary. Direct inspection of model.safetensors.index.json showed every one of the 3 target tensors saved under a nonsense key: language_model...mlp.up_proj.module.weight instead of language_model...mlp.up_proj.weight — and with only weight/scales/biases present, no SU/ SV, meaning even the two tensors meant to be rotated weren’t.

Root cause: model.leaf_modules() defines “leaf” as a Module with no further nn.Module children — not “the position named in the plan”. Before wrapping, a target position (e.g. mlp.up_proj) holds a plain nn.Linear, which qualifies as a leaf named mlp.up_proj. Once wrapped in YaqaCatcher (which itself has a .module child holding the real nn.Linear), that position no longer has zero Module children — leaf_modules() stops treating it as a leaf and instead descends into it, reporting mlp.up_proj.module as the leaf. The restore step used to re-query tree_flatten(model.leaf_modules(), ...) after wrapping, so it saw this shifted key set; orig_modules.get("mlp.up_proj.module", l) always missed (orig_modules is keyed by the real, unshifted plan names), silently fell back to l (the unchanged inner Linear it had just found), and the “restore” was a complete no-op — the model stayed wrapped in YaqaCatcher for the rest of the run, silently, for every run this session that reached this code path.

Reproduced cheaply first, before touching the real file: a 2-layer toy model with the same wrap/restore sequence showed model.a was still Catcher after “restore” — confirmed the bug in under a second, no need for the 27B model to diagnose it.

Consequence for the real smoke test: Part 5c’s all_leaves = tree_flatten(model.leaf_modules(), ...) call (used to decide what to save for every leaf) saw the same shifted keys for all 3 target tensors. if name in q_layers never matched (comparing the real plan name against a shifted key), so every real YAQA-corrected tensor fell through to the fallback-native-quantize branch — wrong bits (--fallback-bits=6 instead of the real plan’s 8/8/5), wrong weight (plain native rounding, no YAQA correction at all), saved under the wrong key. The real Hessian-collection metrics reported earlier in this ledger (75%, 83%, 68.65%, etc.) are unaffected — those read orig_modules[name].weight directly, never through the corrupted live tree — but every real SAVE this session produced before this fix is invalid and has been deleted.

Fix: capture the pre-wrap leaf key list once (orig_leaves_list), reuse it for both the wrap and the restore — never re-derive leaf positions from model.leaf_modules() after the tree has been mutated by wrapping. Added a hard runtime check immediately after restore that verifies every target position is back to its original type, raising loudly instead of silently continuing if it isn’t. Verified the fix on the same toy-model reproduction (restore now correctly shows Linear, and a subsequent quantized replacement lands under the clean a.weight key, no .module. shift) — test_regression.py test [9] guards this permanently, cheaply, without needing the real model. Real smoke test re-run with the fix in progress — result recorded below once verified against actual saved, reloaded output, not before.

Real memory catastrophe found, fixed — then a second real crash found, fixed differently

2026-09-04, after Part 5c’s first real save actually loaded and ran correctly. Before trusting --n-calibration 4 for a real run, computed the real total: all 363 plan targets’ H_I+H_O held simultaneously need 557GB (310GB even excluding lm_head, whose 248,320-wide output dimension alone needs 246.7GB — excluded separately, matching the real reference repo’s own precedent of not running its per-layer correction loop on this tensor either). The earlier claim in this file that “128GB comfortably covers the largest real Hessian” checked only the single largest tensor, never the sum across all targets at once — found only after a real --n-calibration 4 run OOM’d (exit 137).

Fix: --hessian-batch-budget-gb splits targets into memory-bounded batches, each getting its own real forward+backward pass over the full calibration data.

Second real crash, found immediately by testing the fix on real data: batch 2 of a forced 3-batch test failed with RuntimeError: [metal::malloc] Resource limit (499000) exceeded — a hard, real cap Metal puts on GPU resource handles created over a process’s lifetime, not a byte-size limit. Tried mx.clear_cache() (real, native MLX/Metal API) — tested directly, confirmed it does NOT reset this (identical crash, same batch, same spot, re-run twice).

Real fix: each batch now runs as its own separate OS process. New CLI plumbing on 05_full_model_quantize.py (--resume-dir, --batch-only N, --print-num-batches, --assemble-from-resume) plus a new orchestrator, run_full_yaqa.sh, that queries the real batch count, runs each batch as a fresh process (writing results to a resume-cache directory, same real idiom as 08_gptq_apply_plan.py’s own ResumeCache), then does one final --assemble-from-resume pass that loads every batch’s cached result and performs the real save. Proven on real data: re-ran the same 3-batch test with the fix — the exact tensor that crashed twice (self_attn.q_proj) completed cleanly as its own process (78.20% real error reduction). Full 3-batch run completed, assembled, saved, and — the actual bar for “proven,” not just “didn’t crash” — loaded via plain mlx_lm.load() and ran a real forward pass, predicting “Paris” for “The capital of France is.”

A real bug found in the fix itself, found and fixed the same session: the resume-cache manifest write originally happened before err_yaqa/err_naive/pct_reduction/frob_* were computed, so --assemble-from-resume had only NaN placeholders for those fields — the final “N/M tensors reduce error” summary printed “0/M” even though every tensor’s own per-batch log showed real, large reductions (83%, 78.2%, etc.). Wrong summary line, not a wrong saved model — the packed weights were always correct; only the reported count was wrong. Fixed by moving the manifest write to after the real metrics are computed, and verified cheaply (synthetic manifest, no real model load needed) rather than paying for another real run to check a pure bookkeeping fix.

Real, honest gap: the architecture does not match the reference’s actual design

2026-09-04, raised directly by the user, verified by reading the real source rather than arguing from memory. The batching fix above solves the real memory and Metal-resource crashes, but it does so by re-running a full forward+backward pass through the entire 64-layer model once per arbitrary memory-sized batch — for the real full run (~39 batches at the default 8GB budget), that is ~39 redundant passes through every layer.

The real reference (quantize_llama/quantize_finetune_llama.py, read directly, lines ~160-215) does not do this. It processes layers sequentially, one at a time (for i in range(len(model.model.layers))), streaming each layer’s real output forward as a cache for the next layer’s input — each layer’s forward pass runs exactly once, not once per arbitrary batch. (Nuance checked, not assumed: this streaming does NOT feed already-quantized layers forward into later layers’ calibration — each layer’s input still traces back through the original, unquantized chain — so the real design is a genuine efficiency/memory optimization, not a “cascading correction” scheme layered on top.)

This was not caught during design review — it surfaced only when directly challenged on why the port wasn’t processing the network “like MLX OptiQ does, one tensor at a time.” Batching by arbitrary memory size instead of by network structure is a real, disclosed architectural gap: it produces correct numbers (proven above) through a materially less efficient path than the real algorithm’s own design.

Follow-up, same session, checked directly rather than assumed: read the real Hessian-collection driver itself, hessian_llama/get_hess_llama.py (a separate script from quantize_finetune_llama.py, not read before this point). Real finding: the reference’s OWN collection stage does NOT stream layer-by-layer either — it wraps every layer’s tensors simultaneously and runs ONE real forward+backward pass per calibration batch (torch.nn.functional.cross_entropy(logits, fake_target).backward()), accumulating every layer’s Hessian at once, exactly like this port’s original (pre-batching) design. It fits on their real hardware via FSDP (sharding the model and its Hessians across multiple GPUs) plus CPUOffload(offload_params=True) — infrastructure a single 128GB Mac does not have. Their n_seqs default is 65,536 sequences at ctx_size=2048 (~134M tokens) — real production scale, not something a single machine holds regardless of architecture.

Honest conclusion: the 557GB simultaneous-Hessian requirement is not a porting mistake — it is the real algorithm’s actual behavior at real production scale, made feasible upstream by multi-GPU hardware this project doesn’t have. Some form of splitting the work is genuinely unavoidable here, not a wrong design choice. What WAS a real, fixable mistake: splitting arbitrarily instead of along the network’s real structure. Fixed the same session: targets are now sorted by real layer index (extracted from the real dotted tensor name) before batching, so tensors from the same real transformer layer land together wherever the memory budget allows, instead of being scattered by whatever order the plan’s JSON happened to list them in. Real per-layer streaming (avoiding repeated full-model passes entirely, matching neither reference script exactly since both hold everything at once) remains real, deliberate, disclosed future work — not something either reference script actually does either, so not a gap relative to them specifically, but still a real efficiency opportunity for this project’s own single-machine constraint.

Real follow-up: lm_head no longer gets zero correction – real GPTQ, scoped safely

2026-09-04, prompted directly by the user asking why the OOM fix couldn’t be applied to lm_head too. Direct, correct answer worked through with them: batching fixes “too many tensors’ Hessians held together at once” – it cannot shrink one tensor’s own Hessian. lm_head’s H_O is 246.7GB by itself, in a batch of exactly one, so batching was never going to reach it.

The real, different fix: lm_head’s problem is specifically its OUTPUT side (tied to the 248,320-word vocabulary). Its INPUT side (in_features=5120) is completely normal. GPTQ only ever needs the input side (H = x.T@x) – it was never going to hit this wall in the first place. Confirmed directly against the real source: 08_gptq_apply_plan.py has no special exclusion for lm_head at all; it already corrects it automatically as an ordinary nn.Linear.

Checked before reusing 08_gptq_apply_plan.py’s own gptq_quantize_with_plan() directly, not assumed safe: that function wraps every real nn.Linear/SwitchLinear leaf simultaneously (497 of them). Real sum of their H matrices (in_features² each, GPTQ never needs H_O): 125.93GB – dangerously close to this machine’s full 128GB. Reusing it as-is would have risked a third real OOM. Rejected in favor of a properly scoped alternative.

Built: 06_lm_head_gptq.py – wraps only lm_head with Apple’s own real mlx_lm.quant.gptq.Catcher (read directly, unmodified, first-party), collects its real H from a real forward-only pass over calibration data (no backward pass needed – GPTQ never needs gradients), runs the same real column-by-column compensation math as 08_gptq_apply_plan.py’s own function (reproduced here rather than imported, since the original nests it as local functions), and packs the result via mx.quantize – native MLX, same format as every YAQA-corrected tensor, not GPTQ’s own custom bit-packer – so it lands in the exact same manifest schema with zero special-casing needed downstream.

Real safety property, explicitly checked and answered when asked: does this contaminate the other 362 real tensors’ measurements? No – verified this follows directly from the SAME real property already regression-tested (test [10]): every batch reloads the pristine original BF16 model fresh, never a partially-corrected one, so lm_head’s calibration pass (run separately, at a different time) sees exactly the same untouched original weights every other tensor’s measurement was anchored to. Order and timing don’t affect correctness, only scheduling.

Manifest entry format: adds "method": "gptq" (the 362 YAQA entries carry no "method" key, implicitly YAQA) plus frob_gptq/frob_naive in place of the Hessian-weighted err_yaqa/err_naive/pct_reduction fields (a one-sided correction has no real two-sided-weighted metric to report – 05_full_model_quantize.py’s existing .get(..., float("nan")) fallbacks already handle this honestly, verified with a synthetic mixed yaqa+gptq manifest, no code change needed there).

Streamlined as a standing rule, not a one-off: wired into run_full_yaqa.sh to run automatically after every batch completes and before the final assembly, every time – never a manual step to remember. Also made resume-aware (same real per-tensor skip pattern), so an interrupted run doesn’t redo it.

Result: lm_head is real, GPTQ-corrected – not YAQA’s full two-sided correction (structurally impossible on this hardware), but strictly better than the zero-correction fallback it had before, for the same real reason GPTQ-vs-naive already shows real measured improvement on every other tensor in this project.

Real universality: no more model-specific hardcoded constants

2026-09-05. --source and --plan were already real overrides. Found and fixed the remaining one: LM_HEAD_NAME (the vocabulary-output tensor’s dotted name) was still a hardcoded constant in both 05_full_model_quantize.py and 06_lm_head_gptq.py. Now --lm-head-name on both, forwarded automatically by run_full_yaqa.sh. This codebase no longer assumes Qwen3.5/3.8’s own naming for any of its architecture-defining constants — a different MLX model with its own real plan can be pointed at directly.

A real name for the method: YAQA-UMA

Proposed and documented in 00_DOCS/HTML/yaqa_batching_architecture.html: YAQA-UMA — YAQA for Unified Memory Architecture. Not decorative: Apple Silicon’s real, shared CPU/GPU memory design is the specific constraint the batching + resume + separate-process architecture in this ledger was built to work within, as distinct from the reference’s own multi-GPU + CPU-offload design. That page also carries the real, spatial batching-architecture diagram (mermaid), a table of what’s been verified end to end (not assumed), and the real numbers behind the memory problem (557GB needed, 128GB available, 246.7GB for lm_head alone).

NOT YET DONE

Real RCA: a real, verified 302x correction failure on one unrotated tensor (2026-09-05)

Trigger: the live mission-control dashboard’s dynamic quality narrative surfaced a real outlier – language_model.model.layers.14.linear_attn.in_proj_qkv showed pct_reduction = -1788.48% (YAQA’s weighted error ~18.9x worse than naive). User correctly refused to accept the pipeline’s own self-reported number at face value and demanded independent verification.

Independent verification (not reusing any pipeline error-computation code): loaded the real original weight fresh from the source model, ran a fresh mx.quantize for naive, loaded the real cached YAQA result and dequantized it separately, computed Frobenius error myself in a one-off script:

Independent naive Frobenius: 2.543   (manifest: 2.594 -- matches, small dtype rounding diff)
Independent YAQA Frobenius:  767.92  (manifest: 768.0  -- matches, confirms NOT a reporting bug)
mean |original weight value|:  0.0126
mean |naive dequant error|:    0.0003   (2.3% of typical weight size -- normal)
mean |YAQA dequant error|:     0.0291   (231% of typical weight size -- genuinely broken)

Confirmed real, not a metric/reporting artifact.

Root cause, precisely (from a real code review of ldlq2hess_quantize/regularized_block_ldl in yaqa_core.py): the two-sided correction propagates each block’s rounding error forward via Lout/Lin, derived from block_ldl_mlx’s Cholesky-based factorization of the real Hessian. If the Hessian is ill-conditioned (a few directions with much larger eigenvalues than the rest), the diagonal-block inverses used to build L can become large, which makes the propagated-error term amplify the residual instead of damping it – a real, mechanical explanation for a 302x blowup, not a vague “instability.”

Hadamard rotation’s real, documented purpose in this method family (QuIP/QTIP/YAQA) is specifically to prevent this: it spreads the Hessian’s energy evenly across directions, keeping the Cholesky blocks well-conditioned. This tensor’s own log line: NOTE: has a valid Hadamard factor but --incoherence=none -- deploying unrotated by choice, not limitation. – real, deliberate production choice, real quantified cost on this one tensor.

The real gap in the existing safety net: regularized_block_ldl’s retry loop (real code, re-read for this RCA) only checks mx.isfinite – it retries with more damping ONLY on NaN/Inf. This tensor’s block_LDL reported succeeded on attempt 1 – technically finite, still catastrophic. The safety net catches outright numerical failure, not “valid-looking but insane magnitude.” That is the real, fixable gap, not “needs rotation” alone.

Options considered, with real tradeoffs: | Option | Real verdict | |—|—| | A. Fallback to naive for verified losses (err_yaqa >= err_naive in the manifest) | Shippable tonight. Zero risk, uses data already computed, purely additive safety net. | | B. Build the YaqaQuantizedLinear loader, rotate eligible tensors | Off the table for this run, not just “more work”: mlx_lm.utils.load() cannot load a model containing ANY rotated tensor at all (real, hard incompatibility, confirmed directly in the module’s own docstring) – rotating even one tensor makes the whole saved model unloadable by standard tooling until a separate loader exists. Real future work, zero effect on tonight’s ship. | | C. Harden regularized_block_ldl’s own check (retry with more damping if the result would be worse than naive, not just on NaN/Inf) | Real, generalizable fix for the actual root cause – catches this failure mode automatically for any future tensor/model. Needs its own real testing before trusting it on a live run; disclosed as follow-up, not applied tonight. | | D. Raise the global sigma_reg default | Rejected – already documented as tuned for the old single-sequence calibration scale; a blind global change risks the 287 tensors already working well. |

Decision: A now (real, safe, additive), C as real disclosed follow-up work, B not pursued for this run given the hard loader incompatibility.

MLX-native note: the correction sweep is inherently sequential (each block depends on the previous block’s residual – the algorithm itself, not an MLX limitation), so there is no real vectorization opportunity in the core loop. The one genuine, real, MLX-aware improvement identified: a cheap pre-flight conditioning check (e.g. mx.linalg.eigvalsh on the normalized Hessian, or the max diagonal-block-inverse magnitude) before running the full expensive sweep, flagging a dangerously ill-conditioned tensor ahead of time instead of discovering it only after computing a correction that gets discarded anyway.


Real test: mx.clear_streams() as an alternative to one-process-per-batch (2026-09-05)

Tested directly, isolated from the live run (separate resume dir, real model, real data): running multiple real batches inside one process, calling mx.clear_streams() between them, to see if it could avoid the separate-process-per-batch architecture. Result: it does not work — batch 0 completed fine, but clear_streams() destroys the default GPU stream MLX itself needs, so batch 1 crashed immediately with RuntimeError: There is no Stream(gpu, 1) in current thread (a different, earlier failure than the original resource-limit crash it was meant to prevent). Confirmed: the current separate-process- per-batch design remains the best real approach given this machine’s constraints — no cheaper alternative found.

Real follow-up: the actual, verified root cause of the 302x blowup was BF16 precision, not just ill-conditioning (2026-09-06)

This entry corrects and completes the RCA above. The 2026-09-05 entry correctly identified that layers.14.linear_attn.in_proj_qkv has an ill-conditioned Hessian and that the correction amplifies error under those conditions — but it did not yet know why the amplification was so extreme on this specific run. Direct investigation the following day found the actual mechanical cause, and it changed the fix.

Real finding: the Hessian curvature accumulation (yaqa_linear_vjp in 05_full_model_quantize.py) was running in bfloat16, not float32. x and grad_output are this model’s native bf16 compute dtype; the code assumed the matmul Gb = gb.T @ xb would promote to float32 the way float32 + bfloat16 addition does in MLX. It doesn’t — confirmed with a direct, isolated dtype probe: bf16 @ bf16 stays bf16, no auto-promotion, unlike addition.

Real, controlled 3-way measurement on this exact tensor, bf16 vs. fp32 vs. a CPU-float64 gold reference:

bf16 (broken) fp32 (fixed) float64 (gold reference)
H_I min eigenvalue -5.618 (real, large-magnitude PSD violation) ~-3e-4 (numerical noise floor) matches fp32 to 4 decimals
H_O min eigenvalue -7.534 ~-8e-4 matches fp32 to 4 decimals
YAQA weighted error vs. naive -0.05x (i.e. worse than naive) 0.0009x 0.0009x (identical to fp32)
Frobenius error vs. naive 3.84x (worse than naive) 1.0225x 1.0225x (identical to fp32)

A true sum-of-Gram-matrices Hessian is positive semi-definite in exact arithmetic — eigenvalues of -5.6/-7.5 are six to seven orders of magnitude past what float32’s own rounding noise could ever produce, which is what made this a real, provable precision bug rather than “the Hessian is just naturally a bit noisy.” fp32 reproduces the float64 gold reference’s real downstream objective to 4 decimal places — no CPU float64 pipeline is needed in production; float32 already matches it and runs natively on GPU.

The fix: cast x and grad_output to float32 before forming Gb (casting Gb afterward is too late — the precision is already lost in the matmul itself), in both the YAQA curvature path and, separately, the GPTQ path used for lm_head. Added CURVATURE_VERSION = 2 so every pre-fix cached tensor (bf16 curvature) is automatically recomputed on resume, rather than requiring a manual cache wipe.

How this reconciles with the 2026-09-05 ill-conditioning RCA, precisely: both are real and both matter, they’re not competing explanations. Ill-conditioning is why this specific tensor was so sensitive — a well-conditioned Hessian would have absorbed the same bf16 precision loss without a visible effect. BF16 precision loss is the actual mechanical cause of the corrupted computation. The fix that actually worked (FP32) treats the real cause; Option C from the prior entry (harden regularized_block_ldl’s retry logic) is no longer the load-bearing fix for this specific failure, though it remains reasonable defense-in-depth and is still real, disclosed follow-up work.

Real production impact: this wasn’t a one-tensor fix. All 336 tensors already processed under the bf16-curvature bug get auto-invalidated by CURVATURE_VERSION and recomputed correctly on the current production run. Fully documented, plain-English, with diagrams: research_hadamard_blowup/exports/BF16_Curvature_Bug_Plain_English.html (and its from-first-principles technical companion, YAQA_UMA_Hessian_Precision_Root_Cause.html).

Real production build, two real post-build bugs found and fixed, real rebuild + re-verification (2026-09-07)

Missing from this ledger until 2026-09-07 (later) — added retroactively in its real place. The real 363-tensor production run (Step 5-full, above) completed for the first time this day. Two real, measured bugs surfaced only after benchmarking the finished model, not before:

  1. Missing mode: "affine" in config.json — every real per-tensor quantization config dict this script wrote only ever carried {"bits", "group_size"}, never mode, even though every real mx.quantize()/to_quantized() call always used MLX’s own affine default. Packed weight bytes were never wrong — only the metadata describing them was incomplete, matching neither the reference build’s config.json nor what mlx_lm’s own loader expects declared explicitly. Real, measured cost: ~27-87% slower decode across every real benchmarked depth (AR 14.06 vs 17.91 tok/s; best-depth 28.83 vs 54.08 tok/s). Full evidence, every real quantize/dequantize call site checked exhaustively: research_hadamard_blowup/exports/AFFINE_MODE_CODE_REVIEW.html.
  2. Scales/biases saved as float32 instead of bfloat16 — found while re-benchmarking after fix #1 didn’t close the gap. Every one of 727 real scale/bias tensors in the built model was double the correct storage width (the correction math deliberately runs in fp32 for precision; nothing downcast the output metadata back to bf16 before saving, unlike the reference pipeline which runs bf16 throughout). Confirmed via direct tensor forensics (real safetensors headers read on disk, not estimated): 0/1912 bits/group_size/mode mismatches (never a bit-width bug), 727/1912 dtype mismatches (every one F32-vs-BF16). Real, measured cost: +17.2% peak memory, 20-55% decode slowdown. Same investigation found embed_tokens was being quantized (a real, separate confound for apples-to-apples comparison — the reference leaves it uncompressed) — fixed to match. Full RCA, comprehensive tensor table, exact rebuild command: research_hadamard_blowup/exports/RCA_scales_f32_bug.html.

Rebuilt via --assemble-from-resume after each fix — the real 26-hour Hessian correction was never re-run; each rebuild just re-saved the already-corrected resume-cache tensors at the right width/config. Independently re-verified against the actual rebuilt model, not just code-reviewed: tensor forensics re-run shows 0/1910 tensors differing from the reference in dtype/bits/group_size/mode/size; a real 4-way weight-provenance check (original source weight vs. naive quantization vs. resume-cache correction output vs. live shipped model) confirms the real corrected values shipped, not naive quantization standing in for them (live matched resume-cache to bf16-rounding precision, ~0.1-0.6%; differed from naive by 1.9-4.5% absolute). Resume-manifest analysis: 362/363 tensors used real YAQA correction (the 1 exception, lm_head, is the known, expected naive_fallback), real Hessian-weighted error reduction vs. naive ranged 97.49-100%, median 99.67%, zero tensors worse than naive.

Re-benchmarked (mtplx tune): AR and D2 now at real parity with the reference (-0.14%, -1.06%); D1 and D3 still lag (~-20%), with real acceptance-rate degradation concentrated at deeper speculative-decode windows (D3 depth-3: reference 94.74% vs. YAQA 80.60%). Not yet explained by any single localized failure in the manifest analysis above. Real output-quality benchmark (same 5-task protocol as every other model in scripts/yaqa_port/exports/Benchmark_Master_Recap.html) launched 2026-09-07 ~16:04 to test this independently of tok/s; MMLU result already in — 87.7% (100/114, 95% CI ±6.0pp) vs. the v33floor reference’s 89.5%, a real -1.8pp gap within the stated CI. Remaining 4 tasks in progress as of this entry.

Also fixed the same day, project-wide: every HTML doc built via build_html_docs.sh was rendering its own title twice (pandoc’s title-block header plus the body’s own matching first heading) — root-caused and fixed once in the script itself; see CHANGELOG.md’s 2026-09-07 “even later” entry.

MTP sidecar YAQA correction (2026-09-08)

The trunk correction above left the MTP speculative-decode sidecar naive-quantized (plain mx.quantize, identical to the reference baseline) – a real architectural gap, not a design choice: a YAQA-corrected trunk feeding a naive-corrected draft head. Root-caused, planned, implemented, smoke-tested, and launched as a real full build this same day – full detail kept in CHANGELOG.md’s 2026-09-08 entries (not duplicated here) and research_hadamard_blowup/exports/MTP_YAQA_CORRECTION_PLAN.html (the original plan and real evidence). Real launch command and the two real monitoring-gap fixes it surfaced (monitor.sh process-match, generate_dashboard.py missing MTP milestone) are in CHANGELOG.md’s “Launched (2026-09-08, later)” entry.

Two further real gaps found and fixed while actually running the launch, both in CHANGELOG.md’s “later still” entries: a --reuse-cached-fallback skip option added to 06_lm_head_gptq.py (lm_head’s damping search is deterministic, so retrying a prior naive_fallback result wastes ~15-20 real minutes reaching the identical outcome) plus the matching argparse fix 05_full_model_quantize.py needed to not crash on that same forwarded flag; and a real cp -a-into-existing-directory pitfall that silently doubled the resume cache’s disk usage (16GB -> 32GB) when Step 1 of the launch procedure was re-run a second time. Exact, current procedure (all steps, all flags): YAQA_COMMAND_REFERENCE.md’s “Adding YAQA-corrected MTP sidecar” section.


© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.

3 October 2026 — final n=200 comparison and benchmark-page update

GODMODE-v2 completed all five tasks at n=200 (mean 92.58, against 93.04 for HESSIAN-PROBE). The benchmark recap, the floor analysis and the master timeline were updated to match. Chapter 11 of the research ledger records the GODMODE builds and results. Open: Stratified at n=200, the V3 build, and the BF16 reference with KL.