2026-09-03
Author: Hakim Ghelab, VegaLaboratories LTD
This is the resume point. If a session runs out of context, read this file first — it has exactly what’s been verified, the real commands to re-confirm it, and exactly what step comes next. Every command below was actually run; every output shown is real, not illustrative.
What we’re shrinking, and why that’s risky. An AI model is, physically, a huge list of numbers — for this project’s model, about 27 billion of them. Making the model smaller and faster means rounding those numbers to fewer decimal places, the same idea as rounding $19.97 down to $20. Do that carelessly across 27 billion numbers and the model visibly gets dumber. Do it carefully and it barely notices. This whole project is about doing it carefully.
GPTQ — the smart rounding method already working in this
project. When GPTQ rounds one number, it looks at the other
numbers sitting right next to it and nudges them slightly to cancel
out the rounding mistake it just made — so the small local group stays
accurate even though each individual number moved. This is already
built, tested, and documented (see
docs/HTML/GPTQ_Why_It_Matters.html).
YAQA — what this specific port is about. GPTQ only checks “did I mess up this one small local group of numbers.” YAQA checks something bigger: “did I mess up what the entire model’s final answer looks like.” That’s a real, meaningfully harder thing to measure — like proofreading a whole essay’s actual meaning, instead of only spell-checking each sentence in isolation. The real published result: YAQA’s rounding is measurably more accurate than GPTQ’s — about 30% less error, in the paper’s own tests.
What “porting to MLX” means. YAQA was built and only ever tested using Nvidia’s toolkit (CUDA/PyTorch). This project runs on a Mac, using Apple’s own toolkit, MLX. Those two toolkits don’t speak the same language internally — nothing automatically translates one into the other. “Porting” means manually rewriting YAQA’s actual logic so it works inside MLX instead, piece by piece.
Why we don’t just download a finished MLX version. Because one doesn’t exist. Nobody has built this translation before (verified directly — a real search this session found no existing MLX port anywhere). That means every piece has to be built AND proven correct by us, not assumed to work.
The method: prove each small piece before trusting the whole thing. You would never trust a translated legal contract without checking each paragraph. Same idea here — each of the 5 steps below proves ONE piece works, on numbers small enough to check by hand, before it’s ever trusted on the real, 27-billion-number model. This is slower than just writing the whole thing and hoping — but a silent, subtle math mistake here would not crash anything. It would just quietly make the model worse in a way nobody would notice until much later. That’s the exact failure this step-by-step discipline exists to prevent.
Where we actually are right now: All 5 steps are
done and proven — the pipeline works, end to end, across the real, full
model, not just one layer. Steps 1-4 run correctly with real, verified
numbers at every stage (an independent code review on 2026-09-03 caught
two further real problems in what had looked like a finished Step 4 — a
silent Hessian-formula bug and a missing Hadamard incoherence-processing
stage — both fixed and independently verified against the real running
algorithm; YAQA’s correction reduces the real, Hessian-weighted
error metric by 75% versus naive rounding on the real
k_proj layer, see Step 4’s Part 4c write-up below). Step 5
— scaling this to the entire model — completed its first full run on
2026-09-07: all 363 real tensors processed, 362/363 using real YAQA
correction (the 1 exception, lm_head, is the known,
expected naive fallback), real Hessian-weighted error reduction versus
naive ranging 97.49-100% (median 99.67%), zero tensors worse than naive.
Two real post-build bugs (a missing mode: "affine" config
field, and scales/biases saved at the wrong bit width) were found by
benchmarking the finished model, root-caused, fixed, and the rebuild
independently re-verified via real tensor forensics. A further
YAQA-corrected MTP speculative-decode sidecar variant was built and
re-tuned on 2026-09-08 (mtplx tune, run
tune-20260908-050223): AR 17.70 tok/s, D1 39.14 tok/s
(2.21x), D2 51.68 tok/s (2.92x, the real best), D3 48.23 tok/s (2.73x —
slower than D2, not a further win). D1 and D2 land at real near-parity
with the reference build; D3 does not scale further. Real testing on
non-benchmark prompts (difficult, real-world questions) found generation
speed largely unchanged across tuning depths — a genuine, disclosed
limit of the AR>D1>D2>D3 tuning protocol as a
predictor of real-world speed, not a claim this fully resolves. The open
thread is that gap between benchmark and real-world behavior, not
whether the pipeline itself works — it does, and the port is real: the
first known working port of YAQA to MLX.
What “success” looks like at the very end (Step 5): a real, working tool that takes this project’s full-size AI model and produces a smaller version rounded more carefully than what GPTQ already achieves — verified with real before-and-after comparisons on the actual model, not just “the code ran without crashing.”
model.train()) that automatically
reroutes those layers onto a two-way street. Without finding this, Step
4 would have been simply impossible on this actual model — not a bug we
introduced, a real fact about how this model is built.| Term used below | What it actually means |
|---|---|
| Weight | One of the millions of numbers that make up the model’s “knowledge” |
| Quantization | Rounding those numbers to fewer decimal places, to save space |
| Bit-width | How much precision is kept per number after rounding |
| Hessian / sensitivity map | A map of how much a mistake in one number messes up the final answer |
| GPTQ | Smart rounding that checks only within one small local group of numbers |
| YAQA | Smarter rounding that checks against the model’s whole final answer |
| MLX | Apple’s own toolkit for running AI models on a Mac |
| Porting | Rewriting something built for one toolkit so it works on a different one |
| Gradient / backward pass | The calculation that traces “if I nudge this number, how much does the final answer change” |
| Cholesky / block-Cholesky | A mathematical way of breaking one big sensitivity map into smaller, manageable pieces |
| Synthetic / toy example | Small, fake numbers used only to check the logic is correct — never real data |
H_I / H_O |
This project’s own shorthand for “the sensitivity map, input side” / “…output side” |
Everything below this point is the detailed technical log — real commands, real output, real file paths — for anyone who wants the actual evidence behind each claim above.
| Step | What it proves | Status | File |
|---|---|---|---|
| 1 | MLX’s mx.custom_function gives simultaneous access to a
layer’s real forward input AND real incoming gradient — the exact
mechanism YAQA’s real backward pass needs |
✅ PASS | 01_gradient_collection_test.py |
| 2 | Sketch B’s Hessian sketch (H_I, H_O), computed two independent ways, agrees to float32 precision; symmetric; flat-storage round-trips exactly | ✅ PASS | 02_sketch_b_synthetic_test.py |
| 3 | The real block-sequential rounding update (LDLQ_2hess,
not just the paper’s Eq. 5-6 in isolation), with a stand-in quantizer,
cross-verified MLX vs. NumPy |
✅ PASS | 03_rounding_update_test.py |
| 4a | Real Hessian collection (H_I, H_O) on one real layer via a REAL full-model forward+backward pass on this project’s actual hybrid architecture | ✅ PASS | 04_real_layer_test.py |
| 4b | Real block-Cholesky factors (Lin, Lout) from that real H_I/H_O | ✅ PASS | 04_real_layer_test.py |
| 4c | Real quantizer integration + real Hadamard rotation + naive-rounding comparison, full 1024/1024 real output rows — 3 real bugs/gaps found & fixed (diagonal-zeroing, Hessian batch-accumulation, missing Hadamard rotation); real result: 75% lower real trace-based error than naive | ✅ PASS | 04_real_layer_test.py,
hadamard_transform.py |
| 5a | Multi-tensor real Hessian collection (ALL real targets, ONE real forward+backward pass) | ✅ PASS (3-tensor subset) | 05_full_model_quantize.py |
| 5b | Real per-tensor correction (rotated where possible, unrotated fallback where not) + 2 real review-found bugs fixed | ✅ PASS (3-tensor subset) | 05_full_model_quantize.py |
| 5c | Real output model, --incoherence hadamard default
(YaqaQuantizedLinear, derived+verified) or
none (reuses 08_gptq_apply_plan.py’s
save/sidecar/MTP unchanged); real loader for the rotated case NOT
built |
✅ PASS (real production save path, exercised by every real run below) | 05_full_model_quantize.py,
yaqa_core.py |
| 5-full | The real 363-tensor run, end to end | ✅ PASS — built 2026-09-07, then rebuilt twice more
the same day after two real production bugs were found post-build
(missing mode; F32 scales/biases +
embed_tokens) — see the 2026-09-07 section below. Real
output-quality benchmark completed 2026-09-09/10 (see
CHANGELOG.md) — highest 5-test mean of the comparison
group, never worst on any single metric. |
05_full_model_quantize.py,
research_hadamard_blowup/exports/RCA_scales_f32_bug.html,
research_hadamard_blowup/IMPROVEMENT_LEDGER/09_FIRST_REAL_BENCHMARK_RESULTS.html |
Do not skip ahead. Each step gates the next. Step 2 exists specifically because a subtle bug in step 2’s math would be silent — no crash, no obvious symptom, just a wrong Hessian sketch that looks plausible. The whole point of this staged plan is refusing to chain steps on faith.
Interpreter: /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 (Python 3.11.15)
mlx: 0.32.2
mlx-lm: 0.31.3
mlx-optiq: 0.4.32
numpy: 2.4.6
Real source-of-truth locations, already on disk:
research/quantization_methods_2026/yaqa/code/ <- real Cornell-RelaxML/yaqa clone
research/quantization_methods_2026/yaqa/paper/YAQA_2505.22988.pdf <- real 20-page paper PDF
scripts/yaqa_port/ <- this port's own scripts (this ledger's directory)
Real file that everything in Step 1-2 is grounded in, read directly (not summarized):
research/quantization_methods_2026/yaqa/code/hessian_llama/custom_linear_B.py
Picture the AI model as a chain of stations, one after another, with data flowing through — in one side, out the other. Normally you can only watch data flow in ONE direction through a station: forward, left to right. But YAQA’s whole method depends on watching a SECOND signal too — a “how much did this affect the final answer” signal, which flows backward, right to left, only during a special kind of calculation. This script builds a tap on one real station in the chain and checks: can we actually read BOTH signals — the forward one AND the backward one — at the exact same point, at the same time? If we can’t, nothing else in this whole project is possible.
What Step 1 actually proves: a real tap on one real station, reading both the forward signal and the backward signal at the same point.
Where this idea came from (provenance). This isn’t
invented — it’s a direct translation of one specific real mechanism from
YAQA’s own original code, which was itself written for a completely
different toolkit (PyTorch/CUDA). The real file is
research/quantization_methods_2026/yaqa/code/hessian_llama/custom_linear_B.py,
class LinearNoBias. In PyTorch’s toolkit, this “tap” is
built with ctx.save_for_backward(input, weight) during the
forward pass, then read back inside
def backward(ctx, grad_output). Apple’s toolkit (MLX) has a
different mechanism for the exact same job, called
mx.custom_function — and Step 1’s entire purpose is
checking that this different mechanism actually gives us the same two
pieces of information PyTorch’s does, before trusting it for anything
real.
What’s literally inside
01_gradient_collection_test.py, in plain terms: 1.
Defines a tiny stand-in “station” — a single layer that just multiplies
numbers together, the same basic operation every real layer in the model
does. 2. Wires up the tap (mx.custom_function) on that
station. 3. Runs one fake calculation through it, forward then backward,
using the REAL dimensions of an actual layer from this project’s actual
model (not made-up numbers) —
language_model.model.layers.3.self_attn.q_proj,
12288 × 5120. 4. Checks: did the tap actually catch both
signals, with the right shapes and the right values? If yes, prints
STEP 1 PASS. If the shapes or values are wrong, or if MLX
itself refuses to run this at all, the script raises a real Python error
instead of silently continuing — there’s no “partial credit” here.
Real command — copy-paste this exactly to re-run it yourself:
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 01_gradient_collection_test.pyReal captured output (2026-09-03) — and what each line actually means:
loss value: 1613974.8750
grad wrt x shape: (4, 16, 5120) (expected (4, 16, 5120))
grad wrt w shape: (12288, 5120) (expected (12288, 5120))
--- what the custom vjp actually captured ---
captured input shape: (4, 16, 5120) (expected (4, 16, 5120))
captured grad_output shape: (4, 16, 12288) (expected (4, 16, 12288))
STEP 1 PASS: mx.custom_function's vjp gives simultaneous access to both
the real forward input and the real incoming cotangent, on a real layer
shape from this project's own model.
loss value — just a number confirming
the fake calculation actually ran to completion. Its exact value doesn’t
matter here; a Python error instead of a number would mean something
broke.grad wrt x shape /
grad wrt w shape — confirms the “backward”
direction produced numbers of the correct real shape — proof
the tap didn’t corrupt the normal flow of the calculation while it was
listening in.captured input shape /
captured grad_output shape — this is the actual
thing being tested: did our tap really see BOTH the forward signal
(input, shape (4, 16, 5120)) and the backward
signal (grad_output, shape (4, 16, 12288)) at
the same point? Both shapes matching “expected” means yes.STEP 1 PASS — the script’s own final
verdict, only printed if every check above it passed. If this
line is missing and you instead see a Python error/traceback, Step 1 has
failed — something about the tap mechanism doesn’t work as
assumed, and Step 2 should not be attempted until this is understood and
fixed.One real bug found and fixed while building this
(kept here so it’s not rediscovered): mx.custom_function‘s
.vjp callback receives cotangents matching the
output’s pytree structure, not the primals’.
linear_no_bias returns a single array, so
cotangents IS that array directly — not a 1-tuple. The
first version wrote (grad_output,) = cotangents and MLX
raised ValueError: too many values to unpack (expected 1).
Fix: grad_output = cotangents directly.
Step 1 proved we can read the “how much did this affect the answer” signal. Step 2 asks: what do we actually DO with that signal? YAQA’s core idea is to build a sensitivity map — a record of “if I nudge this number, which OTHER numbers does it interact with, and how strongly.” Nudge one number in a pond and the ripples spread outward in a specific pattern; this map is a record of that ripple pattern, computed from real data instead of guessed.
There’s a subtlety worth naming plainly: this one signal actually gets used to build two separate maps, not one — one describing ripples on the “numbers coming IN to this layer” side, one describing ripples on the “numbers going OUT of this layer” side. This step builds a TINY version of both maps — small enough to check by hand — and proves the math is correct using two completely different ways of computing it, before ever trusting it on a real, 27-billion-number model.
What Step 2 actually proves: the same real signal correctly builds two different sensitivity maps, verified two independent ways.
Where this idea came from (provenance). The real
formula lives in the same file as Step 1 —
custom_linear_B.py, inside
LinearNoBias.backward, the it == 0 branch
(lines 57-97) — expressed there as two dense torch.einsum
calls. Nothing here is guessed: the exact algebra was worked out by hand
(shown in full in the script’s own docstring) and reduces to a simple,
checkable recipe — build one small “how it mattered” matrix per example,
then combine those matrices two different ways.
Why Sketch B specifically, not Sketch A — checked
directly by diffing the real source (2026-09-03), not assumed from the
real README’s brief “we recommend Sketch B if you can afford it” line:
hessian_llama/ ships TWO real Hessian-collection files,
custom_linear_A.py and custom_linear_B.py.
Diffing them directly shows Sketch A’s cheap single-pass branch
(it==0) computes only H_I
(in_hess = input.T @ input — literally the same one-sided
form GPTQ’s own Hessian uses) — no H_O computation exists
in that branch at all. Sketch B’s single-pass branch computes
both H_I and H_O together,
via the real two-sided einsum shown above. This port’s whole algorithm
(LDLQ_2hess, Step 3) is two-sided by construction — it
needs both Lin and Lout. Sketch A’s efficient
path structurally cannot supply H_O at all; getting one out
of Sketch A would require its separate power-iteration refinement path
(it>0), which the real source itself flags as immature
(custom_linear_B.py’s own comment: “Additional power
iterations on B are not optimized and should be rewritten with einsums.
Use at your own risk!”). There was no real choice to make here — Sketch
A alone cannot feed the algorithm this project ported.
What’s literally inside
02_sketch_b_synthetic_test.py, in plain terms: 1.
Builds tiny, fake “input” and “how it mattered” numbers — small enough
that a human could, in principle, check the arithmetic by hand. 2.
Computes both sensitivity maps twice, using two
completely different pieces of code that share nothing: once using MLX’s
own einsum (mirroring the real formula’s exact notation),
and once using an ordinary loop with plain matrix multiplication (worked
out independently, from first principles). 3. Compares the two results.
If a bug slipped into either version, the two numbers would very likely
disagree — that disagreement is the actual safety net here, not just
“the code didn’t crash.” 4. Also checks a structural fact both real maps
must satisfy (they must be symmetric — the map from A to B must
equal the map from B to A), as a second, independent sanity check.
The real formula, derived by hand from the actual
source (not the paper’s simplified notation) —
custom_linear_B.py, LinearNoBias.backward, the
it == 0 branch, lines 57-97:
in_hess = einsum('btm,btn,bsm,bsk->nk', grad_output, input, grad_output, input)
out_hess = einsum('btm,btn,bsk,bsn->mk', grad_output, input, grad_output, input)Regrouped algebraically, using the same per-example quantity the real
code itself names elsewhere in this file
(grad_weight = grad_output[i].T @ input[i], line
122-123):
G[b] = grad_output[b].T @ input[b] (out_features, in_features)
in_hess = sum_b G[b].T @ G[b] (in_features, in_features)
out_hess = sum_b G[b] @ G[b].T (out_features, out_features)
Full worked derivation for in_hess is in the script’s
own docstring — verified by hand, not just asserted.
Real command:
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 02_sketch_b_synthetic_test.pyReal captured output (2026-09-03), abbreviated (full matrices in the script’s own run log):
max |diff| (H_I, einsum vs. manual loop): 7.629e-06
max |diff| (H_O, einsum vs. manual loop): 7.629e-06
H_I symmetry error: 0.000e+00
H_O symmetry error: 0.000e+00
sym_to_flat/flat_to_sym round-trip error: 0.000e+00 (10 values for a 4x4 symmetric matrix, vs 16 dense)
STEP 2 PASS: Sketch B's Hessian sketch (H_I, H_O) computed two independent
ways ... agree to float32 precision. Both are symmetric as required. The
real symmetric-storage mechanism round-trips exactly.
What each line actually means: -
max |diff| (H_I/H_O, einsum vs. manual loop)
— this is the actual test. Two completely different pieces of code
computed the same map; this number is the biggest disagreement found
between them, anywhere in the map. 7.629e-06 means
“0.000007-ish” — a difference so small it’s just normal computer
rounding noise, not a real disagreement. If this number had come back
large (say, bigger than 0.01), that would mean the two independent
implementations genuinely disagree — a real bug somewhere, full stop. -
H_I/H_O symmetry error — a second,
unrelated check: does the map satisfy a structural rule real sensitivity
maps must satisfy (symmetric — “A affects B” must equal “B affects A”)?
0.000e+00 means exactly zero violation — this rule holds
perfectly. -
sym_to_flat/flat_to_sym round-trip error —
this project’s own storage trick (saving only half the map, since it’s
symmetric, to use less memory) was checked by compressing then
decompressing it and comparing to the original. 0.000e+00
means nothing was lost. - STEP 2 PASS —
again, only printed if every check above it actually passed. No PASS
line (a Python error instead) means stop and investigate before moving
to Step 3.
Toy dimensions used:
batch=3, seq=5, in_features=4, out_features=3 — small on
purpose, chosen to be hand-tractable, not representative of a real
layer’s scale.
Real verification method (why this counts as
evidence, not just “it ran”): H_I/H_O computed via
mx.einsum (subscripts copied verbatim from the real
torch.einsum calls) AND via a completely separate
plain-matmul per-example Python loop (the hand-derived formula above,
containing no einsum at all). Two structurally different code paths
agreeing to float32 tolerance is real cross-validation.
Now we have the sensitivity maps from Step 2. Step 3 asks: how do we actually USE them to round the model’s numbers more carefully? The real answer turns out to be more specific than “apply one formula” — the weight matrix (think of it as a big spreadsheet grid) gets rounded one small block at a time, in a very particular crisscross order, and each block’s rounding is nudged using the blocks already finished next to it — the same “compensate the neighbors” idea GPTQ uses, but now pulling from two directions at once (using both sensitivity maps from Step 2), not just one.
What Step 3 actually proves: this crisscross, pull-from-already-done-neighbors correction order, computed correctly.
Real finding that changed this step’s scope (2026-09-03): the plan originally assumed Step 3 was “implement Eq. 5-6, a single fixed-point formula.” Reading the real source before writing code found something more specific and more involved — the actual update is a block-sequential 2D scheme, not a closed form:
research/quantization_methods_2026/yaqa/code/lib/algo/ldlq.py, function LDLQ_2hess() (lines 16-59)
It processes the weight matrix in (td_x × td_y) blocks,
walking them in a specific anti-diagonal order (the real
starts list), and at each block combines three cross-terms
built from the already-rounded neighboring blocks
(W - hatW) and the block-Cholesky factors
Lout/Lin — the real code’s own names for the
paper’s L_O/L_I, computed by
lib/utils/math_utils.py:block_LDL from our own validated
H_I/H_O (Step 2). Only at the very end of each
block does the real code call cb.quantize(...) — QTIP’s own
trellis/lattice codebook (lib/codebook/bitshift.py), a
completely different quantization format from this project’s own affine,
group-based one.
Real, checked-against-the-repo finding on a related
question: an external summary (pasted mid-session, phrased like
an AI search overview) claimed the real yaqa repo has a ready-made
--quantizer uniform flag as an alternative to QTIP’s
codebook — i.e., that no custom substitution work would be needed.
Checked directly and this is false: grep’d
the entire real repo for
quantizer/uniform/--codebook
choices — no such flag exists anywhere. The only --codebook
argument (quantize_llama/quantize_finetune_llama.py:34) is
a free-text string resolved against lib/codebook/, and the
only codebook module actually present in the repo is
bitshift.py (genuine trellis/lattice decode functions —
decode_1mad, decode_2mad,
bitshift_codebook — no uniform/affine alternative). The
directional idea (the correction math doesn’t care what
quantizer consumes it) is correct and is exactly what this step’s own
substitution is built on — but it required writing real substitution
code, not flipping an existing flag. Flagging this here so a future
session doesn’t waste time looking for a flag that isn’t there.
What this script does: keeps the real correction
logic (the three-cross-term formula, the exact anti-diagonal block
traversal) faithful to ldlq.py, and replaces ONLY
cb.quantize(...) with a trivial round-to-nearest-grid
stand-in (simple_quantize) — representing where this
project’s OWN eventual quantizer
(mx.quantize/gptq_pack, already built and
validated for the GPTQ integration) will plug in at Step 4, not an
attempt to adopt QTIP’s lattice codebook.
What’s literally inside
03_rounding_update_test.py, in plain terms: 1.
Builds two tiny, fake sensitivity maps (standing in for real ones from
Step 2). 2. Runs the real crisscross block-correction order (the diagram
above) on a tiny, fake weight grid — twice, using two completely
separate pieces of code (an MLX version and a plain NumPy version) that
share no code between them. 3. Compares the two results for an exact
match, and separately checks a structural fact the sensitivity-map math
must satisfy (that certain blocks inside it must come out as an exact
identity — a mathematical fingerprint that’s either exactly right or
reveals a real bug, no middle ground).
Real command:
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 03_rounding_update_test.pyReal captured output (2026-09-03), abbreviated — and what it means:
Lin diagonal block 0 identity error: 0.000e+00
Lin diagonal block 1 identity error: 0.000e+00
max |diff| between independent implementations: 0.000e+00
max deviation from quantization grid: 0.000e+00
STEP 3 PASS: the real block-sequential LDLQ_2hess correction algorithm ...
ported to MLX and cross-verified against an independent NumPy implementation
-- exact agreement, and every output entry actually lands on the quantization
grid. block_LDL's diagonal-identity property also verified directly.
Lin diagonal block N identity error —
the structural fingerprint check mentioned above. 0.000e+00
means it landed exactly right.max |diff| between independent implementations
— the MLX version and the NumPy version, compared block by block, entry
by entry. 0.000e+00 means they produced the literal,
bit-for-bit identical answer — as strong a result as this kind of test
can give, since (unlike Step 2’s tiny floating-point noise) this
arithmetic has no randomness in it at all once the inputs are
fixed.max deviation from quantization grid —
every rounded number should land exactly on an allowed “grid point”
(like landing exactly on $20.00, never $20.03). 0.000e+00
confirms none of them missed.STEP 3 PASS — same rule as every step:
only printed if all checks above it passed. A Python error instead means
stop before Step 4.Real verification method: two independent
implementations — an MLX port and a plain NumPy port, sharing no code —
run on identical tiny random inputs (4×4 weight matrix,
2×2 blocks, the smallest case with a real >1 block grid
so the cross-term logic is actually exercised). Exact
0.000e+00 agreement, since this is deterministic
block-sequential arithmetic, not a stochastic estimate — exact agreement
is the right bar here, and it was met.
Not yet done: real quantizer integration (this still
uses the trivial round-to-grid stand-in, not
mx.quantize/gptq_pack) — that’s part of Step
4, along with running against a real model layer for the first time.
Steps 1-3 proved every piece works on small, fake numbers. Step 4 asks the real question: does any of this survive contact with the actual, real, 27-billion-number model? This is where a genuinely important, real surprise showed up — not a bug in our own code, but a real fact about how this specific AI model is built, that would have quietly blocked everything if we hadn’t found it.
What Step 4 actually proves: the real fix that turns a genuinely blocked measurement into a working one, on the real model.
Target layer, deliberately not q_proj:
language_model.model.layers.3.self_attn.k_proj, weight
shape (1024, 5120) — chosen instead of q_proj
(12288, 5120) because H_O for q_proj would be a dense
12288×12288 matrix; for k_proj it’s 1024×1024, a properly gated first
real-scale test instead of jumping straight to this model’s largest
attention projection.
What’s literally inside 04_real_layer_test.py,
in plain terms: 1. Loads the actual, real 27-billion-number
model — not a stand-in. 2. Puts a real “tap” (Step 1’s mechanism) on one
real layer inside it. 3. Switches the model into the “two-way street”
mode explained in the diagram above (model.train()). 4.
Feeds a handful of real calibration text windows through the
entire real model, computes a real next-word-prediction loss,
and runs that loss backward through the whole thing — which is what
actually fires the tap and builds the two real sensitivity maps (H_I,
H_O) for our one target layer, using the exact math validated in Step 2.
5. Builds the real block-Cholesky factors from those real maps (Step 3’s
math), on real-world-sized numbers for the first time — not tiny toy
ones anymore.
Real command:
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 04_real_layer_test.pyReal captured output (2026-09-03) — and what it means:
Calibration batch: (6, 128)
Running a real forward+backward pass through the WHOLE model...
loss: 2.1989
H_I shape: (5120, 5120), H_O shape: (1024, 1024), batches accumulated: 1
Part 1 PASS: real Hessian collection works end-to-end on a real layer, inside a real
full-model forward+backward pass, with no NaN.
Computing block-Cholesky factors (Lin, Lout) from real H_I/H_O...
Lin shape: (5120, 5120), Lout shape: (1024, 1024)
Part 2 PASS: real block-Cholesky factors computed without error.
Calibration batch: (6, 128) — 6 short
real text samples, 128 words each, deliberately small to keep this first
real-scale run fast.loss: 2.1989 — a real, sane number
(translates to roughly “9 plausible next words on average” — normal for
a real model on real text). This is a soft sanity check: if this number
came back as NaN or something wildly huge, that alone would
signal something is broken, even before checking anything else.H_I shape: (5120, 5120), H_O shape: (1024, 1024)
— confirms the two real sensitivity maps came out at exactly the real
dimensions this real layer requires — not placeholder or mismatched
sizes.Part 1 PASS / Part 2 PASS
— same rule as every step: only printed if the checks above them
actually passed.The loss (2.1989, ≈9 perplexity) is a sane, real cross-entropy value for a real model on real calibration text — worth noting as a soft sanity check, not just “it didn’t crash.”
Two real bugs found and fixed while building this (kept here so they’re not rediscovered):
mx.value_and_grad(loss_fn) differentiated
w.r.t. the wrong thing. By default it computes gradient w.r.t.
loss_fn’s first positional argument — here that was
batch, the real token IDs, which are
discrete/non-differentiable (used for embedding-table gather). MLX
raised this directly:
ValueError: [gather_axis] Cannot calculate VJP with respect to indices. Use stop_gradient on indices...
Fix: use nn.value_and_grad(model, loss_fn)
instead — differentiates w.r.t. the model’s trainable parameters, which
is what actually needs to flow backward through the wrapped layer; the
token IDs were never meant to receive a gradient.
Major, real architectural finding: this model
(mlx_lm.models.qwen3_5) is a hybrid
architecture —
DecoderLayer.is_linear = (layer_idx+1) % full_attention_interval != 0
(full_attention_interval=4), so 3 of every 4 layers
use GatedDeltaNet (“linear_attn”) instead of standard
attention, not just the target layer’s own type.
GatedDeltaNet’s recurrence
(mlx_lm/models/gated_delta.py) runs through a raw
mx.fast.metal_kernel (a hand-written Metal shader) with
no backward rule defined at all. Backward from the
final loss to ANY layer — including our target, layer 3,
which is itself is_linear=False — has to pass through
several GatedDeltaNet layers along the way, so this blocked
the whole pass, not just linear-attention layers specifically. MLX
raised this directly:
ValueError: [Primitive::vjp] Not implemented for CustomKernel.
Real fix, already built into mlx_lm itself, no patching
needed:
gated_delta_update(..., use_kernel=not self.training)
(gated_delta.py:281-282, called from
qwen3_5.py:193) — calling
model.train() switches every
GatedDeltaNet layer onto its own pure-MLX-ops reference
implementation (gated_delta_ops(), genuinely
differentiable), automatically. This is a durable, real fact worth
remembering for any future gradient-based work on this model family:
model.train() is required before any real backward
pass, not just a training-loop nicety.
Minor: mx.linalg.cholesky isn’t
supported directly on the GPU stream —
ValueError: [linalg::cholesky] This op is not yet supported on the GPU.
Same fix already established in the GPTQ integration
(08_gptq_apply_plan.py’s
compute_inverse_hessian): wrap in
with mx.stream(mx.cpu):.
Revision history, honestly: this section was rewritten twice. The first version reported a 32-row partial result with a stand-alone bug fix (diagonal-zeroing) and concluded YAQA had “not yet been shown to beat naive.” An independent code review (2026-09-03) then found two further real problems in that version — a silent formula bug in the Hessian collection, and an entire real preprocessing stage (Hadamard incoherence processing) missing outright, not just simplified. Both are now fixed, independently verified against the real running algorithm (not just its source), and re-run in full. This is the current, complete state — not an addendum to the old one.
What this actually does, in plain English. Steps 1-3
proved the ordering of the smart-rounding correction on toy
numbers, using a fake “round to the nearest step” stand-in for the
actual rounding operation. Part 4c swaps that stand-in for the real
rounding operation — MLX’s own mx.quantize, the exact same
real quantizer this project’s GPTQ path already ships — runs it block by
block on the real weight matrix, and compares the result against just
rounding the same real weight the plain way, no correction at all. Two
things turned out to be missing before this comparison could be trusted:
the real Hessian-collection formula had a bug that only shows up with
more than one calibration sequence, and a whole real preprocessing stage
(a random rotation that makes the weight distribution friendlier to
low-bit rounding) was never ported in the first place. Both are fixed
below, and the real result changes completely once they are: with the
complete real pipeline actually running, YAQA’s correction reduces the
real error metric by 75% versus naive rounding.
Why the earlier version missed this: the earlier
32-row test used mx.isnan as its only correctness check and
a plain Frobenius-norm comparison. Neither check would ever have caught
a formula bug that only manifests at batch > 1, or a
missing preprocessing stage that the real algorithm’s own accuracy
depends on — those required someone reading the real source again, line
by line, specifically hunting for divergence, not just checking that a
given run’s numbers stayed finite.
Real data from the final canonical run, 2026-09-03 (see the annotated output below for the exact numbers and the run that produced them).
Design decisions kept from the plan:
td_x=1 (one output row per block),
td_y=group_size (64), so LDLQ_2hess’s own
blocks line up exactly with mx.quantize’s real
per-row-per-group structure, rather than inventing a new grouping
convention. The repo’s own defaults (td_x=16, td_y=16,
quantize_finetune_llama.py:47-48) target QTIP’s own lattice
codebook, which groups differently — not a fit for this project’s own
affine quantization format.
On row coverage (this changed from the earlier
version): the earlier 32-row partial test relied on a real,
still-correct argument — block_LDL is a sequential
leading-principal-minor factorization, so a leading row/column slice of
its output is exact. That argument stops being sufficient once Hadamard
rotation is added: the rotation is a global transform that
mixes every one of the 1024 real output rows together, so a partial
slice of the rotated result can no longer be validly inverse-rotated
back to the original weight basis, and the real trace-based error metric
is likewise only well-defined over the full output dimension. This run
therefore covers all 1024 real output rows of the real
k_proj layer (81,920 real blocks total), not a slice.
What’s literally inside 04_real_layer_test.py’s
Part 2-4, in plain terms: 1. Generates two real random ±1 “sign
vectors” (SU for the 5120 real input features,
SV for the 1024 real output features) and applies the real
Hadamard incoherence-processing rotation to H_I,
H_O, and the real weight matrix — the exact same rotation
the real pipeline always applies before doing anything else, previously
missing from this port entirely (see bug #2 below). 2. Builds the real
block-Cholesky factors (Lin, Lout) from the
rotated H_I/ H_O, same damping-retry
and diagonal-zeroing logic as before, now applied to the correct
(rotated) inputs. 3. Walks the block grid, all 1024 real output rows, in
the real anti-diagonal order — for each block, computes the real
three-cross-term correction on the rotated weight, then calls
MLX’s own real mx.quantize/mx.dequantize on
the corrected result — “compute the correction, then quantize the
corrected result.” 4. Computes two real comparisons: the real pipeline’s
own trace-based, Hessian- weighted error metric (rotated basis, against
a naive-but-also-rotated baseline), and a plain Frobenius error in the
original weight basis (against naive mx.quantize on the
real, unrotated weight) for continuity with the earlier version of this
test.
Real command:
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 04_real_layer_test.pyReal captured output (2026-09-03, final canonical run, fully fresh — model reloaded, real backward pass rerun, no cache reused) — and what it means:
Applying real Hadamard incoherence-processing rotation...
Hin_rot shape: (5120, 5120), Hout_rot shape: (1024, 1024)
Wr (rotated weight) shape: (1024, 5120)
Computing block-Cholesky factors (Lin, Lout) from the real, rotated H_I/H_O...
Hin: block_LDL succeeded on attempt 1, final sigma_reg=1.0000
Hout: block_LDL succeeded on attempt 1, final sigma_reg=1.0000
Lin max|entry|: 8.9745e-02 Lout max|entry|: 1.5886e+00
Part 2 PASS: real Hadamard rotation + real block-Cholesky factors computed without error.
Part 3: real LDLQ_2hess correction (rotated basis) on 1024 of 1024 real output rows
(all 5120 real input columns), td_x=1, td_y=64 -- 81920 blocks, real mx.quantize
(bits=4, group_size=64) at every block.
sweep 1/1104: max|hatWr| so far = 4.1420e-02
sweep 1100/1104: max|hatWr| so far = 8.2684e-02
Part 3: 100.0% of hatWr entries are finite.
Part 3 PASS: real LDLQ_2hess correction ran end-to-end (rotated basis) on a real
layer with a real quantizer at every block, no NaN.
Part 4: real comparison against naive rounding, two metrics...
[Metric A -- real trace-based error, rotated basis, finetune.py's own formula]
YAQA-corrected: err=1.749097e-03 Naive (rotated, no correction): err=7.077018e-03
YAQA correction reduces the real trace-based error by 75.28% vs. naive rounding.
[Metric B -- plain Frobenius reconstruction error, original (unrotated) weight basis]
YAQA-corrected: Frobenius=3.370495 max|diff|=0.007998
Naive mx.quantize (unrotated, on real W directly): Frobenius=3.314037 max|diff|=0.007405
YAQA correction did NOT reduce Frobenius error vs. plain naive rounding.
Lin max|entry|: 8.97e-02,
Lout max|entry|: 1.59e+00 — small, sane numbers,
smaller even than the pre-rotation version’s already-sane values —
consistent with Hadamard rotation doing its real job of spreading out
concentrated structure.sweep 1 ... sweep 1100: max|hatWr| —
stays small and flat across all 1104 diagonal sweeps (81,920 blocks) —
numerically stable at full scale, not just on a 32-row sample.err=1.75e-3 (YAQA) vs err=7.08e-3 (naive) — a
75.28% reduction in the exact error metric the real
algorithm is designed to minimize (weighted by real per-weight
sensitivity, via the real Hessian maps), computed with identical
rotation preprocessing on both sides so the comparison isolates what the
correction itself buys.Frobenius=3.370 (YAQA) vs
3.314 (naive) — YAQA is about 1.7% worse on raw,
unweighted magnitude error. Reported honestly, not hidden — and
explained in the diagram above: this is the expected signature of an
algorithm that trades a small amount of raw magnitude accuracy for a
large reduction in the metric that actually reflects real model-output
impact, not evidence of a problem.H_I/H_O from a
prior model load — real evidence the result is reproducible, not a
one-off.Real bugs and gaps found and fixed (kept here so they’re not rediscovered):
block_ldl_mlx
(validated in Step 3) correctly sets each diagonal block of its
output to the identity matrix, matching the real block_LDL
function (math_utils.py:14-42) exactly. But the real
caller of block_LDL
(lib/algo/finetune.py:172,183, read directly from the real
source) does one more step: it zeroes the literal
diagonal afterward — Lin[arange(n),arange(n)]=0,
Lout[arange(m),arange(m)]=0 — turning those identity
diagonal blocks into zero blocks before
Lin/Lout are ever used in
LDLQ_2hess. Without this, the correction formula’s
“self-term” spuriously re-injects a near-copy of each block back into
the running correction instead of correctly excluding self-contribution
— on the original (pre-rotation) real layer, this compounded
block-to-block at roughly 1.5× per diagonal sweep, overflowing to
infinity before the loop finished. Fix:
zero the literal diagonal of Lin/Lout after
block_ldl_mlx, exactly matching the real caller code.batch=1, would corrupt
results at batch>1.
yaqa_linear_vjp originally flattened batch and sequence
together into one axis before forming a single pooled
G = grad_output_flat.T @ input_flat, then squared that once
(H_I += G.T@G). The real formula
(hessian_llama/custom_linear_B.py, independently
hand-verified with a real batch of 3 in Step 2) requires a
per-example
G[b] = grad_output[b].T @ input[b] for every sequence
b, then H_I += sum_b(G[b].T@G[b]) — summing
the squared per-example matrices, not squaring the
sum. These differ the moment batch>1:
(sum_b G_b).T@(sum_b G_b) expands to include spurious
cross-sequence terms the real formula never forms. Every real run so far
used data[:1] (batch=1, forced by an earlier OOM), where
the two formulas coincide by construction — completely invisible in
every number this script had produced. Fix: loop over
the real batch axis explicitly when accumulating
H_I/H_O, matching Step 2’s own validated
per-example loop exactly. This did not change any number in this run
(still batch=1), but is now correct for any future run with real
multi-sequence calibration — which is exactly the direction Step 5
needs.finetune.py:145-152,167,178,192-197) always rotates
H_I, H_O, and the real weight through a
random-sign Hadamard transform before
block_LDL/LDLQ_2hess ever runs. This is real
“incoherence processing” from the QuIP#/QTIP lineage this algorithm
descends from — it exists specifically to make weight distributions more
favorable to accurate low-bit rounding, and its absence was not
previously documented as a known limitation; it was simply missing.
Fix: ported the real transform
(hadamard_transform.py, from
lib/utils/matmul_had.py) and wired it into the real order
(normalize → rotate → damp → block_LDL, weight rotation
before the block loop, inverse rotation after). Independently
verified two ways before use, not just read: (a) all 12 of the
real get_hadXX() factor tables were extracted by
mechanically parsing the real source’s literal arrays via AST (not
retyped by hand), and every one independently verified to satisfy
H @ H.T == n·I with every entry exactly ±1
(max error 0.0 across all 12 tables); (b) the real
matmul_hadU/matmul_hadUt functions were
extracted from the real source and actually run in an
isolated torch-CPU environment (no
CUDA/fast_hadamard_transform needed — those are only used
by the separate matmul_hadU_cuda kernel path this port
doesn’t need), and this MLX port’s output matched the real running
output exactly for n=1024
(K=1 path, max diff 0.0) and to float32
rounding precision for n=5120 (K=20 path via
get_had20, max diff 9.5e-7 — the same order as
the real code’s own internal round-trip self-consistency error). This is
the strongest verification standard used anywhere in this port so far:
not just cross-checking two independent implementations of a
formula, but running the real, actual code and
comparing outputs directly.Wscale step
(finetune.py:194-197) rescales the rotated weight’s RMS
magnitude to match the trellis codebook’s fixed lookup-table range. This
port uses mx.quantize’s own affine per-group quantizer
instead of a trellis codebook, which derives its scale/bias fresh from
each block’s own data range every time — a global rescale of its input
is mathematically a no-op for its reconstruction accuracy (scale by
k, quantize, divide back by k reconstructs
identically). Skipped deliberately, not silently.starts block-traversal list
re-processes one diagonal twice (present identically in the real
upstream ldlq.py:23-26, traced through and confirmed
harmless — every reference to a block written on the first pass hits a
position the Lin/Lout diagonal-zeroing already zeroed, so the repeat
pass changes nothing). And HESSIAN_CACHE
(.cache_k_proj_hessians.safetensors, ~114MB) is written but
never auto-deleted — fine at single-layer scale, a real, disclosed
consideration for Step 5’s own design if the same per-script caching
pattern gets reused across many layers.A real engineering aside worth keeping: an earlier
full-6-sequence-batch attempt to raise the real Hessian’s rank (instead
of relying on damping) produced no output at all and exit code
0 from a python3 | tee pipeline — consistent
with the OS OOM-killing python3 mid-run while
tee, on the other end of the pipe, still exits
0 on EOF regardless of why the writer died. Bash
does not propagate an upstream command’s exit code through a pipe
without set -o pipefail — worth remembering for any future
real run in this project that pipes to tee.
What this step actually needs, decided against real project artifacts, not invented:
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara — not
the OptiQ or HybridPareto checkpoints. YAQA, like GPTQ, needs
full-precision weights to compute a correction against; feeding it an
already-quantized model would compound error and make any comparison
meaningless.04_LANGUAGE_PLANS/V3_3_APPLES_TO_APPLES_0005/plan_Qwen38_V3_3_5.028075122774708_BPW.json
— the same real plan the HybridPareto5bpw-V3.3_Stratified
checkpoint was built from. 497 real tensor assignments; 134 are
bits:16 (the KEEP_BF16_BITS sentinel — skip,
matching 08_gptq_apply_plan.py’s own established
convention), leaving 363 real tensors that need real
quantization.stratified_calibration.py module
08_gptq_apply_plan.py already uses (confirmed by reading
both scripts directly) — same methodology, but NOT yet the same scale:
this step’s real runs so far still use the minimal
k_per_domain=1, seq_len=128 from Step 4’s own feasibility
test, not a real production calibration size.optiq.sidecars.attach_sidecars,
optiq.runtime.mtp_convert) is planned to reuse
08_gptq_apply_plan.py’s own already-trusted mechanism
unchanged — not yet wired in (Part 5c, below).yaqa_core.py — Step 4’s validated math (Hadamard
rotation, regularized_block_ldl,
ldlq2hess_quantize, the real trace-based error metric),
extracted into a shared module so Step 5 reuses the exact same proven
code, not a duplicate. Verified after extraction: re-ran
04_real_layer_test.py against its cache and got
bit-identical output to the pre-refactor canonical run
(same err_yaqa=1.749097e-03,
err_naive=7.077018e-03, 75.28%, same Frobenius
numbers to the last digit); full regression suite
(test_regression.py) re-run, all 40+ checks still
pass.05_full_model_quantize.py — the new Step 5 script.What this proves. Wrapping ONE tensor at a time
(Step 4’s approach) would mean 498 separate full-model backward passes
for the full plan — infeasible. Part 5a wraps ALL real target tensors
SIMULTANEOUSLY and collects every one’s
H_I/H_O in ONE real forward+backward pass,
mirroring the real, already-proven mlx_lm GPTQ
Catcher pattern (mlx_lm/quant/gptq.py:40-76,
read directly 2026-09-03: wraps every real
nn.Linear/SwitchLinear at once, one
calibration pass, mx.eval(layers) per batch). This machine
has 128GB RAM (checked directly), comfortably covering even the largest
real Hessian in this plan (H_I for a 17408-in_features MLP
tensor is 17408×17408 float32, ~1.2GB).
Real command:
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
"$PYBIN" 05_full_model_quantize.py --subset defaultReal captured output (2026-09-03) — 3 real tensors,
one of each real architectural type in this hybrid model, each confirmed
to actually need real quantization in the real plan (not
bits:16):
3 real tensors need real quantization (bits != 16) -- SUBSET:
['language_model.model.layers.62.linear_attn.out_proj',
'language_model.model.layers.63.mlp.up_proj',
'language_model.model.layers.63.self_attn.q_proj']
Running ONE real forward+backward pass through the WHOLE model...
loss: 2.1989
3/3 real target tensors' custom vjp fired during the real backward pass.
language_model.model.layers.63.mlp.up_proj: H_I shape=(5120, 5120), H_O shape=(17408, 17408), count=1, NaN in H_I=False, NaN in H_O=False
language_model.model.layers.63.self_attn.q_proj: H_I shape=(5120, 5120), H_O shape=(12288, 12288), count=1, NaN in H_I=False, NaN in H_O=False
language_model.model.layers.62.linear_attn.out_proj: H_I shape=(6144, 6144), H_O shape=(5120, 5120), count=1, NaN in H_I=False, NaN in H_O=False
Part 5a PASS: real Hessian collection works end-to-end for all 3 real target tensors
simultaneously, in ONE real forward+backward pass, no NaN.
This is the first real confirmation that the gradient-capture
mechanism works on a linear_attn/GatedDeltaNet tensor
specifically, not just standard self_attn — a real,
previously-untested case for this hybrid architecture.
Real command (single-tensor test):
"$PYBIN" 05_full_model_quantize.py --subset "language_model.model.layers.62.linear_attn.out_proj"First real run — td_x=1 (matches Step 4’s own
default):
Metric A (real trace-based, rotated): yaqa=2.903688e-05 naive=9.099063e-05 (REDUCES error by 68.09%)
Metric B (Frobenius, original basis): yaqa=8.8815 naive=4.0496
Real wall-clock: 12m 25s (20:43:50 → 20:56:15 BST),
for one tensor, 491,520 real blocks
(out_features × in_features/group_size = 5120 × 6144/64).
Real finding: the full plan is infeasible at this
rate. Total real blocks across all 363 real tensors, computed
directly from the plan’s own param_count/
group_size: 385,064,960. At the measured
per-block rate, that extrapolates to roughly 130 real hours (5+
days) of continuous compute —
language_model.lm_head alone (19.87M blocks, 6-bit) would
take ~6.7 hours by itself. This is a real, hard blocker, not a
nice-to-have optimization.
Real fix: generalize td_x.
yaqa_core.ldlq2hess_quantize previously hardcoded
td_x=1; generalized to accept td_x as a
parameter. This is NOT a new algorithm — Step 3 already validated this
exact formula structure with td_x=td_y=2 (not just 1)
against an independent NumPy reference
(03_rounding_update_test.py, M,N,TD=4,4,2),
and the real code’s own default is td_x=16
(quantize_finetune_llama.py:47, read directly).
mx.quantize’s own real docstring confirms grouping is
per-ROW
("every group_size elements in a row of w are quantized together"),
so a (td_x, group_size) block quantizes identically to
td_x separate (1, group_size) calls —
increasing td_x reduces Python-loop iterations without
changing what mx.quantize computes.
Real re-run, td_x=16, same tensor:
Metric A (real trace-based, rotated): yaqa=2.909955e-05 naive=9.099063e-05 (REDUCES error by 68.02%)
Metric B (Frobenius, original basis): yaqa=8.9201 naive=4.0496
Real wall-clock: 67 seconds (21:00:17 → 21:01:24
BST) — an ~11× real speedup. The naive baseline is
EXACTLY unchanged (9.099063e-05 both runs, as it must be,
since naive doesn’t depend on td_x), and Metric A changed
by only 0.07 percentage points (68.09% → 68.02%) — strong real evidence
the td_x generalization is correct, not just “didn’t
crash.”
A proper independent cross-check was then built and
run (test_regression.py, test 4b), matching Step
3’s own standard rather than resting on the plausibility check above.
First attempt: hand-wrote a NumPy replica of mx.quantize’s
documented affine formula (s=(alpha-beta)/(2^b-1)) and
compared against yaqa_core.ldlq2hess_quantize — failed,
max|diff|=0.399. Root-caused directly (not guessed): the
hand-written replica’s predicted scale (0.1420) did not
match the REAL scale mx.quantize itself returns
(0.1396) for bits=5, on a real standalone
block, checked with no sequential loop involved at all — the real
implementation does something beyond its own docstring’s simplified
formula, and reverse-engineering the exact discrepancy would have been
guessing, not verification. Real fix: rebuilt the
independent check to call REAL
mx.quantize/mx.dequantize from within the
NumPy-orchestrated loop (converting thing to an
mx.array and back each step) instead of replicating Apple’s
primitive by hand — this isolates exactly what needed checking (the
generalized loop’s indexing/slicing/formula correctness), using the same
real, already-trusted quantizer on both sides. Real result:
max|diff|=5.96e-07 (float32 machine epsilon) — the
td_x generalization is now independently verified, not just
plausible.
Honest full-plan extrapolation from this real data point (two bounds, not a guarantee): conservative (no fixed-cost amortization) ≈ 14.6 hours; realistic (the one-time model-load+backward cost is paid ONCE across all 363 tensors, matching Part 5a’s real architecture, not once per tensor) ≈ 3.5-4 hours.
Running the 3-tensor subset with td_x=16 crashed on
mlp.up_proj: get_hadK(17408) has no valid
K-factor (17408 = 2^10 × 17, and 17 isn’t in
the real algorithm’s divisor table
{172,156,140,124,116,108,60,52,36,28,20,12}, nor is 17408
itself a power of 2). This is a real limitation of the real upstream
algorithm itself — matmul_had.py’s own else-branch has the
identical assertion — not a porting bug. Checked directly across all 363
real target tensors in the real plan (shapes read from the real loaded
model, not the plan file, which has no shape info): 190/363
(52%) have at least one dimension with no valid Hadamard factor
— every real
mlp.up_proj/down_proj/gate_proj
(intermediate size 17408) across all 64 layers, plus
lm_head (vocab size 248320).
Real fix, not an invented mechanism:
hadamard_transform.has_hadamard_support(n) (a non-throwing
predicate reusing get_hadK’s exact real dispatch order)
gates whether a given tensor gets real Hadamard rotation. When either
dimension lacks support, 05_full_model_quantize.py falls
back to plain (unrotated) LDLQ_2hess correction —
yaqa_core.normalize_H (just the real mean-normalize step,
no rotation) in place of rotate_symmetric_H, and the real,
unrotated weight directly in place of rotate_weight’s
output. This is exactly the code path that existed and was validated
before the Hadamard fix was added — not new, unproven logic.
Real re-run, all 3 subset tensors, with the fallback (2026-09-03, 21:36:34 → 21:49:57 BST, 13m23s total for all 3 including the one-time model load+backward):
mlp.up_proj (UNROTATED fallback): yaqa=5.31e-07 naive=2.79e-06 REDUCES error by 80.98%
self_attn.q_proj (rotated): yaqa=1.28e-06 naive=3.36e-06 REDUCES error by 61.94%
linear_attn.out_proj (rotated): yaqa=2.91e-05 naive=9.10e-05 REDUCES error by 68.02%
Part 5b: 3/3 real tensors corrected successfully, 0 failed (1/3 ran without rotation).
3/3 tensors: YAQA reduces the real trace-based error vs. naive.
All 3 real architectural types in this hybrid model
(self_attn, linear_attn/ GatedDeltaNet,
mlp) now confirmed working, including the unrotated
fallback path — and the unrotated tensor shows the LARGEST real error
reduction of the three (80.98%), a real, honest result, not assumed or
hoped for.
Reviewed yaqa_core.py,
hadamard_transform.py, and
05_full_model_quantize.py in full against the real
ground-truth source (finetune.py, ldlq.py,
math_utils.py, matmul_had.py,
custom_linear_B.py), requested explicitly given how many
real bugs this session had already found.
Real bug found (HIGH severity):
mx.random.seed(0) was called INSIDE the per-tensor
loop, resetting MLX’s global RNG to the same fixed value before
drawing SU/SV for every tensor. Since the RNG
is deterministic given a seed, every tensor sharing the same
in_features got a bit-identical SU, and every
tensor sharing the same out_features got a bit-identical
SV — e.g. every real self_attn.q_proj across
all 64 real layers (same shape) would have received the exact same
random-sign rotation. Harmless in Step 4 (only one tensor per run — this
line was carried over from there unexamined) but silently defeats the
whole real point of independent per-tensor incoherence processing at
Step 5’s multi-tensor scale. Fix: seed once, before the
loop, letting MLX’s RNG state advance naturally across tensors —
matching the real code’s own per-layer-continuous-stream semantics
(torch.manual_seed(idx), finetune.py:123).
Verified the bug and the fix directly (not just
re-running the model): simulated both behaviors on tiny arrays — old
(reseed before each draw): SU1 == SU2 → True
(bug confirmed real); new (seed once): SU1 == SU2 →
False (fix confirmed working).
Real gap made explicit, not silently deferred:
05_full_model_quantize.py still calls
loss_and_grad_fn(data[:1]), using only 1 of the real 6
stratified-calibration domain windows built — the per-example batch-loop
fix from earlier this session (bug #2) is technically correctly written
for batch>1 but currently inert. Raising
k_per_domain alone would NOT fix this without also removing
the [:1] slice. Deliberately kept at [:1] for
now: an earlier real attempt to use the full 6-sequence batch on a
SINGLE tensor already OOM’d this session (Part 4c) — risking that again
on a multi-tensor run, on top of an unreviewed fix, was judged the wrong
tradeoff right now. Real calibration scale-up (removing the slice,
re-tuning sigma_reg to match, since the two are coupled) is
real, disclosed follow-up work.
Real re-verification, all 3 subset tensors, with the seed fix (2026-09-03, 22:00:55 → 22:13:04 BST, 12m09s):
mlp.up_proj (UNROTATED): 80.98% -- identical to before (unaffected by the fix, no SU/SV used)
self_attn.q_proj (rotated): 61.94% -- identical to before (first rotated tensor processed, same first-draw)
linear_attn.out_proj (rotated): 66.56% -- CHANGED from 68.02% (now gets an independent rotation, not a repeat)
The change in linear_attn.out_proj’s result (68.02% →
66.56%) is itself real evidence the fix works — a genuinely different,
independent rotation now, not a repeated seed.
2026-09-04. While building Part 5c (the real output
model), a direct question surfaced: does deploying a Hadamard-rotated
tensor need special inference-time handling, or can it be saved and
loaded exactly like GPTQ’s output? Answering this required opening files
that had never been read this session — the real repo’s
inference module, not just its training script
(lib/algo/finetune.py, already ported). This should have
been read back when the Hadamard rotation was first added (Step 4), not
discovered only now at the deployment stage — noted honestly, not
glossed over.
Read directly (local clone:
~/HybridOptiQ_FINAL_BENCHMARK/research/quantization_methods_2026/yaqa/code):
lib/linear/quantized_linear.py — the real deployed
module. Stores SU, SV, had_left,
had_right as real buffers, and a packed
trellis tensor (int16). Forward pass calls
self.codebook_class(input, self.trellis, self.SU, self.SV, self.had_left, self.had_right, ...)
— a trellis-coded vector-quantization codebook
(bitshift-packed indices decoded via a tunable LUT), decoded by custom
kernels (qtip-kernels/), with SU/SV/Hadamard rotation
applied to activations on every real forward call.lib/algo/ldlq.py — the real LDLQ_2hess
driver. Confirms the block-sequential correction loop this port already
ported faithfully, but the actual rounding call is
cb.quantize(target) (line 44) — codebook vector
quantization, not mx.quantize’s plain per-group affine
round-to-nearest.What this means for this port:
yaqa_core.ldlq2hess_quantize has always substituted
mx.quantize() (MLX’s native per-group affine quantizer) for
the real cb.quantize() (trellis codebook) at the exact same
point in the correction loop — the block-sequential
LDLQ_2hess math is faithfully ported, but the final
rounding target was changed. This is a real, necessary substitution
(MLX/Metal has no trellis-VQ kernel; building one is its own multi-week
research/engineering project, not attempted here) — but it was made
implicitly back at Step 3/4 when mx.quantize was first
chosen, and should have been stated explicitly then rather than
surfacing indirectly through a deployment question now.
Consequence for rotation: the real repo’s
SU/SV/Hadamard rotation exists specifically to condition weights for a
fixed trellis codebook — a VQ scheme with no per-group scale
adaptation, so uncontrolled outliers are damaging to it.
mx.quantize is per-group affine — it already fits
its own scale/zero-point to each group’s actual observed range, which
absorbs most of what rotation buys. This project’s own GPTQ baseline
(08_gptq_apply_plan.py) is also affine, also unrotated, and
already ships in production — real, direct precedent that an
affine-quantized deployment doesn’t need runtime rotation to work.
Initial decision (v1, 2026-09-04, later reversed the same day): Part 5c would deploy tensors unrotated (Option B) to avoid the real, unverified engineering cost of a rotation-aware runtime module. Reversed after direct pushback (user, backed by an independent ChatGPT technical review) that Option A — deploying WITH rotation — was more tractable than assessed, and that dropping real algorithm fidelity for engineering convenience was the wrong call without first actually trying the derivation. That pushback was correct: the derivation below took under an hour and is fully verified.
Visual comparison of the real reference pipeline vs. this port’s
pipeline: 00_DOCS/HTML/yaqa_pipeline_comparison.html
(project’s real HTML docs folder).
yaqa_core.YaqaQuantizedLinear, derived
and verified2026-09-04, same day as the reversal above.
Hand-derived the real runtime forward pass from
rotate_weight’s own formula (already gold-standard-verified
against real PyTorch). Expanding
rotate_weight(W, SU, SV) = matmul_hadUt(matmul_hadUt(W.T * SV).T * SU)
algebraically:
Wr = H_out^T · diag(SV) · W · diag(SU) · H_in (H_in/H_out
being the orthogonal operators matmul_hadUt implements
along each axis). Solving for W and substituting into
y = x @ W.T gives:
x1 = x * SU
x2 = hadamard(x1)
y1 = x2 @ Wr.T (quantized matmul against the ROTATED, packed weight)
y2 = hadamard(y1)
y = y2 * SV
Verified numerically (not just algebraically): this
exact sequence reproduces x @ W.T (the real, unrotated
forward pass) to max|diff|=2.98e-07 (float32 precision) on
a synthetic case — see test_regression.py test [8] for the
permanent guard.
Real primitive discovery, changes the plan for the
better: MLX ships mx.hadamard_transform (native
Walsh-Hadamard, real Metal kernel, n = m·2^k for
m ∈ {1,12,20,28}) and mx.quantized_matmul
(native quantized matmul) — both public, documented, real functions
(mlx==0.32.2, checked directly via help()).
Two things checked before trusting them: (1)
mx.hadamard_transform matches this port’s own hand-rolled
matmul_hadUt exactly (0.0
diff) at a tested power-of-2 size; (2) real coverage across all 363 plan
targets is identical between MLX-native support and the
real repo’s own 12-divisor table — 173 tensors covered by both, 190 by
neither, zero tensors covered by only one. No coverage
tradeoff from standardizing on the native primitives for the runtime
module.
yaqa_core.YaqaQuantizedLinear implements exactly the
derived sequence above using
mx.hadamard_transform/mx.quantized_matmul.
Full pipeline verified end to end (not just the
algebraic shortcut): real rotate → real
regularized_block_ldl → real
ldlq2hess_quantize → real packed codes → module forward,
compared against the real corrected weight’s own forward pass —
max|diff|=3.58e-07 (float32 precision).
test_regression.py test [8] guards this permanently.
mx.quantize is NOT
idempotent on its own dequantized outputPart 5c’s first draft (Option B) re-quantized the assembled
hatW after ldlq2hess_quantize to get
deployable packed codes, on the assumption this was a lossless,
idempotent operation (each block’s stored value already being the
dequantized result of a real mx.quantize call).
Checked directly, found wrong:
mx.quantize(mx.dequantize(mx.quantize(W))) does NOT
reproduce the same codes as the first quantize —
max|diff|=3.7e-3 on a synthetic 128×256 array (~18% of
packed int32 words changed), stabilizing only after a SECOND
re-quantization (not floating-point noise; a real, one-time rounding
drift). This threatened the correctness of Part 5c’s whole pack step,
for both the rotated and unrotated paths.
Fix: yaqa_core.ldlq2hess_quantize now
captures the real packed (wq, scales, biases) directly from
each block’s own mx.quantize() call during the sweep,
returning them alongside hatWr, instead of reconstructing
them via a second quantize call afterward. Verified safe by first
checking that each block position is visited exactly
once across the real anti-diagonal sweep (checked directly on a
synthetic 5×7 grid — zero double-visits), so no overwrite conflict is
possible. test_regression.py guards both the bug staying
fixed (the returned codes dequantize to EXACTLY hatWr, zero
tolerance) and the original NumPy cross-check (unaffected by the fix,
still passes).
--incoherence {none,hadamard}05_full_model_quantize.py --output <dir> does a
real save. --incoherence hadamard (default): the 173/363
rotatable tensors deploy via YaqaQuantizedLinear using the
real packed codes directly in the rotated basis (no inverse-rotation, no
re-quantize); the other 190 deploy via standard
to_quantized(), same as before.
--incoherence none: every tensor deploys via standard
to_quantized(), matching
08_gptq_apply_plan.py’s save path exactly. Either way,
packing uses the real captured codes (the fix above), not a
re-quantize.
Any plan tensor a --subset run doesn’t cover is left at
full BF16 (disclosed, not silently native-quantized) rather than paying
real cost for a run whose whole point is to validate the save mechanics
cheaply; a full run (no --subset) has no such tensors.
Config dict mirrors 08_gptq_apply_plan.py’s own structure,
including its BUG FIX #3 (top-level bits/
group_size defaults mlx_lm.utils.load()’s
reload path reads unconditionally) — rotated tensors are recorded as
False in that config (so mlx_lm’s own
_quantize() doesn’t try to misreconstruct them as plain
QuantizedLinear), plus a separate
yaqa_rotated_manifest.json written into the output
directory listing every rotated tensor’s real bits/group_size/shape.
Real calibration richness fix in the same change: the
previously-disclosed data[:1] gap (Part 5a) is closed by
chunking calibration into --batch-size-sized real
forward+backward calls, accumulating H_I/H_O
across calls the same way the vjp already accumulates across the batch
dimension within one call — peak activation memory now bounded by
--batch-size regardless of total calibration size (raising
--n-calibration alone, on one giant single-shot call, would
not have been enough — that’s what OOM’d a single tensor in Part
4c).
The real, unresolved gap:
mlx_lm.utils.load() has no concept of
YaqaQuantizedLinear and cannot reconstruct it from
yaqa_rotated_manifest.json — a real, separate loader is
required before a --incoherence hadamard model is actually
runnable. The script prints a loud warning and writes the manifest, but
does not claim the model is loadable.
Production decision (2026-09-04, later the same
day): after direct review of the real compatibility/runtime
tradeoffs — Option A needs a new loader everywhere the model is ever
loaded (not just here), plus a permanent per-forward-pass
Hadamard-transform cost on 173/363 tensors, for an unproven quality
benefit against MLX’s self-adapting affine quantizer —
--incoherence default flipped from hadamard to
none. hadamard remains a real, fully built and
verified research branch (not deleted, not degraded) to benchmark later,
after cascading + stratified-KL work, before deciding whether the loader
is worth building.
2026-09-04, found by actually inspecting a real saved model’s
safetensors keys instead of trusting exit code 0 (the user
explicitly demanded this level of verification after the idempotency bug
above). The first real --output smoke test exited 0 and
printed a normal summary. Direct inspection of
model.safetensors.index.json showed every one of the 3
target tensors saved under a nonsense key:
language_model...mlp.up_proj.module.weight instead of
language_model...mlp.up_proj.weight — and with only
weight/scales/biases present, no SU/
SV, meaning even the two tensors meant to be rotated
weren’t.
Root cause: model.leaf_modules()
defines “leaf” as a Module with no further nn.Module children —
not “the position named in the plan”. Before wrapping, a target position
(e.g. mlp.up_proj) holds a plain nn.Linear,
which qualifies as a leaf named mlp.up_proj. Once wrapped
in YaqaCatcher (which itself has a .module
child holding the real nn.Linear), that position no longer
has zero Module children — leaf_modules() stops treating it
as a leaf and instead descends into it, reporting
mlp.up_proj.module as the leaf. The restore step used to
re-query tree_flatten(model.leaf_modules(), ...)
after wrapping, so it saw this shifted key set;
orig_modules.get("mlp.up_proj.module", l) always missed
(orig_modules is keyed by the real, unshifted plan names),
silently fell back to l (the unchanged inner
Linear it had just found), and the “restore” was a complete
no-op — the model stayed wrapped in YaqaCatcher for the
rest of the run, silently, for every run this session that reached this
code path.
Reproduced cheaply first, before touching the real
file: a 2-layer toy model with the same wrap/restore sequence
showed model.a was still Catcher after
“restore” — confirmed the bug in under a second, no need for the 27B
model to diagnose it.
Consequence for the real smoke test: Part 5c’s
all_leaves = tree_flatten(model.leaf_modules(), ...) call
(used to decide what to save for every leaf) saw the same shifted keys
for all 3 target tensors. if name in q_layers never matched
(comparing the real plan name against a shifted key), so every real
YAQA-corrected tensor fell through to the fallback-native-quantize
branch — wrong bits (--fallback-bits=6 instead of the real
plan’s 8/8/5), wrong weight (plain native rounding, no YAQA correction
at all), saved under the wrong key. The real Hessian-collection
metrics reported earlier in this ledger (75%, 83%, 68.65%, etc.) are
unaffected — those read orig_modules[name].weight
directly, never through the corrupted live tree — but every real SAVE
this session produced before this fix is invalid and has been
deleted.
Fix: capture the pre-wrap leaf key list once
(orig_leaves_list), reuse it for both the wrap and
the restore — never re-derive leaf positions from
model.leaf_modules() after the tree has been mutated by
wrapping. Added a hard runtime check immediately after restore that
verifies every target position is back to its original type, raising
loudly instead of silently continuing if it isn’t. Verified the fix on
the same toy-model reproduction (restore now correctly shows
Linear, and a subsequent quantized replacement lands under
the clean a.weight key, no .module. shift) —
test_regression.py test [9] guards this permanently,
cheaply, without needing the real model. Real smoke test re-run with the
fix in progress — result recorded below once verified against actual
saved, reloaded output, not before.
2026-09-04, after Part 5c’s first real save actually loaded
and ran correctly. Before trusting
--n-calibration 4 for a real run, computed the real total:
all 363 plan targets’ H_I+H_O held
simultaneously need 557GB (310GB even excluding
lm_head, whose 248,320-wide output dimension alone needs
246.7GB — excluded separately, matching the real reference repo’s own
precedent of not running its per-layer correction loop on this tensor
either). The earlier claim in this file that “128GB comfortably covers
the largest real Hessian” checked only the single largest tensor, never
the sum across all targets at once — found only after a real
--n-calibration 4 run OOM’d (exit 137).
Fix: --hessian-batch-budget-gb splits
targets into memory-bounded batches, each getting its own real
forward+backward pass over the full calibration data.
Second real crash, found immediately by testing the fix on
real data: batch 2 of a forced 3-batch test failed with
RuntimeError: [metal::malloc] Resource limit (499000) exceeded
— a hard, real cap Metal puts on GPU resource handles created over a
process’s lifetime, not a byte-size limit. Tried
mx.clear_cache() (real, native MLX/Metal API) —
tested directly, confirmed it does NOT reset this
(identical crash, same batch, same spot, re-run twice).
Real fix: each batch now runs as its own separate OS
process. New CLI plumbing on 05_full_model_quantize.py
(--resume-dir, --batch-only N,
--print-num-batches, --assemble-from-resume)
plus a new orchestrator, run_full_yaqa.sh, that queries the
real batch count, runs each batch as a fresh process (writing results to
a resume-cache directory, same real idiom as
08_gptq_apply_plan.py’s own ResumeCache), then
does one final --assemble-from-resume pass that loads every
batch’s cached result and performs the real save. Proven on real
data: re-ran the same 3-batch test with the fix — the exact
tensor that crashed twice (self_attn.q_proj) completed
cleanly as its own process (78.20% real error reduction). Full 3-batch
run completed, assembled, saved, and — the actual bar for “proven,” not
just “didn’t crash” — loaded via plain
mlx_lm.load() and ran a real forward pass, predicting
“Paris” for “The capital of France is.”
A real bug found in the fix itself, found and fixed the same
session: the resume-cache manifest write originally happened
before
err_yaqa/err_naive/pct_reduction/frob_*
were computed, so --assemble-from-resume had only NaN
placeholders for those fields — the final “N/M tensors reduce error”
summary printed “0/M” even though every tensor’s own per-batch log
showed real, large reductions (83%, 78.2%, etc.). Wrong summary line,
not a wrong saved model — the packed weights were always correct; only
the reported count was wrong. Fixed by moving the manifest write to
after the real metrics are computed, and verified cheaply (synthetic
manifest, no real model load needed) rather than paying for another real
run to check a pure bookkeeping fix.
2026-09-04, raised directly by the user, verified by reading the real source rather than arguing from memory. The batching fix above solves the real memory and Metal-resource crashes, but it does so by re-running a full forward+backward pass through the entire 64-layer model once per arbitrary memory-sized batch — for the real full run (~39 batches at the default 8GB budget), that is ~39 redundant passes through every layer.
The real reference
(quantize_llama/quantize_finetune_llama.py, read directly,
lines ~160-215) does not do this. It processes layers
sequentially, one at a time
(for i in range(len(model.model.layers))), streaming each
layer’s real output forward as a cache for the next layer’s input — each
layer’s forward pass runs exactly once, not once per arbitrary batch.
(Nuance checked, not assumed: this streaming does NOT feed
already-quantized layers forward into later layers’ calibration — each
layer’s input still traces back through the original, unquantized chain
— so the real design is a genuine efficiency/memory optimization, not a
“cascading correction” scheme layered on top.)
This was not caught during design review — it surfaced only when directly challenged on why the port wasn’t processing the network “like MLX OptiQ does, one tensor at a time.” Batching by arbitrary memory size instead of by network structure is a real, disclosed architectural gap: it produces correct numbers (proven above) through a materially less efficient path than the real algorithm’s own design.
Follow-up, same session, checked directly rather than
assumed: read the real Hessian-collection driver itself,
hessian_llama/get_hess_llama.py (a separate script from
quantize_finetune_llama.py, not read before this point).
Real finding: the reference’s OWN collection stage does NOT stream
layer-by-layer either — it wraps every layer’s tensors simultaneously
and runs ONE real forward+backward pass per calibration batch
(torch.nn.functional.cross_entropy(logits, fake_target).backward()),
accumulating every layer’s Hessian at once, exactly like this port’s
original (pre-batching) design. It fits on their real hardware via
FSDP (sharding the model and its Hessians across
multiple GPUs) plus
CPUOffload(offload_params=True) — infrastructure a single
128GB Mac does not have. Their n_seqs default is 65,536
sequences at ctx_size=2048 (~134M tokens) — real production
scale, not something a single machine holds regardless of
architecture.
Honest conclusion: the 557GB simultaneous-Hessian requirement is not a porting mistake — it is the real algorithm’s actual behavior at real production scale, made feasible upstream by multi-GPU hardware this project doesn’t have. Some form of splitting the work is genuinely unavoidable here, not a wrong design choice. What WAS a real, fixable mistake: splitting arbitrarily instead of along the network’s real structure. Fixed the same session: targets are now sorted by real layer index (extracted from the real dotted tensor name) before batching, so tensors from the same real transformer layer land together wherever the memory budget allows, instead of being scattered by whatever order the plan’s JSON happened to list them in. Real per-layer streaming (avoiding repeated full-model passes entirely, matching neither reference script exactly since both hold everything at once) remains real, deliberate, disclosed future work — not something either reference script actually does either, so not a gap relative to them specifically, but still a real efficiency opportunity for this project’s own single-machine constraint.
2026-09-04, prompted directly by the user asking why the OOM
fix couldn’t be applied to lm_head too. Direct,
correct answer worked through with them: batching fixes “too many
tensors’ Hessians held together at once” – it cannot shrink
one tensor’s own Hessian. lm_head’s
H_O is 246.7GB by itself, in a batch of exactly one, so
batching was never going to reach it.
The real, different fix: lm_head’s
problem is specifically its OUTPUT side (tied to the 248,320-word
vocabulary). Its INPUT side (in_features=5120) is
completely normal. GPTQ only ever needs the input side
(H = x.T@x) – it was never going to hit this wall in the
first place. Confirmed directly against the real source:
08_gptq_apply_plan.py has no special exclusion for
lm_head at all; it already corrects it automatically as an
ordinary nn.Linear.
Checked before reusing 08_gptq_apply_plan.py’s
own gptq_quantize_with_plan() directly, not assumed
safe: that function wraps every real
nn.Linear/SwitchLinear leaf simultaneously
(497 of them). Real sum of their H matrices
(in_features² each, GPTQ never needs H_O):
125.93GB – dangerously close to this machine’s full
128GB. Reusing it as-is would have risked a third real OOM. Rejected in
favor of a properly scoped alternative.
Built: 06_lm_head_gptq.py – wraps
only lm_head with Apple’s own real
mlx_lm.quant.gptq.Catcher (read directly, unmodified,
first-party), collects its real H from a real forward-only
pass over calibration data (no backward pass needed – GPTQ never needs
gradients), runs the same real column-by-column compensation math as
08_gptq_apply_plan.py’s own function (reproduced here
rather than imported, since the original nests it as local functions),
and packs the result via mx.quantize – native MLX, same
format as every YAQA-corrected tensor, not GPTQ’s own custom bit-packer
– so it lands in the exact same manifest schema with zero special-casing
needed downstream.
Real safety property, explicitly checked and answered when
asked: does this contaminate the other 362 real tensors’
measurements? No – verified this follows directly from the SAME real
property already regression-tested (test [10]): every batch reloads the
pristine original BF16 model fresh, never a partially-corrected one, so
lm_head’s calibration pass (run separately, at a different
time) sees exactly the same untouched original weights every other
tensor’s measurement was anchored to. Order and timing don’t affect
correctness, only scheduling.
Manifest entry format: adds
"method": "gptq" (the 362 YAQA entries carry no
"method" key, implicitly YAQA) plus
frob_gptq/frob_naive in place of the
Hessian-weighted
err_yaqa/err_naive/pct_reduction
fields (a one-sided correction has no real two-sided-weighted metric to
report – 05_full_model_quantize.py’s existing
.get(..., float("nan")) fallbacks already handle this
honestly, verified with a synthetic mixed yaqa+gptq manifest, no code
change needed there).
Streamlined as a standing rule, not a one-off: wired
into run_full_yaqa.sh to run automatically after every
batch completes and before the final assembly, every time – never a
manual step to remember. Also made resume-aware (same real per-tensor
skip pattern), so an interrupted run doesn’t redo it.
Result: lm_head is real, GPTQ-corrected
– not YAQA’s full two-sided correction (structurally impossible on this
hardware), but strictly better than the zero-correction fallback it had
before, for the same real reason GPTQ-vs-naive already shows real
measured improvement on every other tensor in this project.
2026-09-05. --source and
--plan were already real overrides. Found and fixed the
remaining one: LM_HEAD_NAME (the vocabulary-output tensor’s
dotted name) was still a hardcoded constant in both
05_full_model_quantize.py and
06_lm_head_gptq.py. Now --lm-head-name on
both, forwarded automatically by run_full_yaqa.sh. This
codebase no longer assumes Qwen3.5/3.8’s own naming for any of its
architecture-defining constants — a different MLX model with its own
real plan can be pointed at directly.
Proposed and documented in
00_DOCS/HTML/yaqa_batching_architecture.html:
YAQA-UMA — YAQA for Unified Memory Architecture. Not
decorative: Apple Silicon’s real, shared CPU/GPU memory design is the
specific constraint the batching + resume + separate-process
architecture in this ledger was built to work within, as distinct from
the reference’s own multi-GPU + CPU-offload design. That page also
carries the real, spatial batching-architecture diagram (mermaid), a
table of what’s been verified end to end (not assumed), and the real
numbers behind the memory problem (557GB needed, 128GB available,
246.7GB for lm_head alone).
YaqaQuantizedLinear. Real, necessary follow-up
before a rotated model can be loaded and run at all. Not started.--n-calibration 1) — the mechanism now
supports raising it safely (chunked accumulation, see above), but no run
has actually been done yet at a larger real scale, and
--sigma-reg may need re-tuning alongside it (1.0 was
empirically tuned at the old single-sequence scale).Trigger: the live mission-control dashboard’s
dynamic quality narrative surfaced a real outlier –
language_model.model.layers.14.linear_attn.in_proj_qkv
showed pct_reduction = -1788.48% (YAQA’s weighted error
~18.9x worse than naive). User correctly refused to accept the
pipeline’s own self-reported number at face value and demanded
independent verification.
Independent verification (not reusing any pipeline
error-computation code): loaded the real original weight fresh
from the source model, ran a fresh mx.quantize for naive,
loaded the real cached YAQA result and dequantized it separately,
computed Frobenius error myself in a one-off script:
Independent naive Frobenius: 2.543 (manifest: 2.594 -- matches, small dtype rounding diff)
Independent YAQA Frobenius: 767.92 (manifest: 768.0 -- matches, confirms NOT a reporting bug)
mean |original weight value|: 0.0126
mean |naive dequant error|: 0.0003 (2.3% of typical weight size -- normal)
mean |YAQA dequant error|: 0.0291 (231% of typical weight size -- genuinely broken)
Confirmed real, not a metric/reporting artifact.
Root cause, precisely (from a real code review of
ldlq2hess_quantize/regularized_block_ldl in
yaqa_core.py): the two-sided correction propagates
each block’s rounding error forward via
Lout/Lin, derived from
block_ldl_mlx’s Cholesky-based factorization of the real
Hessian. If the Hessian is ill-conditioned (a few directions with much
larger eigenvalues than the rest), the diagonal-block inverses used to
build L can become large, which makes the propagated-error
term amplify the residual instead of damping it – a real,
mechanical explanation for a 302x blowup, not a vague “instability.”
Hadamard rotation’s real, documented purpose in this method family
(QuIP/QTIP/YAQA) is specifically to prevent this: it spreads the
Hessian’s energy evenly across directions, keeping the Cholesky blocks
well-conditioned. This tensor’s own log line:
NOTE: has a valid Hadamard factor but --incoherence=none -- deploying unrotated by choice, not limitation.
– real, deliberate production choice, real quantified cost on this one
tensor.
The real gap in the existing safety net:
regularized_block_ldl’s retry loop (real code, re-read for
this RCA) only checks mx.isfinite – it retries with more
damping ONLY on NaN/Inf. This tensor’s block_LDL reported
succeeded on attempt 1 – technically finite, still
catastrophic. The safety net catches outright numerical failure, not
“valid-looking but insane magnitude.” That is the real, fixable gap, not
“needs rotation” alone.
Options considered, with real tradeoffs: | Option |
Real verdict | |—|—| | A. Fallback to naive for verified losses
(err_yaqa >= err_naive in the manifest) |
Shippable tonight. Zero risk, uses data already
computed, purely additive safety net. | | B. Build the
YaqaQuantizedLinear loader, rotate eligible tensors |
Off the table for this run, not just “more work”:
mlx_lm.utils.load() cannot load a model containing ANY
rotated tensor at all (real, hard incompatibility, confirmed directly in
the module’s own docstring) – rotating even one tensor makes the whole
saved model unloadable by standard tooling until a separate loader
exists. Real future work, zero effect on tonight’s ship. | | C. Harden
regularized_block_ldl’s own check (retry with more damping
if the result would be worse than naive, not just on NaN/Inf) | Real,
generalizable fix for the actual root cause – catches this failure mode
automatically for any future tensor/model. Needs its own real testing
before trusting it on a live run; disclosed as follow-up, not applied
tonight. | | D. Raise the global sigma_reg default |
Rejected – already documented as tuned for the old single-sequence
calibration scale; a blind global change risks the 287 tensors already
working well. |
Decision: A now (real, safe, additive), C as real disclosed follow-up work, B not pursued for this run given the hard loader incompatibility.
MLX-native note: the correction sweep is inherently
sequential (each block depends on the previous block’s residual – the
algorithm itself, not an MLX limitation), so there is no real
vectorization opportunity in the core loop. The one genuine, real,
MLX-aware improvement identified: a cheap pre-flight conditioning check
(e.g. mx.linalg.eigvalsh on the normalized Hessian, or the
max diagonal-block-inverse magnitude) before running the full expensive
sweep, flagging a dangerously ill-conditioned tensor ahead of time
instead of discovering it only after computing a correction that gets
discarded anyway.
mx.clear_streams() as an alternative to
one-process-per-batch (2026-09-05)Tested directly, isolated from the live run (separate resume dir,
real model, real data): running multiple real batches inside one
process, calling mx.clear_streams() between them, to see if
it could avoid the separate-process-per-batch architecture. Result: it
does not work — batch 0 completed fine, but clear_streams()
destroys the default GPU stream MLX itself needs, so batch 1 crashed
immediately with
RuntimeError: There is no Stream(gpu, 1) in current thread
(a different, earlier failure than the original resource-limit crash it
was meant to prevent). Confirmed: the current separate-process-
per-batch design remains the best real approach given this machine’s
constraints — no cheaper alternative found.
This entry corrects and completes the RCA above. The
2026-09-05 entry correctly identified that
layers.14.linear_attn.in_proj_qkv has an ill-conditioned
Hessian and that the correction amplifies error under those
conditions — but it did not yet know why the amplification was
so extreme on this specific run. Direct investigation the following day
found the actual mechanical cause, and it changed the fix.
Real finding: the Hessian curvature accumulation
(yaqa_linear_vjp in 05_full_model_quantize.py)
was running in bfloat16, not float32. x
and grad_output are this model’s native bf16 compute dtype;
the code assumed the matmul Gb = gb.T @ xb would promote to
float32 the way float32 + bfloat16 addition does in MLX. It
doesn’t — confirmed with a direct, isolated dtype probe:
bf16 @ bf16 stays bf16, no auto-promotion, unlike
addition.
Real, controlled 3-way measurement on this exact tensor, bf16 vs. fp32 vs. a CPU-float64 gold reference:
| bf16 (broken) | fp32 (fixed) | float64 (gold reference) | |
|---|---|---|---|
H_I min eigenvalue |
-5.618 (real, large-magnitude PSD violation) | ~-3e-4 (numerical noise floor) | matches fp32 to 4 decimals |
H_O min eigenvalue |
-7.534 | ~-8e-4 | matches fp32 to 4 decimals |
| YAQA weighted error vs. naive | -0.05x (i.e. worse than naive) | 0.0009x | 0.0009x (identical to fp32) |
| Frobenius error vs. naive | 3.84x (worse than naive) | 1.0225x | 1.0225x (identical to fp32) |
A true sum-of-Gram-matrices Hessian is positive semi-definite in exact arithmetic — eigenvalues of -5.6/-7.5 are six to seven orders of magnitude past what float32’s own rounding noise could ever produce, which is what made this a real, provable precision bug rather than “the Hessian is just naturally a bit noisy.” fp32 reproduces the float64 gold reference’s real downstream objective to 4 decimal places — no CPU float64 pipeline is needed in production; float32 already matches it and runs natively on GPU.
The fix: cast x and
grad_output to float32 before forming
Gb (casting Gb afterward is too late — the
precision is already lost in the matmul itself), in both the YAQA
curvature path and, separately, the GPTQ path used for
lm_head. Added CURVATURE_VERSION = 2 so every
pre-fix cached tensor (bf16 curvature) is automatically recomputed on
resume, rather than requiring a manual cache wipe.
How this reconciles with the 2026-09-05 ill-conditioning RCA,
precisely: both are real and both matter, they’re not competing
explanations. Ill-conditioning is why this specific tensor was
so sensitive — a well-conditioned Hessian would have absorbed the same
bf16 precision loss without a visible effect. BF16 precision loss is the
actual mechanical cause of the corrupted computation. The fix that
actually worked (FP32) treats the real cause; Option C from the prior
entry (harden regularized_block_ldl’s retry logic) is no
longer the load-bearing fix for this specific failure, though it remains
reasonable defense-in-depth and is still real, disclosed follow-up
work.
Real production impact: this wasn’t a one-tensor
fix. All 336 tensors already processed under the bf16-curvature bug get
auto-invalidated by CURVATURE_VERSION and recomputed
correctly on the current production run. Fully documented,
plain-English, with diagrams:
research_hadamard_blowup/exports/BF16_Curvature_Bug_Plain_English.html
(and its from-first-principles technical companion,
YAQA_UMA_Hessian_Precision_Root_Cause.html).
Missing from this ledger until 2026-09-07 (later) — added retroactively in its real place. The real 363-tensor production run (Step 5-full, above) completed for the first time this day. Two real, measured bugs surfaced only after benchmarking the finished model, not before:
mode: "affine" in
config.json — every real per-tensor quantization
config dict this script wrote only ever carried
{"bits", "group_size"}, never mode, even
though every real mx.quantize()/to_quantized()
call always used MLX’s own affine default. Packed weight bytes were
never wrong — only the metadata describing them was incomplete,
matching neither the reference build’s config.json nor what
mlx_lm’s own loader expects declared explicitly. Real,
measured cost: ~27-87% slower decode across every real benchmarked depth
(AR 14.06 vs 17.91 tok/s; best-depth 28.83 vs 54.08 tok/s). Full
evidence, every real quantize/dequantize call site checked exhaustively:
research_hadamard_blowup/exports/AFFINE_MODE_CODE_REVIEW.html.bits/group_size/mode mismatches
(never a bit-width bug), 727/1912 dtype mismatches (every one
F32-vs-BF16). Real, measured cost: +17.2% peak memory, 20-55% decode
slowdown. Same investigation found embed_tokens was being
quantized (a real, separate confound for apples-to-apples comparison —
the reference leaves it uncompressed) — fixed to match. Full RCA,
comprehensive tensor table, exact rebuild command:
research_hadamard_blowup/exports/RCA_scales_f32_bug.html.Rebuilt via --assemble-from-resume after each
fix — the real 26-hour Hessian correction was never re-run;
each rebuild just re-saved the already-corrected resume-cache tensors at
the right width/config. Independently re-verified against the
actual rebuilt model, not just code-reviewed: tensor forensics
re-run shows 0/1910 tensors differing from the reference in
dtype/bits/group_size/mode/size; a real 4-way weight-provenance check
(original source weight vs. naive quantization vs. resume-cache
correction output vs. live shipped model) confirms the real corrected
values shipped, not naive quantization standing in for them (live
matched resume-cache to bf16-rounding precision, ~0.1-0.6%; differed
from naive by 1.9-4.5% absolute). Resume-manifest analysis: 362/363
tensors used real YAQA correction (the 1 exception,
lm_head, is the known, expected
naive_fallback), real Hessian-weighted error reduction
vs. naive ranged 97.49-100%, median 99.67%, zero tensors worse than
naive.
Re-benchmarked (mtplx tune): AR and D2
now at real parity with the reference (-0.14%, -1.06%); D1 and D3 still
lag (~-20%), with real acceptance-rate degradation concentrated at
deeper speculative-decode windows (D3 depth-3: reference 94.74% vs. YAQA
80.60%). Not yet explained by any single localized failure in the
manifest analysis above. Real output-quality benchmark (same 5-task
protocol as every other model in
scripts/yaqa_port/exports/Benchmark_Master_Recap.html)
launched 2026-09-07 ~16:04 to test this independently of tok/s; MMLU
result already in — 87.7% (100/114, 95% CI ±6.0pp) vs. the v33floor
reference’s 89.5%, a real -1.8pp gap within the stated CI. Remaining 4
tasks in progress as of this entry.
Also fixed the same day, project-wide: every HTML
doc built via build_html_docs.sh was rendering its own
title twice (pandoc’s title-block header plus the body’s own matching
first heading) — root-caused and fixed once in the script itself; see
CHANGELOG.md’s 2026-09-07 “even later” entry.
The trunk correction above left the MTP speculative-decode sidecar
naive-quantized (plain mx.quantize, identical to the
reference baseline) – a real architectural gap, not a design choice: a
YAQA-corrected trunk feeding a naive-corrected draft head. Root-caused,
planned, implemented, smoke-tested, and launched as a real full build
this same day – full detail kept in CHANGELOG.md’s
2026-09-08 entries (not duplicated here) and
research_hadamard_blowup/exports/MTP_YAQA_CORRECTION_PLAN.html
(the original plan and real evidence). Real launch command and the two
real monitoring-gap fixes it surfaced (monitor.sh
process-match, generate_dashboard.py missing MTP milestone)
are in CHANGELOG.md’s “Launched (2026-09-08, later)”
entry.
Two further real gaps found and fixed while actually running the
launch, both in CHANGELOG.md’s “later still” entries: a
--reuse-cached-fallback skip option added to
06_lm_head_gptq.py (lm_head’s damping search is
deterministic, so retrying a prior naive_fallback result
wastes ~15-20 real minutes reaching the identical outcome) plus the
matching argparse fix 05_full_model_quantize.py needed to
not crash on that same forwarded flag; and a real
cp -a-into-existing-directory pitfall that silently doubled
the resume cache’s disk usage (16GB -> 32GB) when Step 1 of the
launch procedure was re-run a second time. Exact, current procedure (all
steps, all flags): YAQA_COMMAND_REFERENCE.md’s “Adding
YAQA-corrected MTP sidecar” section.
docs/HTML/YAQA_Why_It_Matters.html — public-facing “why
it matters” explainer (dark theme, diagrams). Separate from this ledger
— that page explains the concepts; this ledger tracks
implementation progress.research/quantization_methods_2026/YAQA_MLX_Port.md /
exports/YAQA_MLX_Port.html — the original feasibility
research and staged plan this ledger is now executing against.research/quantization_methods_2026/PORTING_PLAN.md —
the original 5-step plan definition.© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.
GODMODE-v2 completed all five tasks at n=200 (mean 92.58, against 93.04 for HESSIAN-PROBE). The benchmark recap, the floor analysis and the master timeline were updated to match. Chapter 11 of the research ledger records the GODMODE builds and results. Open: Stratified at n=200, the V3 build, and the BF16 reference with KL.