2026-09-07
Author: Hakim Ghelab, VegaLaboratories LTD
This document answers, with direct evidence and no hand-waving, the
four real questions raised after the speed investigation: what
05_full_model_quantize.py actually is, whether the missing
"mode": "affine" config field means the real YAQA
correction math was ever wrong, whether embed_tokens is
safe to quantize and what precision it’s actually at compared to the
reference model, and what was actually fixed.
What “the reference model” means, throughout this
document: this project builds two things side by side — (1)
this YAQA port, the code this document is reviewing, and (2) a
separately pre-existing, already-trusted build
(HybridPareto5bpw-V3.3_Stratified, produced by this
project’s own older, already-validated
05_build_final_model_V5*.py scripts) that YAQA’s real
output gets compared against for speed and quality. “The reference
model” / “the reference build” always means that second, pre-existing
model — never a third-party or external model. It matters here
specifically because it’s the only way to tell whether a difference in
YAQA’s output (like embed_tokens’s bit-width, below) is a
real bug or just two independent scripts making their own separate,
both-legitimate default choice.
05_full_model_quantize.py is the core script of
this entire YAQA port — it is where the Hessian collection, the
real two-sided correction, and the final model assembly all happen.
Confirmed directly from its own file header.mx.quantize() and mx.dequantize() call in
the whole pipeline — every single one, checked exhaustively, not sampled
— used the same implicit default (mode="affine"). The
missing "mode" key lived only in the final
config.json metadata, written in a completely separate part
of the script (Part 5c) that runs after every real correction
has already been computed and saved.embed_tokens was never a YAQA correction target
at all, and is not quantized at all today. It isn’t in the real
plan file, and a real, genuine bug that once quantized it anyway (via a
generic fallback path meant for out-of-plan Linear layers)
was found and fixed the same day — it’s now full-precision BF16 in both
the YAQA build and the reference, verified directly from both models’
real tensor headers.05_full_model_quantize.py?Its own file header, read directly, not paraphrased:
#!/usr/bin/env python3
"""
Author: Hakim Ghelab, VegaLaboratories LTD
YAQA port -- Step 5 of 5 (staged plan): full model.
This is Step 5 of the real, staged YAQA port plan — the final, full-model production script. It is not a side tool. It is where:
H_in/H_out
for every target tensor)happen, in that order, for every real production run this whole session has been monitoring.
This is the question that actually matters, and it deserves the most rigorous answer, not a reassurance.
Part 5b (the correction itself) computes real
numbers — it runs the Hessian-weighted block-LDL sweep, decides how to
round each weight, and produces real packed
(weight, scales, biases) arrays via real calls to
mx.quantize().
Part 5c (assembly and save), which runs
after Part 5b is completely finished, takes those
already-computed arrays and writes them to disk, alongside a
config.json that describes what was done —
including a "mode" field that’s supposed to say which
quantization scheme was used.
The bug was only in what Part 5c wrote down — never in what Part 5b computed. Here is the real, exhaustive evidence for that claim.
mx.quantize() call in the entire pipelineChecked directly, every file, every call site — not a sample:
yaqa_core.py:172 w_q, scales, biases = mx.quantize(thing, bits=bits, group_size=group_size)
yaqa_core.py:246 wq_final, scales_final, biases_final = mx.quantize(hatW, bits=bits, group_size=group_size)
yaqa_core.py:416 naive_wq, naive_s, naive_b = mx.quantize(W, bits=bits, group_size=group_size)
05_full_model_quantize.py:907 naive_wq, naive_scales, naive_biases = mx.quantize(Wr_full, bits=bits, group_size=group_size)
05_full_model_quantize.py:940 naive_wq, naive_scales, naive_biases = mx.quantize(Wr_full, bits=bits, group_size=group_size)
06_lm_head_gptq.py:102 _, scales, biases = mx.quantize(Wl, bits=bits, group_size=group_size)
06_lm_head_gptq.py:113 wq, scales, biases = mx.quantize(Wc, bits=bits, group_size=group_size)
06_lm_head_gptq.py:422 naive_wq, naive_scales, naive_biases = mx.quantize(W, bits=bits, group_size=group_size)
Every single one passes only bits= and
group_size=. None of them ever passes a
mode= argument. That means every one of them used MLX’s own
real default — which, checked directly from
mlx.core.quantize’s actual signature (not assumed):
quantize(w, group_size=None, bits=None, mode: str = 'affine', ...)
mode="affine" is the default. Every real quantize call
in this entire codebase — the one inside YAQA’s own correction sweep at
yaqa_core.py:172 included — used it.
mx.dequantize() call — checked the same wayThis matters just as much: if quantize and dequantize ever used different modes anywhere, that would be a genuine correctness bug (comparing apples to oranges when measuring how good a correction is). Checked, exhaustively:
yaqa_core.py:173 hatWr[...] = mx.dequantize(w_q, scales, biases, bits=bits, group_size=group_size)
yaqa_core.py:417 naive_dq = mx.dequantize(naive_wq, naive_s, naive_b, bits=bits, group_size=group_size)
05_full_model_quantize.py:908 naive_hatWr = mx.dequantize(naive_wq, naive_scales, naive_biases, bits=bits, group_size=group_size)
05_full_model_quantize.py:941 naive_hatWr = mx.dequantize(naive_wq, naive_scales, naive_biases, bits=bits, group_size=group_size)
06_lm_head_gptq.py:115 ...mx.dequantize(wq, scales, biases, bits=bits, group_size=group_size)...
06_lm_head_gptq.py:155 corrected_W = mx.dequantize(wq, scales, biases, bits=bits, group_size=group_size)
06_lm_head_gptq.py:423 naive_W = mx.dequantize(naive_wq, naive_scales, naive_biases, bits=bits, group_size=group_size)
Same result: every dequantize call, no exceptions, uses only
bits=/group_size= — the same real default,
mode="affine", matching every quantize call above. No
mismatch anywhere.
05_full_model_quantize.py, Part 5c, the code that builds
what gets written into config.json:
plan_config[name] = {"bits": bits, "group_size": group_size}Four places in the script built this dict this way — never including
"mode". This is pure bookkeeping code, running
after Part 5b’s real correction has already produced its real,
correctly-affine-quantized numbers. It doesn’t touch a single weight
value. It only decides what text gets written into a JSON file
describing weights that were already finished and saved.
The correction math, the actual rounding decisions, the actual packed
weight values — all of it was computed with consistent, correct,
default-affine quantize/dequantize the entire time. The bug was a
metadata omission, downstream of the real work, and it has now been
fixed both in the already-built model’s config.json
(patched directly, real backup kept) and in the script itself (all four
real locations now explicitly write "mode": "affine").
embed_tokens, was it YAQA-corrected, and what
precision is it at?What it is: the model’s input token embedding table — a lookup table mapping each vocabulary token to its starting vector, not a Linear projection layer. YAQA’s two-sided Hessian correction is built for the kind of weight matrix that Linear layers have; it was never designed or applied to embedding tables.
Was it YAQA-corrected? No. Checked directly against the real plan file this whole production run used:
real plan entries matching "embed_tokens": []
Zero matches. embed_tokens was never a target in the
real Pareto/MILP sensitivity plan at all — for either model, including
the reference.
Real, current answer, verified directly against both models’
actual on-disk tensor headers (checked 2026-09-17, not assumed from
config.json alone): embed_tokens is full-precision BF16 in
both builds, identical.
REFERENCE embed_tokens.weight: {'dtype': 'BF16', 'shape': [248320, 5120]}
YAQA embed_tokens.weight: {'dtype': 'BF16', 'shape': [248320, 5120]}
Neither has a matching .scales/.biases
sidecar tensor — the real, definitive signal (in this project’s own
format) that a tensor was never quantized at all, as opposed to
quantized at some specific bit-width. There is currently no difference
between the two builds for this layer.
This wasn’t always true, and the real history matters
here. embed_tokens used to fall through to this
script’s generic “outside the plan” fallback branch — any leaf module
the real plan doesn’t assign — and get quantized there like any other
out-of-plan tensor, at whatever --fallback-bits said
(4-bit, packed as U32, 636MB, on the run this was first measured
against). This was found to be a real, genuine bug, fixed the same day
(2026-09-07, same-day follow-up, confirmed directly from this script’s
own real code comment at
05_full_model_quantize.py:1699-1714): the reference build’s
own real embed_tokens.weight was checked directly against
its safetensors header and found untouched BF16, 2.54GB — meaning
quantizing it in the YAQA build was never something the real plan asked
for, just an accidental side effect of embed_tokens (an
nn.Embedding leaf) falling into the same catch-all branch
built for ordinary out-of-plan Linear layers. The
fix was structural, not a flag: the script now checks
isinstance(leaf, nn.Embedding) before that fallback branch
and routes any embedding leaf around quantization entirely, matching the
reference exactly. This is why the real, current, on-disk state
(BF16/BF16, verified above) differs from what an earlier version of this
very document once reported.
Useful background even though this specific project ended up not
quantizing embed_tokens at all (below):
embed_tokens is a lookup table. A forward pass reads
exactly one row per input token (a gather), never a full matrix multiply
across the whole table the way a Linear layer’s weights get used. That
structural difference is why quantization error here behaves differently
from a Linear layer’s — it only ever affects the specific tokens
actually used in a given input, not something that compounds across
every position the way error in an attention/MLP weight can. This is the
general, field-standard reasoning (seen across GPTQ, AWQ,
GGUF/llama.cpp, and other common quantization toolkits) for why
embedding tables are usually treated as lower-risk to quantize
than Linear layers, and it’s also the same reasoning this project
already applies to lm_head (the output-side equivalent of
an embedding table) — see
GPTQ_LM_Head_Safety_Gate_Plan.html — which gets its own
dedicated, more careful correction path specifically because errors
there are amplified rather than merely local. The one real caveat, for
any future model where this question comes up again: if a model
ties its embedding and output-projection weights (shares the
same matrix for both), errors in that shared table are no longer purely
local — they show up twice, at input and at the final logits.
Two genuinely separate real bugs got fixed by this investigation — both listed here with the exact real chronology, so it’s unambiguous which script(s) changed and in what order.
1. The "mode": "affine" metadata bug (Question
2) — fully fixed: 1. The already-built production model’s
config.json — patched directly (real backup:
config.json.bak-before-mode-fix),
"mode": "affine" added to all 364 real per-tensor entries
plus the top-level field. Verified the patched model still lazy-loads
correctly via mlx_lm.utils.load(). 2.
05_full_model_quantize.py — all four real locations that
build a per-tensor quantization config dict now explicitly include
"mode": "affine", plus the top-level
config["quantization"]["mode"] is now set too, matching the
reference build’s own real structure. 3. Full regression suite still
green after both changes.
2. The embed_tokens quantization bug (Question
3) — also a real bug, fixed in two real steps, not one: 1.
First (still 2026-09-07, later the same day):
--fallback-bits 4 was passed on the next full rebuild, so
embed_tokens would at least match the reference build’s own
real bit-width for other out-of-plan tensors — an attempt at parity, not
yet the real fix. 2. That still didn’t match, because the reference
build doesn’t quantize embed_tokens at all
— confirmed by reading its real safetensors header directly (untouched
BF16, 2.54GB). The actual, structural fix:
05_full_model_quantize.py now checks
isinstance(leaf, nn.Embedding) before the generic
out-of-plan fallback branch, and routes any embedding leaf around
quantization entirely. Verified today (2026-09-17) against both models’
real, current tensor headers: embed_tokens.weight is BF16
in both, with no .scales/.biases sidecars in
either — genuinely matching now, not just close.
What chasing this down led to, and why it matters
here: investigating whether the embed_tokens
bit-width explained the production model’s measured slowdown led to
discovering it didn’t — the real, dominant cause turned out to
be a second, unrelated, larger bug: this project’s own
.scales/.biases metadata tensors (727 of them,
confirmed by reading every one’s real file header on disk) were being
saved as float32 instead of bfloat16,
silently doubling the per-group metadata bytes MLX’s quantized-matmul
kernel has to read on every single forward pass — a real, measurable
speed cost, and a completely separate bug from anything else in this
document. embed_tokens’s bit-width turned out to be a real
but secondary factor once that larger bug was fixed. Full investigation
and fix: RCA_scales_f32_bug.html.
© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.