← Back to index

The "Missing Affine Mode" Bug — Full Code Review, Zero Ambiguity

2026-09-07

Author: Hakim Ghelab, VegaLaboratories LTD

This document answers, with direct evidence and no hand-waving, the four real questions raised after the speed investigation: what 05_full_model_quantize.py actually is, whether the missing "mode": "affine" config field means the real YAQA correction math was ever wrong, whether embed_tokens is safe to quantize and what precision it’s actually at compared to the reference model, and what was actually fixed.

What “the reference model” means, throughout this document: this project builds two things side by side — (1) this YAQA port, the code this document is reviewing, and (2) a separately pre-existing, already-trusted build (HybridPareto5bpw-V3.3_Stratified, produced by this project’s own older, already-validated 05_build_final_model_V5*.py scripts) that YAQA’s real output gets compared against for speed and quality. “The reference model” / “the reference build” always means that second, pre-existing model — never a third-party or external model. It matters here specifically because it’s the only way to tell whether a difference in YAQA’s output (like embed_tokens’s bit-width, below) is a real bug or just two independent scripts making their own separate, both-legitimate default choice.


The verdict, up front


Question 1 — What is 05_full_model_quantize.py?

Its own file header, read directly, not paraphrased:

#!/usr/bin/env python3
"""
Author: Hakim Ghelab, VegaLaboratories LTD

YAQA port -- Step 5 of 5 (staged plan): full model.

This is Step 5 of the real, staged YAQA port plan — the final, full-model production script. It is not a side tool. It is where:

happen, in that order, for every real production run this whole session has been monitoring.


Question 2 — Did the missing “affine” mode mean the correction was wrong?

This is the question that actually matters, and it deserves the most rigorous answer, not a reassurance.

The real distinction: two completely separate parts of the script

Part 5b (the correction itself) computes real numbers — it runs the Hessian-weighted block-LDL sweep, decides how to round each weight, and produces real packed (weight, scales, biases) arrays via real calls to mx.quantize().

Part 5c (assembly and save), which runs after Part 5b is completely finished, takes those already-computed arrays and writes them to disk, alongside a config.json that describes what was done — including a "mode" field that’s supposed to say which quantization scheme was used.

The bug was only in what Part 5c wrote down — never in what Part 5b computed. Here is the real, exhaustive evidence for that claim.

Every real mx.quantize() call in the entire pipeline

Checked directly, every file, every call site — not a sample:

yaqa_core.py:172   w_q, scales, biases = mx.quantize(thing, bits=bits, group_size=group_size)
yaqa_core.py:246   wq_final, scales_final, biases_final = mx.quantize(hatW, bits=bits, group_size=group_size)
yaqa_core.py:416   naive_wq, naive_s, naive_b = mx.quantize(W, bits=bits, group_size=group_size)

05_full_model_quantize.py:907   naive_wq, naive_scales, naive_biases = mx.quantize(Wr_full, bits=bits, group_size=group_size)
05_full_model_quantize.py:940   naive_wq, naive_scales, naive_biases = mx.quantize(Wr_full, bits=bits, group_size=group_size)

06_lm_head_gptq.py:102   _, scales, biases = mx.quantize(Wl, bits=bits, group_size=group_size)
06_lm_head_gptq.py:113   wq, scales, biases = mx.quantize(Wc, bits=bits, group_size=group_size)
06_lm_head_gptq.py:422   naive_wq, naive_scales, naive_biases = mx.quantize(W, bits=bits, group_size=group_size)

Every single one passes only bits= and group_size=. None of them ever passes a mode= argument. That means every one of them used MLX’s own real default — which, checked directly from mlx.core.quantize’s actual signature (not assumed):

quantize(w, group_size=None, bits=None, mode: str = 'affine', ...)

mode="affine" is the default. Every real quantize call in this entire codebase — the one inside YAQA’s own correction sweep at yaqa_core.py:172 included — used it.

Every real mx.dequantize() call — checked the same way

This matters just as much: if quantize and dequantize ever used different modes anywhere, that would be a genuine correctness bug (comparing apples to oranges when measuring how good a correction is). Checked, exhaustively:

yaqa_core.py:173   hatWr[...] = mx.dequantize(w_q, scales, biases, bits=bits, group_size=group_size)
yaqa_core.py:417   naive_dq = mx.dequantize(naive_wq, naive_s, naive_b, bits=bits, group_size=group_size)

05_full_model_quantize.py:908   naive_hatWr = mx.dequantize(naive_wq, naive_scales, naive_biases, bits=bits, group_size=group_size)
05_full_model_quantize.py:941   naive_hatWr = mx.dequantize(naive_wq, naive_scales, naive_biases, bits=bits, group_size=group_size)

06_lm_head_gptq.py:115   ...mx.dequantize(wq, scales, biases, bits=bits, group_size=group_size)...
06_lm_head_gptq.py:155   corrected_W = mx.dequantize(wq, scales, biases, bits=bits, group_size=group_size)
06_lm_head_gptq.py:423   naive_W = mx.dequantize(naive_wq, naive_scales, naive_biases, bits=bits, group_size=group_size)

Same result: every dequantize call, no exceptions, uses only bits=/group_size= — the same real default, mode="affine", matching every quantize call above. No mismatch anywhere.

Where the real bug actually lived

05_full_model_quantize.py, Part 5c, the code that builds what gets written into config.json:

plan_config[name] = {"bits": bits, "group_size": group_size}

Four places in the script built this dict this way — never including "mode". This is pure bookkeeping code, running after Part 5b’s real correction has already produced its real, correctly-affine-quantized numbers. It doesn’t touch a single weight value. It only decides what text gets written into a JSON file describing weights that were already finished and saved.

The real conclusion

The correction math, the actual rounding decisions, the actual packed weight values — all of it was computed with consistent, correct, default-affine quantize/dequantize the entire time. The bug was a metadata omission, downstream of the real work, and it has now been fixed both in the already-built model’s config.json (patched directly, real backup kept) and in the script itself (all four real locations now explicitly write "mode": "affine").


Question 3 — What is embed_tokens, was it YAQA-corrected, and what precision is it at?

What it is: the model’s input token embedding table — a lookup table mapping each vocabulary token to its starting vector, not a Linear projection layer. YAQA’s two-sided Hessian correction is built for the kind of weight matrix that Linear layers have; it was never designed or applied to embedding tables.

Was it YAQA-corrected? No. Checked directly against the real plan file this whole production run used:

real plan entries matching "embed_tokens": []

Zero matches. embed_tokens was never a target in the real Pareto/MILP sensitivity plan at all — for either model, including the reference.

Real, current answer, verified directly against both models’ actual on-disk tensor headers (checked 2026-09-17, not assumed from config.json alone): embed_tokens is full-precision BF16 in both builds, identical.

REFERENCE embed_tokens.weight: {'dtype': 'BF16', 'shape': [248320, 5120]}
YAQA      embed_tokens.weight: {'dtype': 'BF16', 'shape': [248320, 5120]}

Neither has a matching .scales/.biases sidecar tensor — the real, definitive signal (in this project’s own format) that a tensor was never quantized at all, as opposed to quantized at some specific bit-width. There is currently no difference between the two builds for this layer.

This wasn’t always true, and the real history matters here. embed_tokens used to fall through to this script’s generic “outside the plan” fallback branch — any leaf module the real plan doesn’t assign — and get quantized there like any other out-of-plan tensor, at whatever --fallback-bits said (4-bit, packed as U32, 636MB, on the run this was first measured against). This was found to be a real, genuine bug, fixed the same day (2026-09-07, same-day follow-up, confirmed directly from this script’s own real code comment at 05_full_model_quantize.py:1699-1714): the reference build’s own real embed_tokens.weight was checked directly against its safetensors header and found untouched BF16, 2.54GB — meaning quantizing it in the YAQA build was never something the real plan asked for, just an accidental side effect of embed_tokens (an nn.Embedding leaf) falling into the same catch-all branch built for ordinary out-of-plan Linear layers. The fix was structural, not a flag: the script now checks isinstance(leaf, nn.Embedding) before that fallback branch and routes any embedding leaf around quantization entirely, matching the reference exactly. This is why the real, current, on-disk state (BF16/BF16, verified above) differs from what an earlier version of this very document once reported.

Is quantizing an embedding table actually safe, as a general question?

Useful background even though this specific project ended up not quantizing embed_tokens at all (below): embed_tokens is a lookup table. A forward pass reads exactly one row per input token (a gather), never a full matrix multiply across the whole table the way a Linear layer’s weights get used. That structural difference is why quantization error here behaves differently from a Linear layer’s — it only ever affects the specific tokens actually used in a given input, not something that compounds across every position the way error in an attention/MLP weight can. This is the general, field-standard reasoning (seen across GPTQ, AWQ, GGUF/llama.cpp, and other common quantization toolkits) for why embedding tables are usually treated as lower-risk to quantize than Linear layers, and it’s also the same reasoning this project already applies to lm_head (the output-side equivalent of an embedding table) — see GPTQ_LM_Head_Safety_Gate_Plan.html — which gets its own dedicated, more careful correction path specifically because errors there are amplified rather than merely local. The one real caveat, for any future model where this question comes up again: if a model ties its embedding and output-projection weights (shares the same matrix for both), errors in that shared table are no longer purely local — they show up twice, at input and at the final logits.


What was actually fixed, concretely

Two genuinely separate real bugs got fixed by this investigation — both listed here with the exact real chronology, so it’s unambiguous which script(s) changed and in what order.

1. The "mode": "affine" metadata bug (Question 2) — fully fixed: 1. The already-built production model’s config.json — patched directly (real backup: config.json.bak-before-mode-fix), "mode": "affine" added to all 364 real per-tensor entries plus the top-level field. Verified the patched model still lazy-loads correctly via mlx_lm.utils.load(). 2. 05_full_model_quantize.py — all four real locations that build a per-tensor quantization config dict now explicitly include "mode": "affine", plus the top-level config["quantization"]["mode"] is now set too, matching the reference build’s own real structure. 3. Full regression suite still green after both changes.

2. The embed_tokens quantization bug (Question 3) — also a real bug, fixed in two real steps, not one: 1. First (still 2026-09-07, later the same day): --fallback-bits 4 was passed on the next full rebuild, so embed_tokens would at least match the reference build’s own real bit-width for other out-of-plan tensors — an attempt at parity, not yet the real fix. 2. That still didn’t match, because the reference build doesn’t quantize embed_tokens at all — confirmed by reading its real safetensors header directly (untouched BF16, 2.54GB). The actual, structural fix: 05_full_model_quantize.py now checks isinstance(leaf, nn.Embedding) before the generic out-of-plan fallback branch, and routes any embedding leaf around quantization entirely. Verified today (2026-09-17) against both models’ real, current tensor headers: embed_tokens.weight is BF16 in both, with no .scales/.biases sidecars in either — genuinely matching now, not just close.

What chasing this down led to, and why it matters here: investigating whether the embed_tokens bit-width explained the production model’s measured slowdown led to discovering it didn’t — the real, dominant cause turned out to be a second, unrelated, larger bug: this project’s own .scales/.biases metadata tensors (727 of them, confirmed by reading every one’s real file header on disk) were being saved as float32 instead of bfloat16, silently doubling the per-group metadata bytes MLX’s quantized-matmul kernel has to read on every single forward pass — a real, measurable speed cost, and a completely separate bug from anything else in this document. embed_tokens’s bit-width turned out to be a real but secondary factor once that larger bug was fixed. Full investigation and fix: RCA_scales_f32_bug.html.


© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.