← Back to index

Root Cause Analysis · Qwen3.8-27B-heretic-ara YAQA-5bpw-v2

Why the YAQA build ran 20–55% slower

Full forensic trace of the decode-speed regression against the Stratified reference build: what broke, exactly where in the code, what it did to every tensor in the model, and what it did not touch.

Correction math verified correct — 26h correction never re-run · 2026-09-07
Rebuilt, live, and independently re-verified — AR/D2 now at real parity
D1/D3 gap remains — real quality benchmark still open (§08)

01TL;DR

What broke

Every quantized layer's scale/bias metadata was saved as 32-bit floats instead of 16-bit — double the size for no reason.

Where

5 save points in 05_full_model_quantize.py and 06_lm_head_gptq.py — the correction step, not the save-config step.

Was the correction wrong?

No. 0 mismatches in bits / group_size / mode across all 1912 tensors. Only the storage width of scales/biases was affected.

Do we redo the 26h correction?

No. Already-computed values are correct — they just need re-saving at the right size. Minutes, not days.

02The bug, explained plainly

No jargon version first, then the exact mechanism.

Every compressed (quantized) weight in the model comes with a small "decoder key" — for every group of 64 numbers, a scale and a bias that tell the chip how to turn the tiny compressed integers back into real weight values. That decoder key normally takes up 16 bits per number (the same as the rest of the model). In this build, it was accidentally saved at 32 bits per number — twice the size, everywhere, in every layer.

The chip has to read that decoder key from memory constantly while generating each token. Doubling its size doubles that memory traffic on every single layer, for every single token — which is exactly why the slowdown showed up uniformly across AR, D1, D2, and D3, instead of in just one spot.

Broken (YAQA build)
scales / biases → float32
×2 bytes/elem
Correct (reference & fixed build)
scales / biases → bfloat16

Why it happened: the correction math (the real, expensive Hessian-weighted rounding decisions) deliberately runs in 32-bit precision — that's correct and necessary for numerically stable results. But nothing ever shrank the resulting scale/bias arrays back down to 16-bit before writing them into the final model file. The reference build runs its whole pipeline in 16-bit from the start, so it never had this problem — there was nothing to shrink.

The packed weight codes themselves (the actual compressed integers) are a fixed-width integer format regardless of this bug — they were never affected, in this build or the reference.

03Where it lives in the code

Five save points, all with the same shape of mistake: a float32-precision correction result gets assigned to a layer's scales/biases with no downcast before serialization.

Line 1129–114005_full_model_quantize.py
--assemble-from-resume path
The path your last two real builds actually used. Loads payload["scales"] / payload["biases"] straight from the fp32 resume cache and assigns them with zero cast. This is the one that produced the model you benchmarked.
Line 101005_full_model_quantize.py
live correction, unrotated
Had a cast — q_layer.set_dtype(W_full.dtype) — but W_full.dtype is the correction dtype (float32), not the deploy dtype. The cast fired, just to the wrong target: a no-op.
Line 985–99105_full_model_quantize.py
live correction, rotated (Hadamard)
No cast attempted at all. Not on the production path (--incoherence none is the real default) but fixed for correctness.
Line 1264 / 127205_full_model_quantize.py
out-of-plan fallback (to_quantized)
Defensive fix — would only manifest if the source leaf's weight were fp32 at that point. On the actual build inspected, embed_tokens's scales/biases were already bf16 here, so this path did not contribute to the measured regression; fixed anyway to close the gap for any future run where it could.
Line 419–43906_lm_head_gptq.py
W = orig_module.weight.astype(mx.float32) for correction precision, then quantized and saved with no downcast. Confirmed on disk: lm_head scales+biases were 79.5 MB (float32) vs 39.7 MB (bfloat16) in the reference — exactly 2x.
FIX

All five now cast explicitly to bfloat16 right before the array is assigned to the layer / written to disk. Full regression suite (test_regression.py) re-run after the change: all tests still pass.

04Tensor forensics — every mismatched tensor, measured on disk

Read directly from each model's real safetensors headers (dtype, shape, byte size) and each model's real config.json quantization dict — no estimates. Rows grouped by tensor-name pattern (63–64 near-identical decoder layers collapsed into one row); counts and totals are exact sums, not samples.

Tensors compared
1912 / 1910
YAQA / reference
bits / group_size / mode mismatches
0
quantization identity is correct everywhere
dtype mismatches (scales+biases)
727
every one is F32 in YAQA, BF16 in reference
Excess bytes from the bug
1.54 GB
reclaimed by the fix, weights unchanged
Tensor patternCount YAQA dtypeRef dtype YAQA bytesRef bytesΔ bytes
embed_tokens.weight1U32BF16635,699,2002,542,796,800−1,907,097,600
layers.N.mlp.down_proj.biases63F32BF16350,945,280175,472,640+175,472,640
layers.N.mlp.down_proj.scales63F32BF16350,945,280175,472,640+175,472,640
layers.N.mlp.gate_proj.biases63F32BF16350,945,280175,472,640+175,472,640
layers.N.mlp.gate_proj.scales63F32BF16350,945,280175,472,640+175,472,640
layers.N.mlp.up_proj.biases63F32BF16350,945,280175,472,640+175,472,640
layers.N.mlp.up_proj.scales63F32BF16350,945,280175,472,640+175,472,640
layers.N.linear_attn.in_proj_qkv.biases44F32BF16144,179,20072,089,600+72,089,600
layers.N.linear_attn.in_proj_qkv.scales44F32BF16144,179,20072,089,600+72,089,600
layers.N.linear_attn.in_proj_z.biases44F32BF1686,507,52043,253,760+43,253,760
layers.N.linear_attn.in_proj_z.scales44F32BF1686,507,52043,253,760+43,253,760
layers.N.linear_attn.out_proj.biases42F32BF1682,575,36041,287,680+41,287,680
layers.N.linear_attn.out_proj.scales42F32BF1682,575,36041,287,680+41,287,680
lm_head.biases1F32BF1679,462,40039,731,200+39,731,200
lm_head.scales1F32BF1679,462,40039,731,200+39,731,200
embed_tokens.biases (no ref equiv.)1BF16—39,731,2000+39,731,200
embed_tokens.scales (no ref equiv.)1BF16—39,731,2000+39,731,200
layers.N.self_attn.q_proj.biases16F32BF1662,914,56031,457,280+31,457,280
layers.N.self_attn.q_proj.scales16F32BF1662,914,56031,457,280+31,457,280
layers.N.self_attn.o_proj.biases14F32BF1627,525,12013,762,560+13,762,560
layers.N.self_attn.o_proj.scales14F32BF1627,525,12013,762,560+13,762,560
layers.N.self_attn.k_proj.biases7F32BF162,293,7601,146,880+1,146,880
layers.N.self_attn.k_proj.scales7F32BF162,293,7601,146,880+1,146,880
layers.N.self_attn.v_proj.biases6F32BF161,966,080983,040+983,040
layers.N.self_attn.v_proj.scales6F32BF161,966,080983,040+983,040

All 25 pattern-rows shown — nothing omitted. Every row with a positive Δ and F32→BF16 tags is the same bug, repeated once per layer. embed_tokens.weight and the two embed_tokens.biases/scales rows are a separate, non-bug story — see below.

Total bytes as F32 — YAQA
3.08 GB
reference: 0 bytes
Total bytes as BF16 — YAQA
2.96 GB
reference: 6.97 GB
Total bytes as U32 (packed codes)
14.78 GB
reference: 14.15 GB — embed_tokens delta only

05What happened with embed_tokens — and why it's not the cause

What it is: the input token lookup table — converts a token ID into its starting vector. It's a lookup table, not a weight matrix a language-model layer multiplies through, so YAQA's Hessian-based correction was never designed for it and never targets it. It isn't in the quantization plan for either model.

What each build does with it: the reference build leaves it uncompressed — full bfloat16, 2.54 GB. This build compresses it with plain 4-bit quantization to save space — 636 MB. That's a real, deliberate difference in size, but a lookup-table gather is cheap either way and it is not what's driving the AR/D1/D2/D3 regression — the regression is uniform across every matmul-heavy layer, which points at the scales/biases bug, not at one lookup table.

CORRECTION

An earlier pass through this investigation guessed the reference build used a different --fallback-bits default for embed_tokens. That guess was wrong and is superseded here: the reference doesn't quantize it at all.

06Measured impact (pre-fix)

Real mtplx tune numbers, decode tok/s, on the broken (F32 scales/biases) build vs. the Stratified reference. Superseded by the real post-fix re-benchmark in §08 — kept here as the "before" baseline.

AR
17.91
reference
14.15
−21.0%
D1
36.35
reference
21.53
−40.8%
D2
47.93
reference
28.46
−40.6%
D3
54.08
reference
24.51
−54.7%
Stratified (reference) YAQA, broken build (this RCA)
Peak memory, broken build
27.85 GB
reference: 23.77 GB — +17.2%
Output validation
2/2 pass
every depth — output was always correct, just slow

07Getting the fix into a model you can benchmark

Neither option requires re-running the 26-hour correction — the corrected values on disk were already right, they just needed to be re-saved at the right width. Option B, below, is the one that was actually run — this is the exact, real command executed for the model currently live at Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.

ACTUALLY RUN

Option B — full rebuild with the fixed script

Re-ran the same --assemble-from-resume command as the original broken build, now producing bf16 scales/biases and full-precision embed_tokens from the fixed code paths (both fixes landed in the same script before this was run — see §05 below for embed_tokens). --output points at the same path as before; the prior broken build was moved aside first (.OLD-before-fallback4-fix) so nothing was overwritten blind.

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port

/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 -u 05_full_model_quantize.py \
  --assemble-from-resume \
  --resume-dir "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume" \
  --output "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32" \
  --n-calibration 4 \
  --incoherence none \
  --fallback-bits 4

Option A — fast patch, no recompute (not used, kept for reference)

Available if you ever need to patch an already-built model without a rebuild: loads it, casts every .scales/.biases tensor float32→bfloat16 and swaps embed_tokens for the real unquantized source weight, writes to a new directory, source untouched. Not the path used for the current live model.

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port

/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 09_patch_scales_bf16.py \
  --source "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32" \
  --output "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-bf16fix"
STATUS

Rebuilt, and independently verified. Full regression suite green before and after. See §08 for what was checked on the actual rebuilt model, not just the code.

08Post-rebuild verification — real checks against the actual live model

Everything below was run against the model that's actually live on disk right now, after the Option B rebuild above — not against code, not against the resume cache alone.

Structural re-check (tensor forensics, re-run): 0 / 1910 tensors differ from the reference in dtype, bits, group_size, mode, or byte size. BF16 total bytes and U32 total bytes both match the reference exactly, to the byte. The earlier F32 bug and the embed_tokens mismatch are both fully gone from the live model, not just from the code.

Weight-provenance check (chain of custody): pulled the same real tensor four independent ways — the original unquantized source weight, a plain naive quantization of it (no correction), the resume cache's real correction output, and what's actually sitting in the live model's safetensors file. Live matched the resume cache to bf16-rounding precision (~0.1–0.6%) and differed substantially from naive (1.9–4.5% absolute) — proof the real correction ran and its exact output shipped, not naive quantization standing in for it.

Correction-quality manifest analysis: 362/363 tensors used real YAQA correction (the 1 exception, lm_head, is the known, expected naive_fallback). Real Hessian-weighted error reduction vs. naive ranged 97.49%–100% across every corrected tensor, median 99.67%, zero tensors where correction was worse than naive. Mild, real decline in correction quality in later layers (self_attn.k_proj, linear_attn.out_proj, layers ~37–58) — measurable, not catastrophic, no cliff near the output layers.

Real mtplx tune re-benchmark, after the rebuild above, vs. before:

ModeReferenceYAQA (pre-fix, broken)YAQA (post-scales-fix)Δ scales-fix alone
(vs. Reference)
YAQA (post-MTP-fix, final)Δ final
(vs. Reference)
Δ total journey
(vs. broken)
AR17.91414.15317.889−0.14%17.699−1.20%+25.05%
D136.34521.53229.007−20.19%39.140+7.69%+81.78%
D247.93128.45847.422−1.06%51.681+7.82%+81.60%
D354.07624.50943.300−19.93%48.233−10.81%+96.80%
Peak memory (GB)23.76527.84823.765+0.00%23.765+0.00%−14.66%

AR and D2 now at real parity with the reference. D1 and D3 still down ~20%, with real acceptance-rate degradation at deeper speculative depths (D3 depth-3 acceptance: reference 94.74% vs. YAQA 80.60%). The manifest analysis above found no localized failure to explain this — current working hypothesis is that YAQA's correction and the reference build's own correction algorithm land on different, individually-small residual errors that compound differently through depth. Not yet independently confirmed by an actual output-quality benchmark — that's the open next step.

OPEN

Still pending: a real behavioral/quality benchmark (not just tok/s) against this rebuilt model, to determine whether the D1/D3 gap reflects a real quality trade-off worth keeping (as opposed to tok/s alone).

UPDATE

The working hypothesis above turned out to be wrong — the real cause was found and mostly fixed. It wasn't residual-error compounding: the YAQA trunk correction here never touched the MTP (speculative-decode draft head) sidecar at all, which was still plain mx.quantize, identical to what the naive reference does — a real trunk/MTP mismatch, not a diffuse one. Correcting the MTP sidecar the same way (full writeup: MTP_YAQA_CORRECTION_PLAN.html) is two separate fixes, each independently tested and re-measured: the scales/bias fix (§07–08, this page) landed AR/D2 at reference parity but left D1/D3 still down; the MTP fix (linked page) then closed the rest. At D1 and D2, the final post-MTP-fix speed doesn't just approach the reference — it beats both the reference and the scales-fix-only state outright: D1 39.140 tok/s exceeds both the scales-fix-only 29.007 and the reference 36.345 (+7.7% above reference); D2 51.681 tok/s exceeds both the scales-fix-only 47.422 and the reference 47.931 (+7.8% above reference). Measured against the original broken build (before either fix), the total journey is +81.78% at D1 and +81.60% at D2. D3 improved substantially (43.300 → 48.233, and +96.80% vs. the original broken build) but, unlike D1/D2, still sits −10.81% behind the reference, with D3 depth-3 acceptance up 80.60% → 92.00%. D3 is real, measured, and still open — not yet fully closed.