Root Cause Analysis · Qwen3.8-27B-heretic-ara YAQA-5bpw-v2
Full forensic trace of the decode-speed regression against the Stratified reference build: what broke, exactly where in the code, what it did to every tensor in the model, and what it did not touch.
What broke
Every quantized layer's scale/bias metadata was saved as 32-bit floats instead of 16-bit — double the size for no reason.
Where
5 save points in 05_full_model_quantize.py and 06_lm_head_gptq.py — the correction step, not the save-config step.
Was the correction wrong?
No. 0 mismatches in bits / group_size / mode across all 1912 tensors. Only the storage width of scales/biases was affected.
Do we redo the 26h correction?
No. Already-computed values are correct — they just need re-saving at the right size. Minutes, not days.
No jargon version first, then the exact mechanism.
Every compressed (quantized) weight in the model comes with a small "decoder key" — for every group of 64 numbers, a scale and a bias that tell the chip how to turn the tiny compressed integers back into real weight values. That decoder key normally takes up 16 bits per number (the same as the rest of the model). In this build, it was accidentally saved at 32 bits per number — twice the size, everywhere, in every layer.
The chip has to read that decoder key from memory constantly while generating each token. Doubling its size doubles that memory traffic on every single layer, for every single token — which is exactly why the slowdown showed up uniformly across AR, D1, D2, and D3, instead of in just one spot.
Why it happened: the correction math (the real, expensive Hessian-weighted rounding decisions) deliberately runs in 32-bit precision — that's correct and necessary for numerically stable results. But nothing ever shrank the resulting scale/bias arrays back down to 16-bit before writing them into the final model file. The reference build runs its whole pipeline in 16-bit from the start, so it never had this problem — there was nothing to shrink.
The packed weight codes themselves (the actual compressed integers) are a fixed-width integer format regardless of this bug — they were never affected, in this build or the reference.
Five save points, all with the same shape of mistake: a float32-precision correction result gets assigned to a layer's scales/biases with no downcast before serialization.
payload["scales"] / payload["biases"] straight from the fp32 resume cache and assigns them with zero cast. This is the one that produced the model you benchmarked.q_layer.set_dtype(W_full.dtype) — but W_full.dtype is the correction dtype (float32), not the deploy dtype. The cast fired, just to the wrong target: a no-op.--incoherence none is the real default) but fixed for correctness.embed_tokens's scales/biases were already bf16 here, so this path did not contribute to the measured regression; fixed anyway to close the gap for any future run where it could.W = orig_module.weight.astype(mx.float32) for correction precision, then quantized and saved with no downcast. Confirmed on disk: lm_head scales+biases were 79.5 MB (float32) vs 39.7 MB (bfloat16) in the reference — exactly 2x.All five now cast explicitly to bfloat16 right before the array is assigned to the layer / written to disk. Full regression suite (test_regression.py) re-run after the change: all tests still pass.
Read directly from each model's real safetensors headers (dtype, shape, byte size) and each model's real config.json quantization dict — no estimates. Rows grouped by tensor-name pattern (63–64 near-identical decoder layers collapsed into one row); counts and totals are exact sums, not samples.
| Tensor pattern | Count | YAQA dtype | Ref dtype | YAQA bytes | Ref bytes | Δ bytes |
|---|---|---|---|---|---|---|
| embed_tokens.weight | 1 | U32 | BF16 | 635,699,200 | 2,542,796,800 | −1,907,097,600 |
| layers.N.mlp.down_proj.biases | 63 | F32 | BF16 | 350,945,280 | 175,472,640 | +175,472,640 |
| layers.N.mlp.down_proj.scales | 63 | F32 | BF16 | 350,945,280 | 175,472,640 | +175,472,640 |
| layers.N.mlp.gate_proj.biases | 63 | F32 | BF16 | 350,945,280 | 175,472,640 | +175,472,640 |
| layers.N.mlp.gate_proj.scales | 63 | F32 | BF16 | 350,945,280 | 175,472,640 | +175,472,640 |
| layers.N.mlp.up_proj.biases | 63 | F32 | BF16 | 350,945,280 | 175,472,640 | +175,472,640 |
| layers.N.mlp.up_proj.scales | 63 | F32 | BF16 | 350,945,280 | 175,472,640 | +175,472,640 |
| layers.N.linear_attn.in_proj_qkv.biases | 44 | F32 | BF16 | 144,179,200 | 72,089,600 | +72,089,600 |
| layers.N.linear_attn.in_proj_qkv.scales | 44 | F32 | BF16 | 144,179,200 | 72,089,600 | +72,089,600 |
| layers.N.linear_attn.in_proj_z.biases | 44 | F32 | BF16 | 86,507,520 | 43,253,760 | +43,253,760 |
| layers.N.linear_attn.in_proj_z.scales | 44 | F32 | BF16 | 86,507,520 | 43,253,760 | +43,253,760 |
| layers.N.linear_attn.out_proj.biases | 42 | F32 | BF16 | 82,575,360 | 41,287,680 | +41,287,680 |
| layers.N.linear_attn.out_proj.scales | 42 | F32 | BF16 | 82,575,360 | 41,287,680 | +41,287,680 |
| lm_head.biases | 1 | F32 | BF16 | 79,462,400 | 39,731,200 | +39,731,200 |
| lm_head.scales | 1 | F32 | BF16 | 79,462,400 | 39,731,200 | +39,731,200 |
| embed_tokens.biases (no ref equiv.) | 1 | BF16 | — | 39,731,200 | 0 | +39,731,200 |
| embed_tokens.scales (no ref equiv.) | 1 | BF16 | — | 39,731,200 | 0 | +39,731,200 |
| layers.N.self_attn.q_proj.biases | 16 | F32 | BF16 | 62,914,560 | 31,457,280 | +31,457,280 |
| layers.N.self_attn.q_proj.scales | 16 | F32 | BF16 | 62,914,560 | 31,457,280 | +31,457,280 |
| layers.N.self_attn.o_proj.biases | 14 | F32 | BF16 | 27,525,120 | 13,762,560 | +13,762,560 |
| layers.N.self_attn.o_proj.scales | 14 | F32 | BF16 | 27,525,120 | 13,762,560 | +13,762,560 |
| layers.N.self_attn.k_proj.biases | 7 | F32 | BF16 | 2,293,760 | 1,146,880 | +1,146,880 |
| layers.N.self_attn.k_proj.scales | 7 | F32 | BF16 | 2,293,760 | 1,146,880 | +1,146,880 |
| layers.N.self_attn.v_proj.biases | 6 | F32 | BF16 | 1,966,080 | 983,040 | +983,040 |
| layers.N.self_attn.v_proj.scales | 6 | F32 | BF16 | 1,966,080 | 983,040 | +983,040 |
All 25 pattern-rows shown — nothing omitted. Every row with a positive Δ and F32→BF16 tags is the same bug, repeated once per layer. embed_tokens.weight and the two embed_tokens.biases/scales rows are a separate, non-bug story — see below.
What it is: the input token lookup table — converts a token ID into its starting vector. It's a lookup table, not a weight matrix a language-model layer multiplies through, so YAQA's Hessian-based correction was never designed for it and never targets it. It isn't in the quantization plan for either model.
What each build does with it: the reference build leaves it uncompressed — full bfloat16, 2.54 GB. This build compresses it with plain 4-bit quantization to save space — 636 MB. That's a real, deliberate difference in size, but a lookup-table gather is cheap either way and it is not what's driving the AR/D1/D2/D3 regression — the regression is uniform across every matmul-heavy layer, which points at the scales/biases bug, not at one lookup table.
An earlier pass through this investigation guessed the reference build used a different --fallback-bits default for embed_tokens. That guess was wrong and is superseded here: the reference doesn't quantize it at all.
Real mtplx tune numbers, decode tok/s, on the broken (F32 scales/biases) build vs. the Stratified reference. Superseded by the real post-fix re-benchmark in §08 — kept here as the "before" baseline.
Neither option requires re-running the 26-hour correction — the corrected values on disk were already right, they just needed to be re-saved at the right width. Option B, below, is the one that was actually run — this is the exact, real command executed for the model currently live at Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.
Re-ran the same --assemble-from-resume command as the original broken build, now producing bf16 scales/biases and full-precision embed_tokens from the fixed code paths (both fixes landed in the same script before this was run — see §05 below for embed_tokens). --output points at the same path as before; the prior broken build was moved aside first (.OLD-before-fallback4-fix) so nothing was overwritten blind.
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 -u 05_full_model_quantize.py \ --assemble-from-resume \ --resume-dir "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume" \ --output "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32" \ --n-calibration 4 \ --incoherence none \ --fallback-bits 4
Available if you ever need to patch an already-built model without a rebuild: loads it, casts every .scales/.biases tensor float32→bfloat16 and swaps embed_tokens for the real unquantized source weight, writes to a new directory, source untouched. Not the path used for the current live model.
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 09_patch_scales_bf16.py \ --source "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32" \ --output "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-bf16fix"
Rebuilt, and independently verified. Full regression suite green before and after. See §08 for what was checked on the actual rebuilt model, not just the code.
Everything below was run against the model that's actually live on disk right now, after the Option B rebuild above — not against code, not against the resume cache alone.
Structural re-check (tensor forensics, re-run): 0 / 1910 tensors differ from the reference in dtype, bits, group_size, mode, or byte size. BF16 total bytes and U32 total bytes both match the reference exactly, to the byte. The earlier F32 bug and the embed_tokens mismatch are both fully gone from the live model, not just from the code.
Weight-provenance check (chain of custody): pulled the same real tensor four independent ways — the original unquantized source weight, a plain naive quantization of it (no correction), the resume cache's real correction output, and what's actually sitting in the live model's safetensors file. Live matched the resume cache to bf16-rounding precision (~0.1–0.6%) and differed substantially from naive (1.9–4.5% absolute) — proof the real correction ran and its exact output shipped, not naive quantization standing in for it.
Correction-quality manifest analysis: 362/363 tensors used real YAQA correction (the 1 exception, lm_head, is the known, expected naive_fallback). Real Hessian-weighted error reduction vs. naive ranged 97.49%–100% across every corrected tensor, median 99.67%, zero tensors where correction was worse than naive. Mild, real decline in correction quality in later layers (self_attn.k_proj, linear_attn.out_proj, layers ~37–58) — measurable, not catastrophic, no cliff near the output layers.
Real mtplx tune re-benchmark, after the rebuild above, vs. before:
| Mode | Reference | YAQA (pre-fix, broken) | YAQA (post-scales-fix) | Δ scales-fix alone (vs. Reference) | YAQA (post-MTP-fix, final) | Δ final (vs. Reference) | Δ total journey (vs. broken) |
|---|---|---|---|---|---|---|---|
| AR | 17.914 | 14.153 | 17.889 | −0.14% | 17.699 | −1.20% | +25.05% |
| D1 | 36.345 | 21.532 | 29.007 | −20.19% | 39.140 | +7.69% | +81.78% |
| D2 | 47.931 | 28.458 | 47.422 | −1.06% | 51.681 | +7.82% | +81.60% |
| D3 | 54.076 | 24.509 | 43.300 | −19.93% | 48.233 | −10.81% | +96.80% |
| Peak memory (GB) | 23.765 | 27.848 | 23.765 | +0.00% | 23.765 | +0.00% | −14.66% |
AR and D2 now at real parity with the reference. D1 and D3 still down ~20%, with real acceptance-rate degradation at deeper speculative depths (D3 depth-3 acceptance: reference 94.74% vs. YAQA 80.60%). The manifest analysis above found no localized failure to explain this — current working hypothesis is that YAQA's correction and the reference build's own correction algorithm land on different, individually-small residual errors that compound differently through depth. Not yet independently confirmed by an actual output-quality benchmark — that's the open next step.
Still pending: a real behavioral/quality benchmark (not just tok/s) against this rebuilt model, to determine whether the D1/D3 gap reflects a real quality trade-off worth keeping (as opposed to tok/s alone).
The working hypothesis above turned out to be wrong — the real cause was found and mostly fixed. It wasn't residual-error compounding: the YAQA trunk correction here never touched the MTP (speculative-decode draft head) sidecar at all, which was still plain mx.quantize, identical to what the naive reference does — a real trunk/MTP mismatch, not a diffuse one. Correcting the MTP sidecar the same way (full writeup: MTP_YAQA_CORRECTION_PLAN.html) is two separate fixes, each independently tested and re-measured: the scales/bias fix (§07–08, this page) landed AR/D2 at reference parity but left D1/D3 still down; the MTP fix (linked page) then closed the rest. At D1 and D2, the final post-MTP-fix speed doesn't just approach the reference — it beats both the reference and the scales-fix-only state outright: D1 39.140 tok/s exceeds both the scales-fix-only 29.007 and the reference 36.345 (+7.7% above reference); D2 51.681 tok/s exceeds both the scales-fix-only 47.422 and the reference 47.931 (+7.8% above reference). Measured against the original broken build (before either fix), the total journey is +81.78% at D1 and +81.60% at D2. D3 improved substantially (43.300 → 48.233, and +96.80% vs. the original broken build) but, unlike D1/D2, still sits −10.81% behind the reference, with D3 depth-3 acceptance up 80.60% → 92.00%. D3 is real, measured, and still open — not yet fully closed.