2026-09-06
Not to be confused with: a second, unrelated BF16/F32 bug found 2026-09-07, in a completely different part of the pipeline — this document is about the curvature computation (the Hessian used during correction) running in too little precision. The later bug is about the scales/biases metadata (saved after correction is already done) being written in too much precision, which cost real inference speed. Different stage, different symptom, different fix. See RCA_scales_f32_bug.html for that one.
Author: Hakim Ghelab Organization:
VegaLaboratories LTD Date: 2026-09-06
Status: Root cause confirmed. Fix implemented and
tested. Full-model reprocessing complete as of
2026-09-07 — the real 363-tensor production run finished, and
every tensor previously processed under the buggy bf16 curvature was
auto-invalidated by CURVATURE_VERSION and recomputed
correctly. That same production build went on to surface two further,
unrelated real bugs (missing mode: "affine"; scales/biases
saved as float32) — see RCA_scales_f32_bug.html for those,
and CHANGELOG.md’s 2026-09-07 entries for the full real
chronology.
This document explains, in plain language, what went wrong, how we found it, and what we fixed. It assumes no prior knowledge of the math. Every technical term is defined the first time it appears. If you only read one document from tonight’s investigation, read this one.
H_I (input curvature) and H_O (output
curvature) — built by running real example text through the model and
recording how the tensor’s output responds. High curvature in a
direction means “getting this direction wrong hurts the model a lot”;
near-zero curvature means “getting this direction wrong barely
matters.”Out of 363 tensors processed by YAQA, most tensors get 50-90% less
error than naive quantization. But 14 tensors ended up
worse than naive quantization — meaning this project’s smarter
method actively made those specific tensors worse than the simplest
possible approach would have. One tensor in particular
(layers.14.linear_attn.in_proj_qkv) came out ~300
times worse than naive on the raw distance metric, and its
overall quality score was ~18x worse than the naive
baseline (a “-1788%” reduction — a negative number means YAQA
moved the wrong way).
Ruling things out, in order, each backed by a real, direct test:
We directly measured the eigenvalues of the curvature matrices for the broken tensor. They should all be zero or positive (see PSD in the glossary). Instead:
| Precision used to build curvature | Worst (most negative) eigenvalue | Real-world result |
|---|---|---|
| BF16 (what production was actually doing) | -5.6 and -7.5 | Broken: quality score 18x worse than naive |
| FP32 (32-bit — never actually tested until tonight) | -0.0003 and -0.0008 | Fixed: matches the 64-bit reference |
| Float64 (64-bit, the most precise option) | -0.0000000000016 (essentially zero) | Fixed (reference/gold standard) |
The critical realization: production was building curvature using BF16 math, not FP32 as everyone (including this project’s own earlier documentation) had assumed. BF16 only has 2-3 accurate decimal digits — nowhere near enough for a curvature matrix this close to singular (see effective rank — this tensor’s curvature is effectively described by only ~2 real directions out of thousands). Plain FP32 (7 accurate digits) turns out to already be enough to fix it — its remaining tiny negative eigenvalues are exactly what you’d expect from ordinary FP32 rounding, not a sign of anything still broken.
We confirmed this isn’t just “the eigenvalues look better” by running the entire correction end-to-end and scoring the real result:
| Precision | Quality score vs. naive | Distance from original vs. naive |
|---|---|---|
| BF16 (broken) | -0.05x (worse than naive) | 3.84x (much worse than naive) |
| FP32 (fixed) | 0.0009x (~1000x better than naive) | 1.02x (essentially matches naive) |
| Float64 (reference) | 0.0009x (identical to FP32 to 4 decimal places) | 1.02x (identical) |
FP32 and float64 land on the same answer. We do not need 64-bit precision in production — we only needed to stop accidentally using 16-bit precision.
One line, in the one place curvature is actually built. Before:
xb = x[b] # 16-bit number
gb = grad_output[b] # 16-bit number
Gb = gb.T @ xb # curvature math done in 16-bit — too coarseAfter:
xb = x[b].astype(mx.float32) # promoted to 32-bit BEFORE the curvature math
gb = grad_output[b].astype(mx.float32)
Gb = gb.T @ xb # curvature math done in 32-bit — confirmed sufficientThe model itself keeps running in fast 16-bit mode throughout — only the separate curvature bookkeeping (which is small compared to the model itself) is promoted to 32-bit. This runs natively on the GPU with no slowdown of note (32-bit curvature collection took 95.7 seconds vs. 16-bit’s 116.8 seconds on the same real tensor — not slower at all in this measurement).
Every one of the 363 tensors went through the same buggy 16-bit curvature step. The 14 tensors that failed outright are simply the ones where the damage was large enough to visibly fail. The other ~349 tensors may have received a valid-looking but quietly-suboptimal correction — we have not yet checked this at scale, only confirmed it for one tensor pair (one broken, one that looked fine). Until we check more broadly, we do not know how many of the “fine-looking” tensors were actually only fine by luck.
A cache-version marker has been added so that once we decide to reprocess, the system can never silently reuse a result computed with the old, buggy 16-bit curvature — every old result is now tagged with which curvature version produced it, and a mismatch forces recomputation rather than silent reuse.
Confirming the fix on a “normal” (previously fine-looking) tensor, not just the one broken tensor — currently in progress.
Measuring the real, end-to-end effect on the model’s actual output (KL divergence — see glossary) — not yet run.
Deciding how many of the 363 tensors need to be reprocessed with the fix, and doing it — not yet started. Nothing in the live model cache has been touched.
lm_head question, raised 2026-09-06 (new,
separate from the YAQA fix above): lm_head is
corrected by a completely different code path — real GPTQ (see glossary
entry to add: GPTQ is a different, one-sided quantization-correction
method, reused from Apple’s own MLX library — used for
lm_head specifically because its full two-sided YAQA
correction is structurally impossible: its output-side curvature alone
would need 246.7GB, see PORT_LEDGER.md line 1600 for the
original, real, already-settled memory-driven decision to use GPTQ there
instead of YAQA). That GPTQ path’s own Hessian (H = x.T@x,
from Apple’s unmodified mlx_lm.quant.gptq.Catcher) is
accumulated in bf16 the same structural way the YAQA bug was — confirmed
by reading Apple’s actual source, not yet confirmed to actually matter
in practice for this specific tensor. GPTQ’s own regularization
(1e-2 * mean(diag(H)), i.e. ~1% of the mean diagonal) is
markedly weaker than YAQA’s (~100%), and GPTQ requires a full matrix
inverse of H rather than only a factorization — both
make it structurally more exposed to poor conditioning if
lm_head’s real Hessian turns out to be ill-conditioned, not
less. Not yet measured either way.
Exact command to check this (not yet run):
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
"$PYBIN" -u research_hadamard_blowup/lm_head_precision_check.pyThis is a small, isolated diagnostic
(research_hadamard_blowup/lm_head_precision_check.py) —
collects lm_head’s real Hessian under both bf16 and fp32
accumulation (forward-only, no backward pass needed for GPTQ) and
reports effective rank / eigenvalue conditioning for both, without
running the actual GPTQ correction. Does not touch production code or
the live model cache.
The original run’s output folder is
Qwen3.8-27B-heretic-ara-YAQA-5bpw (its resume folder — the
working directory the pipeline caches partial progress in — is
automatically named
Qwen3.8-27B-heretic-ara-YAQA-5bpw.yaqa_resume by the
orchestrator script itself; nothing in this project’s commands has to
specify that name manually). That original folder and its resume folder
are left completely untouched — the reprocessing run
below targets a brand-new output name so the old (BF16-curvature)
results stay available for a full side-by-side comparison against the
new (FP32-curvature) results, tensor by tensor.
Confirmed against the actual orchestrator script
(run_full_yaqa.sh) before writing this down — the
script derives its own resume folder as
<output_dir>.yaqa_resume automatically; it does not
take a separate --resume-dir override safely (passing one
manually would create a mismatch between where the script logs its
progress and where the real tensor data actually goes). The correct way
to target a new location is simply to give it a new output folder
name:
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
./run_full_yaqa.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32 \
--n-calibration 4 --incoherence noneThis will: - Automatically create and use
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume
as its working folder — brand new, starts empty. - Process all 363
tensors fresh, using the fixed FP32 curvature code. - Leave the original
Qwen3.8-27B-heretic-ara-YAQA-5bpw output and its
.yaqa_resume folder completely untouched, for comparison. -
Not been run as of this document — this is the exact command, not a
report of something already done.
YAQA’s smarter quantization needs to know which parts of each tensor matter — that information is called “curvature.” The code that builds curvature was accidentally running its math in the model’s fast, low-precision 16-bit format instead of a more precise 32-bit format, and for a handful of tensors whose curvature is extremely concentrated in just a couple of real directions, 16-bit precision wasn’t enough — it produced mathematically impossible results that corrupted the correction. Switching that one calculation to 32-bit (not needing the much slower 64-bit) fixes it completely, confirmed by matching the most precise reference available to four decimal places, at no meaningful speed cost.
© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.