← Back to index

The BF16 Curvature Bug — Plain-English Explanation

2026-09-06

Not to be confused with: a second, unrelated BF16/F32 bug found 2026-09-07, in a completely different part of the pipeline — this document is about the curvature computation (the Hessian used during correction) running in too little precision. The later bug is about the scales/biases metadata (saved after correction is already done) being written in too much precision, which cost real inference speed. Different stage, different symptom, different fix. See RCA_scales_f32_bug.html for that one.

Author: Hakim Ghelab Organization: VegaLaboratories LTD Date: 2026-09-06 Status: Root cause confirmed. Fix implemented and tested. Full-model reprocessing complete as of 2026-09-07 — the real 363-tensor production run finished, and every tensor previously processed under the buggy bf16 curvature was auto-invalidated by CURVATURE_VERSION and recomputed correctly. That same production build went on to surface two further, unrelated real bugs (missing mode: "affine"; scales/biases saved as float32) — see RCA_scales_f32_bug.html for those, and CHANGELOG.md’s 2026-09-07 entries for the full real chronology.

This document explains, in plain language, what went wrong, how we found it, and what we fixed. It assumes no prior knowledge of the math. Every technical term is defined the first time it appears. If you only read one document from tonight’s investigation, read this one.


1. Glossary — every term used below, defined once


2. What went wrong — the visible symptom

Out of 363 tensors processed by YAQA, most tensors get 50-90% less error than naive quantization. But 14 tensors ended up worse than naive quantization — meaning this project’s smarter method actively made those specific tensors worse than the simplest possible approach would have. One tensor in particular (layers.14.linear_attn.in_proj_qkv) came out ~300 times worse than naive on the raw distance metric, and its overall quality score was ~18x worse than the naive baseline (a “-1788%” reduction — a negative number means YAQA moved the wrong way).

3. The investigation — three wrong turns before the real answer

Ruling things out, in order, each backed by a real, direct test:

  1. First suspicion: the missing Hadamard rotation (a preprocessing step some quantization methods use). Tested directly: adding rotation only partially helped, and a rigorous mathematical check (rotation cannot change a matrix’s eigenvalues) proved it could never have fully fixed this by itself. Ruled out.
  2. Second suspicion: not enough example text was used to build the curvature (only 4 examples per topic, i.e. 24 examples total). Tested directly: rebuilding curvature with 2x more example text (48 examples) gave nearly identical results. Ruled out — this is not a “not enough data” problem.
  3. Third suspicion: the curvature math itself is numerically broken (this investigation’s actual title going in) — specifically, that FP32 (32-bit) isn’t precise enough and the fix requires float64 (64-bit). This is where the real answer was found, but not the one we expected — see below.

4. The real root cause

We directly measured the eigenvalues of the curvature matrices for the broken tensor. They should all be zero or positive (see PSD in the glossary). Instead:

Precision used to build curvature Worst (most negative) eigenvalue Real-world result
BF16 (what production was actually doing) -5.6 and -7.5 Broken: quality score 18x worse than naive
FP32 (32-bit — never actually tested until tonight) -0.0003 and -0.0008 Fixed: matches the 64-bit reference
Float64 (64-bit, the most precise option) -0.0000000000016 (essentially zero) Fixed (reference/gold standard)

The critical realization: production was building curvature using BF16 math, not FP32 as everyone (including this project’s own earlier documentation) had assumed. BF16 only has 2-3 accurate decimal digits — nowhere near enough for a curvature matrix this close to singular (see effective rank — this tensor’s curvature is effectively described by only ~2 real directions out of thousands). Plain FP32 (7 accurate digits) turns out to already be enough to fix it — its remaining tiny negative eigenvalues are exactly what you’d expect from ordinary FP32 rounding, not a sign of anything still broken.

We confirmed this isn’t just “the eigenvalues look better” by running the entire correction end-to-end and scoring the real result:

Precision Quality score vs. naive Distance from original vs. naive
BF16 (broken) -0.05x (worse than naive) 3.84x (much worse than naive)
FP32 (fixed) 0.0009x (~1000x better than naive) 1.02x (essentially matches naive)
Float64 (reference) 0.0009x (identical to FP32 to 4 decimal places) 1.02x (identical)

FP32 and float64 land on the same answer. We do not need 64-bit precision in production — we only needed to stop accidentally using 16-bit precision.

5. The fix

One line, in the one place curvature is actually built. Before:

xb = x[b]                 # 16-bit number
gb = grad_output[b]       # 16-bit number
Gb = gb.T @ xb             # curvature math done in 16-bit — too coarse

After:

xb = x[b].astype(mx.float32)   # promoted to 32-bit BEFORE the curvature math
gb = grad_output[b].astype(mx.float32)
Gb = gb.T @ xb                  # curvature math done in 32-bit — confirmed sufficient

The model itself keeps running in fast 16-bit mode throughout — only the separate curvature bookkeeping (which is small compared to the model itself) is promoted to 32-bit. This runs natively on the GPU with no slowdown of note (32-bit curvature collection took 95.7 seconds vs. 16-bit’s 116.8 seconds on the same real tensor — not slower at all in this measurement).

6. What this means for the rest of the model

Every one of the 363 tensors went through the same buggy 16-bit curvature step. The 14 tensors that failed outright are simply the ones where the damage was large enough to visibly fail. The other ~349 tensors may have received a valid-looking but quietly-suboptimal correction — we have not yet checked this at scale, only confirmed it for one tensor pair (one broken, one that looked fine). Until we check more broadly, we do not know how many of the “fine-looking” tensors were actually only fine by luck.

A cache-version marker has been added so that once we decide to reprocess, the system can never silently reuse a result computed with the old, buggy 16-bit curvature — every old result is now tagged with which curvature version produced it, and a mismatch forces recomputation rather than silent reuse.

7. What is still open (not yet done, as of this document)


7b. The exact reprocessing command (for when reprocessing is authorized)

The original run’s output folder is Qwen3.8-27B-heretic-ara-YAQA-5bpw (its resume folder — the working directory the pipeline caches partial progress in — is automatically named Qwen3.8-27B-heretic-ara-YAQA-5bpw.yaqa_resume by the orchestrator script itself; nothing in this project’s commands has to specify that name manually). That original folder and its resume folder are left completely untouched — the reprocessing run below targets a brand-new output name so the old (BF16-curvature) results stay available for a full side-by-side comparison against the new (FP32-curvature) results, tensor by tensor.

Confirmed against the actual orchestrator script (run_full_yaqa.sh) before writing this down — the script derives its own resume folder as <output_dir>.yaqa_resume automatically; it does not take a separate --resume-dir override safely (passing one manually would create a mismatch between where the script logs its progress and where the real tensor data actually goes). The correct way to target a new location is simply to give it a new output folder name:

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
./run_full_yaqa.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32 \
  --n-calibration 4 --incoherence none

This will: - Automatically create and use /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume as its working folder — brand new, starts empty. - Process all 363 tensors fresh, using the fixed FP32 curvature code. - Leave the original Qwen3.8-27B-heretic-ara-YAQA-5bpw output and its .yaqa_resume folder completely untouched, for comparison. - Not been run as of this document — this is the exact command, not a report of something already done.


8. Flow diagram — before and after

AFTER — the fixBEFORE — the bugreal, direct 3-waymeasurement found thisCurvature math promoted to32-bit BEFORE it startsModel still runs in fast 16-bit BF16Curvature matrix comes outcorrect (negative values only atordinary 32-bit rounding noise, ~0.0003)YAQA's correction step trustsa curvature matrix that isactually trustworthyResult: matches the most precise64-bit reference to 4 decimal placesCurvature math ALSO done in 16-bit(nobody intended this)Model runs in fast 16-bit BF16Curvature matrix comes outwith impossible negative values(eigenvalues as low as -7.5)YAQA's correction step truststhis broken curvatureResult: 14 tensors end upWORSE than the simplest possible method

9. One-paragraph summary, if you read nothing else

YAQA’s smarter quantization needs to know which parts of each tensor matter — that information is called “curvature.” The code that builds curvature was accidentally running its math in the model’s fast, low-precision 16-bit format instead of a more precise 32-bit format, and for a handful of tensors whose curvature is extremely concentrated in just a couple of real directions, 16-bit precision wasn’t enough — it produced mathematically impossible results that corrupted the correction. Switching that one calculation to 32-bit (not needing the much slower 64-bit) fixes it completely, confirmed by matching the most precise reference available to four decimal places, at no meaningful speed cost.


© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.