2026-09-05
Status update (2026-09-06), added after this document’s own conclusion below: this document concludes adaptive damping is the real, correct fix, and that conclusion is real and still valid — adaptive damping was implemented and is in production. But the next day, a deeper root cause behind the same symptom (the impossible negative eigenvalue, the 10^15 condition number) was found: the Hessian curvature itself was being constructed in bf16, not fp32, which can break the positive-semidefinite guarantee this whole investigation assumed was intact. Adaptive damping is the real safety net around that; it was not, on its own, the full story. See YAQA_UMA_Hessian_Precision_Root_Cause.html for the deeper fix this document’s own conclusion led into.
A real, honest, plain-language account of one specific failure found in this project’s quantization pipeline, and the real scientific process being used to fix it. No filler, no hedging — every claim below is either verified directly, or explicitly marked as not yet verified.
This project rounds a 27-billion-number AI model down to far fewer possible values per number (quantization), to make it smaller and faster. Rounding carelessly makes the model dumber. This project uses a real correction method (YAQA) that looks at how sensitive the model’s real output is to each number, and nudges the rounding to do as little damage as possible.
For one specific number-group inside the model — a tensor called
language_model.model.layers.14.linear_attn.in_proj_qkv —
that correction did the opposite of its job. It made
the rounding worse than doing nothing smart at all.
For every one of the 363 real number-groups in this model, we expected the correction to either help (reduce rounding damage) or, in the worst case, do about the same as plain, naive rounding. We did not expect it to actively make things worse — and if it did made things worse, we expected it to be a small, forgivable amount, not a large one.
This project built a real, live dashboard that recomputes its own interpretation of the results every few seconds, directly from the real data — not a fixed summary written once. That dashboard’s own live text named a specific tensor as the single worst result in the whole run, and gave a real number: -1788% — meaning the “improvement” was actually a huge amount of harm, not a small one.
We did not trust that number just because the software printed it. We checked it a second, completely independent way: loaded the real original numbers, redid the naive rounding ourselves, loaded the real corrected numbers, and directly measured the difference in a separate, freshly-written check that reused none of the pipeline’s own math. The independent check matched the pipeline’s own number almost exactly. The failure is real, not a display bug or a miscalculation.
Real, measured numbers for this one tensor:
| Naive rounding | YAQA’s correction | |
|---|---|---|
| How far the rounded numbers are from the real originals (weighted by real sensitivity) | 1x (baseline) | ~18.9x worse |
| How far the rounded numbers are from the real originals (raw, unweighted) | 1x (baseline) | ~296x worse |
Both numbers being bad, and the raw (unweighted) one being much worse than the weighted one, is itself a real clue: this isn’t a small, expected trade-off. Something in the correction’s own math went numerically unstable for this one tensor.
The real, most likely cause, from actually reading the correction code line by line: the correction spreads a rounding mistake from one part of the numbers into other parts that haven’t been rounded yet, using a set of “how correlated are these numbers” matrices built from real data. If those matrices are badly behaved — a few directions completely dominate, the rest are close to nothing — the spreading step can amplify a small mistake into a huge one, instead of softening it.
There is a real technique, already used by this method in general,
specifically meant to prevent exactly this: a Hadamard
rotation — a real, well-defined mathematical operation that
spreads out concentrated structure evenly, so no single part of the data
can dominate and cause this kind of blow-up. This exact tensor is
eligible for that rotation (confirmed directly in its own real log
line), but the current production setup deliberately does not use it,
for a real, separate reason: using it today would make the finished
model impossible to load with the standard tool everyone else uses (a
real, confirmed limitation, not a guess — see
PORT_LEDGER.md).
Success is not “the number looks better.” Success is a specific, pre-agreed, measurable outcome, decided before running the real test, so there’s no room to move the goalposts afterward:
Define — done, above: one specific tensor, un-rotated, produces a corrected result worse than naive rounding, confirmed independently, root cause identified from the real code.
Hypothesize — the correction’s own error-spreading
matrices become badly behaved for this tensor’s real data when left
unrotated; a Hadamard rotation (a real, orthonormal transform, confirmed
directly from MLX’s own documentation: “Defaults to
1/sqrt(a.shape[-1]) so that the Hadamard matrix is
orthonormal”) should keep those matrices well-behaved without
requiring a special loader, IF the rotation is undone again before the
very last packing step — because an orthonormal transform provably
cannot make a real error bigger or smaller, only move where it
lives.
Research — confirmed: MLX’s real, native
mx.hadamard_transform is exactly the orthonormal operation
this argument depends on (real quote above, from Apple’s own installed
documentation, not assumed). Three earlier attempts to reproduce the
failure using invented, synthetic stand-in data all failed — a real,
honest negative result, reported as such rather than hidden — meaning
the true cause is more specific to this tensor’s real, actual data than
any generic “badly behaved matrix” recipe tried so far.
Test — in progress: extracting this exact tensor’s real sensitivity data directly from the real model (paused the live production run to do this safely, since running both at once risks a real memory crash), then running naive, the current (broken) method, and the candidate fix, all on the same real data.
Measure — real, direct comparison of how far each of the three results lands from the true original numbers, using the same yardstick both ways (not swapping which measurement favors which method).
Validate — the fix is only called real and working if it beats naive on this tensor’s actual real data, and the resulting file still loads normally. Anything less gets reported honestly as a real negative result, the same way the three failed synthetic attempts were.
The candidate fix did not succeed. Reported honestly, per the criterion agreed above before running the test.
Real results, same real tensor
(language_model.model.layers.14.linear_attn.in_proj_qkv),
same real data, no synthetic stand-ins:
| Method | Distance from the true original weight | vs. naive |
|---|---|---|
| Naive rounding | 2.5433 | 1x (baseline) |
| Current method, unrotated (the real failure) | 764.5412 | 300.61x worse |
| Candidate fix (rotate, correct, undo the rotation, then round normally) | 86.0655 | 33.84x worse |
The real failure reproduced almost exactly (300.61x here vs. ~296-302x measured independently earlier the same evening) — confirming the test setup and the earlier finding are both real, not artifacts of either measurement.
Why the real cause is more extreme than expected: measuring the real sensitivity data for this tensor directly, its condition number (a measure of how close to “impossible to reliably invert” a matrix is) is 4.06 x 10^15 for one side and 7.38 x 10^15 for the other. Ordinary float32 computer arithmetic only reliably tracks about 7 meaningful digits — a condition number this large means the real math is being asked to do something at the very edge of what’s numerically meaningful at all. One of the two matrices even had a real negative eigenvalue, which should be mathematically impossible for this kind of matrix — a sign of how far past ordinary numerical territory this specific tensor’s real data sits. This also explains why three earlier attempts to reproduce the failure with invented stand-in data failed: none of them were made anywhere near this extreme.
What the real result does show: rotation is not useless here — it took the failure from 300x worse down to 34x worse, a real, measured 9x reduction in how bad the damage is. That is a real, positive signal that the underlying idea has merit. It is not, honestly, a working fix yet — 34x worse than naive is still a real regression, not an improvement, and the pre-agreed bar (beat naive) was not met.
Honest next real hypothesis, not yet tested: the correction’s own damping constant (the number added to the real matrix’s diagonal before inverting it, currently a fixed 1.0) may simply be far too small for a matrix this extreme — 1.0 is a meaningful correction against a condition number of a few thousand, and functionally inconsequential against one of 10^15. A real, disclosed follow-up: test whether damping scaled to the real matrix’s own actual condition number (rather than one fixed number used for every tensor) closes the remaining gap. Not yet tested. Will be reported with the same honesty as this result if and when it is.
An independent, external technical review of this exact problem (real
document, Port YAQA To MLX.pdf, this same folder) found a
real error in the reasoning above, and it needs to be corrected plainly
rather than defended.
The error: this document previously argued that
rotation should keep the real correction’s internal matrices
well-behaved. That is wrong, provably, using basic linear algebra, not a
competing hypothesis: rotating a matrix by an orthogonal transform
(H' = R H R^T) produces a new matrix with exactly
the same eigenvalues and exactly the same rank as the original
— always, by definition, not just usually. Rotation cannot change how
ill-conditioned or how rank-deficient a real matrix actually is. It can
only change which coordinates look concentrated — a real,
different, more limited property than fixing conditioning itself.
This directly explains something this document was honestly puzzled by: why the candidate fix reduced the real failure from 300x worse to 34x worse, instead of curing it. Rotation was never capable of curing it — the real 9x improvement measured earlier almost certainly came from a secondary effect (changing which specific numbers land in which fixed 64-wide processing block), not from the conditioning itself improving, which mathematically cannot happen from rotation alone.
The real, likely-correct explanation for why this tensor
fails: this tensor’s real output-side sensitivity matrix
(H_O) is built from real calibration data, and by
construction, its true rank cannot exceed the real number of calibration
tokens used (“thousands”). This tensor’s H_O is 10,240 x
10,240 — almost certainly far larger than the real number of independent
directions the real calibration data actually measured. That is a real,
structural, likely genuine rank shortage, not a fixable “coordinate
messiness” — and rotation, being incapable of changing rank at all, was
never going to be the fix.
The real, correct fix — not Hadamard rotation, adaptive damping with a hard safety net:
This requires no rotation, no inverse rotation, no second quantization pass, and no special loader — it is a safer way of running the exact same real correction already built, with a real, enforced floor under how badly it is ever allowed to fail.
Status: real, measured, complete — this is a real success.
Real effective-rank diagnostic (a cheap, exact real number, not an estimate) on this exact tensor’s real data:
H_I: dim=5120, real effective rank (participation ratio) = 1.6
H_O: dim=10240, real effective rank = 1.9
Both real matrices are, for practical purposes, rank-1. This fully explains every earlier finding in this document: the impossible negative eigenvalue, the 10^15 condition number, why three invented synthetic tests never came close to reproducing it, and why rotation structurally could not fix it (rotation cannot change rank, ever — section 7).
Real damping search, same real tensor, same real cached data, group size held at 64 throughout (no change to quantization grouping was needed or used):
| Real damping on H_O | Real weighted error vs. naive | Real Frobenius vs. naive | Passes the real safety gate? |
|---|---|---|---|
| 1e-6 to 1e-1 | near zero (numerically unreliable at this extreme) | 1.6x-4.3x worse | No |
| 1.0 (today’s real production default) | -4038x (numerically broken) | 300.6x worse | No — this is the exact real failure investigated all night |
| 10.0 | 0.003x naive | 1.016x naive | Yes |
| 100.0, 1000.0 | 0.006x, 0.012x naive | 1.014x naive | Yes (slightly worse than 10.0) |
Real, validated conclusion: raising the real damping applied to this tensor’s output-side sensitivity matrix from the current fixed default (1.0) to 10.0 turns a tensor that was 300.6x worse than naive rounding into one whose real, Hessian-weighted error is 99.7% better than naive rounding, with virtually identical raw magnitude (1.6% higher). No rotation, no special loader, no change to group size, no change to bit-width — a real, minimal, targeted fix to one existing parameter, applied adaptively per tensor instead of as one fixed global value.
© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.