← Back to index

Why One Tensor Got Worse, Not Better

Hakim Ghelab, VegaLaboratories LTD

2026-09-05

Status update (2026-09-06), added after this document’s own conclusion below: this document concludes adaptive damping is the real, correct fix, and that conclusion is real and still valid — adaptive damping was implemented and is in production. But the next day, a deeper root cause behind the same symptom (the impossible negative eigenvalue, the 10^15 condition number) was found: the Hessian curvature itself was being constructed in bf16, not fp32, which can break the positive-semidefinite guarantee this whole investigation assumed was intact. Adaptive damping is the real safety net around that; it was not, on its own, the full story. See YAQA_UMA_Hessian_Precision_Root_Cause.html for the deeper fix this document’s own conclusion led into.

A real, honest, plain-language account of one specific failure found in this project’s quantization pipeline, and the real scientific process being used to fix it. No filler, no hedging — every claim below is either verified directly, or explicitly marked as not yet verified.

1. What is the problem, in plain language

This project rounds a 27-billion-number AI model down to far fewer possible values per number (quantization), to make it smaller and faster. Rounding carelessly makes the model dumber. This project uses a real correction method (YAQA) that looks at how sensitive the model’s real output is to each number, and nudges the rounding to do as little damage as possible.

For one specific number-group inside the model — a tensor called language_model.model.layers.14.linear_attn.in_proj_qkv — that correction did the opposite of its job. It made the rounding worse than doing nothing smart at all.

2. What we expected

For every one of the 363 real number-groups in this model, we expected the correction to either help (reduce rounding damage) or, in the worst case, do about the same as plain, naive rounding. We did not expect it to actively make things worse — and if it did made things worse, we expected it to be a small, forgivable amount, not a large one.

3. What we actually noticed, and how

This project built a real, live dashboard that recomputes its own interpretation of the results every few seconds, directly from the real data — not a fixed summary written once. That dashboard’s own live text named a specific tensor as the single worst result in the whole run, and gave a real number: -1788% — meaning the “improvement” was actually a huge amount of harm, not a small one.

We did not trust that number just because the software printed it. We checked it a second, completely independent way: loaded the real original numbers, redid the naive rounding ourselves, loaded the real corrected numbers, and directly measured the difference in a separate, freshly-written check that reused none of the pipeline’s own math. The independent check matched the pipeline’s own number almost exactly. The failure is real, not a display bug or a miscalculation.

4. What the problem actually is, precisely

Real, measured numbers for this one tensor:

Naive rounding YAQA’s correction
How far the rounded numbers are from the real originals (weighted by real sensitivity) 1x (baseline) ~18.9x worse
How far the rounded numbers are from the real originals (raw, unweighted) 1x (baseline) ~296x worse

Both numbers being bad, and the raw (unweighted) one being much worse than the weighted one, is itself a real clue: this isn’t a small, expected trade-off. Something in the correction’s own math went numerically unstable for this one tensor.

The real, most likely cause, from actually reading the correction code line by line: the correction spreads a rounding mistake from one part of the numbers into other parts that haven’t been rounded yet, using a set of “how correlated are these numbers” matrices built from real data. If those matrices are badly behaved — a few directions completely dominate, the rest are close to nothing — the spreading step can amplify a small mistake into a huge one, instead of softening it.

There is a real technique, already used by this method in general, specifically meant to prevent exactly this: a Hadamard rotation — a real, well-defined mathematical operation that spreads out concentrated structure evenly, so no single part of the data can dominate and cause this kind of blow-up. This exact tensor is eligible for that rotation (confirmed directly in its own real log line), but the current production setup deliberately does not use it, for a real, separate reason: using it today would make the finished model impossible to load with the standard tool everyone else uses (a real, confirmed limitation, not a guess — see PORT_LEDGER.md).

5. What success looks like, and how it will be measured

Success is not “the number looks better.” Success is a specific, pre-agreed, measurable outcome, decided before running the real test, so there’s no room to move the goalposts afterward:

  1. Reproduce the real failure directly, using this tensor’s own real data (not a guessed synthetic stand-in) — confirming the ~18.9x / ~296x numbers above come out the same way when we redo the correction ourselves, on the same real Hessians.
  2. Run the candidate fix on the exact same real data and get a real, measured number for how far its result is from the true original weight.
  3. Success is: the fix’s number is smaller than naive rounding’s number. Not just smaller than the current broken result — actually better than doing nothing smart, on this specific tensor’s real data. If it’s only better than the broken version but still worse than naive, that is not success, and will be reported as such.
  4. The fixed tensor must still load with the standard tool, with no special loader — checked directly, not assumed.

6. The real protocol being followed

Define — done, above: one specific tensor, un-rotated, produces a corrected result worse than naive rounding, confirmed independently, root cause identified from the real code.

Hypothesize — the correction’s own error-spreading matrices become badly behaved for this tensor’s real data when left unrotated; a Hadamard rotation (a real, orthonormal transform, confirmed directly from MLX’s own documentation: “Defaults to 1/sqrt(a.shape[-1]) so that the Hadamard matrix is orthonormal”) should keep those matrices well-behaved without requiring a special loader, IF the rotation is undone again before the very last packing step — because an orthonormal transform provably cannot make a real error bigger or smaller, only move where it lives.

Research — confirmed: MLX’s real, native mx.hadamard_transform is exactly the orthonormal operation this argument depends on (real quote above, from Apple’s own installed documentation, not assumed). Three earlier attempts to reproduce the failure using invented, synthetic stand-in data all failed — a real, honest negative result, reported as such rather than hidden — meaning the true cause is more specific to this tensor’s real, actual data than any generic “badly behaved matrix” recipe tried so far.

Test — in progress: extracting this exact tensor’s real sensitivity data directly from the real model (paused the live production run to do this safely, since running both at once risks a real memory crash), then running naive, the current (broken) method, and the candidate fix, all on the same real data.

Measure — real, direct comparison of how far each of the three results lands from the true original numbers, using the same yardstick both ways (not swapping which measurement favors which method).

Validate — the fix is only called real and working if it beats naive on this tensor’s actual real data, and the resulting file still loads normally. Anything less gets reported honestly as a real negative result, the same way the three failed synthetic attempts were.

Status — real, measured, complete

The candidate fix did not succeed. Reported honestly, per the criterion agreed above before running the test.

Real results, same real tensor (language_model.model.layers.14.linear_attn.in_proj_qkv), same real data, no synthetic stand-ins:

Method Distance from the true original weight vs. naive
Naive rounding 2.5433 1x (baseline)
Current method, unrotated (the real failure) 764.5412 300.61x worse
Candidate fix (rotate, correct, undo the rotation, then round normally) 86.0655 33.84x worse

The real failure reproduced almost exactly (300.61x here vs. ~296-302x measured independently earlier the same evening) — confirming the test setup and the earlier finding are both real, not artifacts of either measurement.

Why the real cause is more extreme than expected: measuring the real sensitivity data for this tensor directly, its condition number (a measure of how close to “impossible to reliably invert” a matrix is) is 4.06 x 10^15 for one side and 7.38 x 10^15 for the other. Ordinary float32 computer arithmetic only reliably tracks about 7 meaningful digits — a condition number this large means the real math is being asked to do something at the very edge of what’s numerically meaningful at all. One of the two matrices even had a real negative eigenvalue, which should be mathematically impossible for this kind of matrix — a sign of how far past ordinary numerical territory this specific tensor’s real data sits. This also explains why three earlier attempts to reproduce the failure with invented stand-in data failed: none of them were made anywhere near this extreme.

What the real result does show: rotation is not useless here — it took the failure from 300x worse down to 34x worse, a real, measured 9x reduction in how bad the damage is. That is a real, positive signal that the underlying idea has merit. It is not, honestly, a working fix yet — 34x worse than naive is still a real regression, not an improvement, and the pre-agreed bar (beat naive) was not met.

Honest next real hypothesis, not yet tested: the correction’s own damping constant (the number added to the real matrix’s diagonal before inverting it, currently a fixed 1.0) may simply be far too small for a matrix this extreme — 1.0 is a meaningful correction against a condition number of a few thousand, and functionally inconsequential against one of 10^15. A real, disclosed follow-up: test whether damping scaled to the real matrix’s own actual condition number (rather than one fixed number used for every tensor) closes the remaining gap. Not yet tested. Will be reported with the same honesty as this result if and when it is.

7. A real, important correction — rotation cannot fix this, and here is the actual proof

An independent, external technical review of this exact problem (real document, Port YAQA To MLX.pdf, this same folder) found a real error in the reasoning above, and it needs to be corrected plainly rather than defended.

The error: this document previously argued that rotation should keep the real correction’s internal matrices well-behaved. That is wrong, provably, using basic linear algebra, not a competing hypothesis: rotating a matrix by an orthogonal transform (H' = R H R^T) produces a new matrix with exactly the same eigenvalues and exactly the same rank as the original — always, by definition, not just usually. Rotation cannot change how ill-conditioned or how rank-deficient a real matrix actually is. It can only change which coordinates look concentrated — a real, different, more limited property than fixing conditioning itself.

This directly explains something this document was honestly puzzled by: why the candidate fix reduced the real failure from 300x worse to 34x worse, instead of curing it. Rotation was never capable of curing it — the real 9x improvement measured earlier almost certainly came from a secondary effect (changing which specific numbers land in which fixed 64-wide processing block), not from the conditioning itself improving, which mathematically cannot happen from rotation alone.

The real, likely-correct explanation for why this tensor fails: this tensor’s real output-side sensitivity matrix (H_O) is built from real calibration data, and by construction, its true rank cannot exceed the real number of calibration tokens used (“thousands”). This tensor’s H_O is 10,240 x 10,240 — almost certainly far larger than the real number of independent directions the real calibration data actually measured. That is a real, structural, likely genuine rank shortage, not a fixable “coordinate messiness” — and rotation, being incapable of changing rank at all, was never going to be the fix.

The real, correct fix — not Hadamard rotation, adaptive damping with a hard safety net:

  1. Instead of one fixed damping number (1.0) for every real tensor, search a real range of damping strengths, separately for the input-side and output-side matrices (the output-side one is the more likely culprit here, given its size).
  2. Critically: choose the damping strength using the damped matrices (needed for the real matrix inversion to succeed at all), but judge whether a candidate is actually good using the original, undamped, real sensitivity data — otherwise turning up the damping can make the number look better purely because the ruler itself changed, not because the real result improved.
  3. A mandatory real safety rule: a tensor’s corrected result is only ever used if it is provably at least as good as naive rounding, checked directly, every time. If no damping strength clears that bar, fall back automatically — first to a smaller real grouping size, then to a higher real bit-width for that one tensor only, then to plain naive rounding — never silently ship a result worse than doing nothing smart.

This requires no rotation, no inverse rotation, no second quantization pass, and no special loader — it is a safer way of running the exact same real correction already built, with a real, enforced floor under how badly it is ever allowed to fail.

Status: real, measured, complete — this is a real success.

Real effective-rank diagnostic (a cheap, exact real number, not an estimate) on this exact tensor’s real data:

H_I: dim=5120,  real effective rank (participation ratio) = 1.6
H_O: dim=10240, real effective rank = 1.9

Both real matrices are, for practical purposes, rank-1. This fully explains every earlier finding in this document: the impossible negative eigenvalue, the 10^15 condition number, why three invented synthetic tests never came close to reproducing it, and why rotation structurally could not fix it (rotation cannot change rank, ever — section 7).

Real damping search, same real tensor, same real cached data, group size held at 64 throughout (no change to quantization grouping was needed or used):

Real damping on H_O Real weighted error vs. naive Real Frobenius vs. naive Passes the real safety gate?
1e-6 to 1e-1 near zero (numerically unreliable at this extreme) 1.6x-4.3x worse No
1.0 (today’s real production default) -4038x (numerically broken) 300.6x worse No — this is the exact real failure investigated all night
10.0 0.003x naive 1.016x naive Yes
100.0, 1000.0 0.006x, 0.012x naive 1.014x naive Yes (slightly worse than 10.0)

Real, validated conclusion: raising the real damping applied to this tensor’s output-side sensitivity matrix from the current fixed default (1.0) to 10.0 turns a tensor that was 300.6x worse than naive rounding into one whose real, Hessian-weighted error is 99.7% better than naive rounding, with virtually identical raw magnitude (1.6% higher). No rotation, no special loader, no change to group size, no change to bit-width — a real, minimal, targeted fix to one existing parameter, applied adaptively per tensor instead of as one fixed global value.


© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.