← Back to index ← Back to research index
Draft — not live
Improvement Ledger · Draft 1 · not yet promoted to the live site · 2026-09-10

What our numbers actually are, and how they're actually computed — redone with one real tensor, start to finish

The live pages defined too many terms up front and then jumped straight to results without ever working one real example all the way through. This draft fixes that: every concept below is worked on one real tensor — layers.3.self_attn.q_proj — with the exact real numbers from this project's own files, so every step is checkable by hand.

How to read this: This is a review draft, not a live page. Sit with each section before moving to the next — nothing later depends on anything not already computed earlier. If a section still doesn't click, that's a bug in this document, not something to push through.

0. The three words everything else is built from: vector, matrix, dimension

Not this project's numbers yet — just what these words mean, with the smallest possible example.

Vector

A vector is just an ordered list of numbers. [3, 7, 1] is a vector with 3 numbers in it.

Dimension

The dimension of a vector is simply how many numbers are in it. [3, 7, 1] has dimension 3. In this project, one linear layer's weight might have dimension 5,120 — meaning each input (or output) is a list of 5,120 numbers, not 3. Same idea, just a much longer list.

Matrix

A matrix is a grid of numbers — rows and columns. A 3×3 matrix has 3 rows and 3 columns, 9 numbers total. Here's a real, tiny, made-up-for-teaching 3×3 matrix we'll reuse for the rest of this section:

4
1
0
1
3
1
0
1
2

The purple diagonal cells (4, 3, 2, top-left to bottom-right) matter a lot in a moment — that's the diagonal of the matrix.

1. What "the Hessian" means in this project, specifically

Not the textbook definition in the abstract — what H_I and H_O actually are here, traced to the real log line.

When YAQA builds this model, for every tensor (every weight matrix) it runs real calibration data through the model and records two matrices for that tensor:

  • H_I (Hessian, input side) — a matrix that captures how sensitive this tensor's output is to small changes in each of its input directions.
  • H_O (Hessian, output side) — the same idea, but for the tensor's output directions instead of input directions.

Both H_I and H_O are matrices, same as the 3×3 toy grid above — just much bigger (thousands of rows/columns instead of 3), because a real tensor has thousands of input/output dimensions, not 3.

Real, from disk
The actual log line for our example tensor, unedited, from ~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume/logs/batch_*.log:
language_model.model.layers.3.self_attn.q_proj: H_I shape=(5120, 5120), H_O shape=(12288, 12288), count=24, NaN in H_I=False, NaN in H_O=False
So for this one tensor: H_I is a 5120×5120 matrix (5120 input dimensions), H_O is a 12288×12288 matrix (12288 output dimensions). count=24 confirms this was built from 24 real calibration windows — the "N=24 stratified" you've been asking about, not a smaller/different run.

2. Trace, Frobenius norm, and "effective rank" — computed by hand on the toy matrix

These three ideas are what turn a giant matrix into the one small number ("effective rank") this project actually uses. Worked on the 3×3 toy matrix from Section 0, every arithmetic step shown.

Trace

The trace of a matrix is just: add up the diagonal numbers. Nothing else — ignore every other cell.

trace = 4 + 3 + 2 = 9

Frobenius norm (squared)

The Frobenius norm squared is: square every single number in the matrix (every cell, not just the diagonal), then add all of those squares up.

4² + 1² + 0² + 1² + 3² + 1² + 0² + 1² + 2² = 16+1+0 + 1+9+1 + 0+1+4 = 33

Effective rank — the formula this project actually uses

This project's own definition (confirmed in CASCADE_INVESTIGATION_FULL_REASONING.html's glossary):

effective_rank = trace(H)² / ||H||F² = trace(H)² / (Frobenius norm squared)

Plugging in our toy matrix's real numbers from above:

effective_rank = 9² / 33 = 81 / 33 = 2.45

Why this number, and why it's often a fraction, not a whole number: if a matrix's "risk" were spread perfectly evenly across all 3 of its dimensions, effective_rank would come out to exactly 3 (the true dimension). If instead one single direction dominates everything else, effective_rank comes out much closer to 1 — even though the matrix is still technically 3×3. It's a continuous measure of "how many directions is the risk really spread across," not a count of literal dimensions. That's exactly why the real log line below shows H_I=1.1/5120 — 1.1, not a whole number: this tensor's real risk is concentrated in barely more than one effective direction, out of the 5,120 it technically has.

3. hi_frac, ho_frac, hess_score — computed on the real tensor, not the toy one

Same effective-rank idea as Section 2, now on the real H_I/H_O for layers.3.self_attn.q_proj, using the real numbers YAQA already computed and logged (we don't recompute the giant 5120×5120 matrix ourselves — YAQA already did that arithmetic; we're just reading its answer).

Real, from disk
The real effective-rank line for our tensor, unedited:
language_model.model.layers.3.self_attn.q_proj: real effective rank -- H_I=1.1/5120, H_O=1.8/12288
Step 1 — hi_frac and ho_frac

The code (hessian_story_lib.py, load_hessian()) divides each effective rank by its real dimension:

hi_frac = H_I effective rank / H_I dimension = 1.1 / 5120 = 0.000215 ho_frac = H_O effective rank / H_O dimension = 1.8 / 12288 = 0.0001465
Step 2 — hess_score

The code then just averages those two fractions (hessian_story_lib.py line 197: hess_score = (hi_frac + ho_frac) / 2.0):

hess_score = (0.000215 + 0.0001465) / 2 = 0.00018

This is the exact real number you'll see for this tensor everywhere else on the site — it isn't a separately-computed "danger score," it's literally this average, nothing more.

What "lower = worse" means, concretely, for this tensor

0.00018 is a very low hess_score (close to the dangerous end of the 0–1 scale). In plain terms: this tensor's real risk is concentrated in a very small number of directions out of the thousands it has — a narrow, sharp weak spot, not a risk spread broadly and safely across the whole tensor. That's what "narrow, dangerous weak spot" means on the site's charts — not a vague label, this specific arithmetic.

4. Isolated KL vs. cascaded KL — same real tensor, two different real measurements

Two genuinely different experiments on the same tensor, both real, both on disk, both traced below.

Isolated KL

Run the real model forward on real calibration text. Quantize only this one tensor (layers.3.self_attn.q_proj) to a candidate bit-width; leave every other tensor in the model at full BF16 precision. Compare the model's output probability distribution before and after — the gap between those two distributions, measured with KL divergence, is the isolated KL sensitivity.

Real, from disk
03_LANGUAGE_SENSITIVITY/analysis_Stratified/tensor_report.csv, this tensor's real row: isolated_kl = 0.03073
Cascaded KL

Same idea, but every tensor before this one in the network (layers 0, 1, 2, and the rest of layer 3 up to this point) is already quantized to whatever bit-width the current plan assigned it — not BF16. Only tensors after this one stay at BF16. So this measurement includes real upstream rounding error accumulating into what this tensor sees, which isolated measurement can never show (isolated always starts from a pristine BF16 network).

Real, from disk
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE/cascaded_checkpoint_round1.json, this tensor's real entry (every candidate bit-width, at once — this is one JSON object, not four separate measurements):
"sensitivities": {"4": 0.04022, "5": 0.03080, "6": 0.02055, "8": 0.02045, "16": 0.0}
The site's charts always plot the "4" value specifically (a fixed reference bit-width, so every tensor's cascaded number is comparable to every other tensor's) — here, casc_kl4 = 0.04022.
Comparing the two, for this one real tensor
MeasurementValueWhat it means for this tensor
Isolated KL0.03073Risk if only this tensor is degraded, everything else pristine
Cascaded KL (round 1, at 4-bit)0.04022Risk once real upstream rounding error (layers 0–3) is already baked in

The cascaded number is higher — for this specific tensor, real upstream error made things measurably worse, exactly the effect isolated measurement is structurally blind to.

5. Pearson vs. Spearman correlation — hand-computed on a 3-point example

Not this project's real 360-tensor data yet — a toy example small enough to compute by hand, so the difference between the two statistics is undeniable before we apply either to anything real.

Tensorhess_scoreKL
A0.0010.05
B0.0050.10
C0.0100.30
Spearman — ranks only

Rank each column separately, smallest = rank 1:

Tensorhess_score rankKL rank
A1 (smallest)1 (smallest)
B22
C3 (largest)3 (largest)

Every rank matches exactly → Spearman = +1.00, a perfect correlation, even though the real numbers (0.05, 0.10, 0.30) don't grow in equal steps.

Pearson — the actual numbers, fit to a straight line

Pearson checks how well the real values sit on one straight line. Going from A→B, KL rises by 0.05 for a hess_score rise of 0.004. Going from B→C, KL rises by 0.20 for the same size hess_score rise (0.005). That's not a constant rate — not a straight line — so Pearson comes out high but not a perfect 1.00 (real value for this exact toy set, computed and verified directly: +0.964 — covariance of the two columns, divided by the product of their standard deviations).

The takeaway

Both agree here (both strongly positive) because this toy example was built to agree. In this project's real data, Pearson and Spearman also came out close every time we checked (e.g. −0.34 vs −0.40, +0.31 vs +0.31) — and that agreement between two differently-built statistics is itself part of why the finding is trusted, not just one method's quirk.

6. Putting it together: the real finding, on real numbers, no gaps

Everything above, now applied to this project's actual 360-tensor dataset — the same real files, same methods, no toy numbers left.

ComparisonnPearsonSpearman
hess_score vs. isolated KL (stratified, N=24)360−0.34−0.40
hess_score vs. cascaded round 1 KL (stratified, complete 497/497)360+0.31+0.31
hess_score vs. cascaded round 2 KL (stratified, live, 83.3% so far)295+0.37+0.40

Read the way Section 5 taught: the isolated-KL row is a negative Spearman — meaning low hess_score (dangerous) tends to rank with high isolated KL (dangerous), the direction you'd want from a working risk predictor. The two cascaded rows flip to positive — that relationship reverses once real upstream error is included, and gets stronger, not weaker, as round 2 fills in.

Separately, and just as real: for the 414 tensors measured in both round 1 and round 2 so far, 402 of them (97.1%) show higher cascaded sensitivity in round 2 than round 1 — the raw numbers are still climbing, not settling. Both facts are true at once. What that means for whether the plan is actually converging is still open — this section states the two verified facts precisely; it doesn't yet resolve which story they add up to.

Hakim Ghelab, VegaLaboratories LTD · draft 1, improvement ledger · every number on this page traced to a real file on disk, listed inline · not yet promoted to the live site