The live pages defined too many terms up front and then jumped straight to results without ever working one real example all the way through. This draft fixes that: every concept below is worked on one real tensor — layers.3.self_attn.q_proj — with the exact real numbers from this project's own files, so every step is checkable by hand.
Not this project's numbers yet — just what these words mean, with the smallest possible example.
A vector is just an ordered list of numbers. [3, 7, 1] is a vector with 3 numbers in it.
The dimension of a vector is simply how many numbers are in it. [3, 7, 1] has dimension 3. In this project, one linear layer's weight might have dimension 5,120 — meaning each input (or output) is a list of 5,120 numbers, not 3. Same idea, just a much longer list.
A matrix is a grid of numbers — rows and columns. A 3×3 matrix has 3 rows and 3 columns, 9 numbers total. Here's a real, tiny, made-up-for-teaching 3×3 matrix we'll reuse for the rest of this section:
The purple diagonal cells (4, 3, 2, top-left to bottom-right) matter a lot in a moment — that's the diagonal of the matrix.
Not the textbook definition in the abstract — what H_I and H_O actually are here, traced to the real log line.
When YAQA builds this model, for every tensor (every weight matrix) it runs real calibration data through the model and records two matrices for that tensor:
Both H_I and H_O are matrices, same as the 3×3 toy grid above — just much bigger (thousands of rows/columns instead of 3), because a real tensor has thousands of input/output dimensions, not 3.
~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume/logs/batch_*.log:
count=24 confirms this was built from 24 real calibration windows — the "N=24 stratified" you've been asking about, not a smaller/different run.
These three ideas are what turn a giant matrix into the one small number ("effective rank") this project actually uses. Worked on the 3×3 toy matrix from Section 0, every arithmetic step shown.
The trace of a matrix is just: add up the diagonal numbers. Nothing else — ignore every other cell.
The Frobenius norm squared is: square every single number in the matrix (every cell, not just the diagonal), then add all of those squares up.
This project's own definition (confirmed in CASCADE_INVESTIGATION_FULL_REASONING.html's glossary):
Plugging in our toy matrix's real numbers from above:
Why this number, and why it's often a fraction, not a whole number: if a matrix's "risk" were spread perfectly evenly across all 3 of its dimensions, effective_rank would come out to exactly 3 (the true dimension). If instead one single direction dominates everything else, effective_rank comes out much closer to 1 — even though the matrix is still technically 3×3. It's a continuous measure of "how many directions is the risk really spread across," not a count of literal dimensions. That's exactly why the real log line below shows H_I=1.1/5120 — 1.1, not a whole number: this tensor's real risk is concentrated in barely more than one effective direction, out of the 5,120 it technically has.
Same effective-rank idea as Section 2, now on the real H_I/H_O for layers.3.self_attn.q_proj, using the real numbers YAQA already computed and logged (we don't recompute the giant 5120×5120 matrix ourselves — YAQA already did that arithmetic; we're just reading its answer).
The code (hessian_story_lib.py, load_hessian()) divides each effective rank by its real dimension:
The code then just averages those two fractions (hessian_story_lib.py line 197: hess_score = (hi_frac + ho_frac) / 2.0):
This is the exact real number you'll see for this tensor everywhere else on the site — it isn't a separately-computed "danger score," it's literally this average, nothing more.
0.00018 is a very low hess_score (close to the dangerous end of the 0–1 scale). In plain terms: this tensor's real risk is concentrated in a very small number of directions out of the thousands it has — a narrow, sharp weak spot, not a risk spread broadly and safely across the whole tensor. That's what "narrow, dangerous weak spot" means on the site's charts — not a vague label, this specific arithmetic.
Two genuinely different experiments on the same tensor, both real, both on disk, both traced below.
Run the real model forward on real calibration text. Quantize only this one tensor (layers.3.self_attn.q_proj) to a candidate bit-width; leave every other tensor in the model at full BF16 precision. Compare the model's output probability distribution before and after — the gap between those two distributions, measured with KL divergence, is the isolated KL sensitivity.
03_LANGUAGE_SENSITIVITY/analysis_Stratified/tensor_report.csv, this tensor's real row: isolated_kl = 0.03073Same idea, but every tensor before this one in the network (layers 0, 1, 2, and the rest of layer 3 up to this point) is already quantized to whatever bit-width the current plan assigned it — not BF16. Only tensors after this one stay at BF16. So this measurement includes real upstream rounding error accumulating into what this tensor sees, which isolated measurement can never show (isolated always starts from a pristine BF16 network).
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE/cascaded_checkpoint_round1.json, this tensor's real entry (every candidate bit-width, at once — this is one JSON object, not four separate measurements):
"4" value specifically (a fixed reference bit-width, so every tensor's cascaded number is comparable to every other tensor's) — here, casc_kl4 = 0.04022.
| Measurement | Value | What it means for this tensor |
|---|---|---|
| Isolated KL | 0.03073 | Risk if only this tensor is degraded, everything else pristine |
| Cascaded KL (round 1, at 4-bit) | 0.04022 | Risk once real upstream rounding error (layers 0–3) is already baked in |
The cascaded number is higher — for this specific tensor, real upstream error made things measurably worse, exactly the effect isolated measurement is structurally blind to.
Not this project's real 360-tensor data yet — a toy example small enough to compute by hand, so the difference between the two statistics is undeniable before we apply either to anything real.
| Tensor | hess_score | KL |
|---|---|---|
| A | 0.001 | 0.05 |
| B | 0.005 | 0.10 |
| C | 0.010 | 0.30 |
Rank each column separately, smallest = rank 1:
| Tensor | hess_score rank | KL rank |
|---|---|---|
| A | 1 (smallest) | 1 (smallest) |
| B | 2 | 2 |
| C | 3 (largest) | 3 (largest) |
Every rank matches exactly → Spearman = +1.00, a perfect correlation, even though the real numbers (0.05, 0.10, 0.30) don't grow in equal steps.
Pearson checks how well the real values sit on one straight line. Going from A→B, KL rises by 0.05 for a hess_score rise of 0.004. Going from B→C, KL rises by 0.20 for the same size hess_score rise (0.005). That's not a constant rate — not a straight line — so Pearson comes out high but not a perfect 1.00 (real value for this exact toy set, computed and verified directly: +0.964 — covariance of the two columns, divided by the product of their standard deviations).
Both agree here (both strongly positive) because this toy example was built to agree. In this project's real data, Pearson and Spearman also came out close every time we checked (e.g. −0.34 vs −0.40, +0.31 vs +0.31) — and that agreement between two differently-built statistics is itself part of why the finding is trusted, not just one method's quirk.
Everything above, now applied to this project's actual 360-tensor dataset — the same real files, same methods, no toy numbers left.
| Comparison | n | Pearson | Spearman |
|---|---|---|---|
| hess_score vs. isolated KL (stratified, N=24) | 360 | −0.34 | −0.40 |
| hess_score vs. cascaded round 1 KL (stratified, complete 497/497) | 360 | +0.31 | +0.31 |
| hess_score vs. cascaded round 2 KL (stratified, live, 83.3% so far) | 295 | +0.37 | +0.40 |
Read the way Section 5 taught: the isolated-KL row is a negative Spearman — meaning low hess_score (dangerous) tends to rank with high isolated KL (dangerous), the direction you'd want from a working risk predictor. The two cascaded rows flip to positive — that relationship reverses once real upstream error is included, and gets stronger, not weaker, as round 2 fills in.
Separately, and just as real: for the 414 tensors measured in both round 1 and round 2 so far, 402 of them (97.1%) show higher cascaded sensitivity in round 2 than round 1 — the raw numbers are still climbing, not settling. Both facts are true at once. What that means for whether the plan is actually converging is still open — this section states the two verified facts precisely; it doesn't yet resolve which story they add up to.