← Back to index ← Back to research index
Cascaded Sensitivity Investigation · Full Reasoning, Ground Up · 2026-09-10

Does the free Hessian signal predict real risk? The whole investigation, hand-verifiable, start to finish.

Written so a reader who has never seen this project before can follow every definition, every formula, and every real number to the finding it produces — with nothing assumed and nothing skipped.

How to read this page: Section 0 defines every technical term used anywhere below, in plain language, before it is used. Every later section uses only terms already defined. Every real number is traced to a real file on disk. Nothing is illustrative or invented.

0. Glossary — every term, defined once, before it is used anywhere else

Tensor
One named block of numbers inside the AI model — e.g. "layer 14's query-projection weights." The whole model is made of hundreds of these blocks. Quantizing (shrinking) the model means shrinking each tensor separately.
Matrix
A rectangular grid of numbers. A Hessian (defined next) is one kind of matrix.
Hessian (in this project, written H)
For one tensor, a real matrix built by actually running the model on real text and recording, mathematically, "if this tensor's weights are nudged in some direction, how much does the model's real output change?" Bigger numbers in the matrix mean that direction is more dangerous to disturb. It is computed from real forward and backward passes through the real model — not estimated, not assumed.
Dimension (of a tensor / of its Hessian), written N
How many independent numbers make up that tensor's input or output. A tensor with 5120 input dimensions has a Hessian that is a 5120×5120 matrix. Real examples used later: 5120, 6144, 10240, 17408.
Eigenvalue / eigenvector
Every real square matrix like a Hessian can be described as a set of independent "directions" (eigenvectors), each with its own single number (eigenvalue) saying how much curvature/risk exists in that one direction. A matrix with N dimensions has exactly N of these direction/number pairs. If the eigenvalues are all similar, risk is spread evenly. If one eigenvalue is huge and the rest are near zero, risk is concentrated in a single direction.
Trace of a matrix, trace(H)
Add up the numbers running down the matrix's diagonal. Also exactly equals the sum of all N eigenvalues — two equivalent ways to compute the same real number.
Frobenius norm squared, ||H||F²
Square every single number in the matrix, then add them all up. Also exactly equals the sum of the squares of all N eigenvalues.
Effective rank / participation ratio
A single real number, computed as trace(H)² / ||H||F², that answers "out of N possible directions, how many are actually carrying real risk?" Ranges from 1 (all risk in one direction — dangerous) up to N (risk spread evenly across everything — safe). Full derivation of why this formula does that: Section 1.
Danger score (hi_frac / ho_frac in this project's own code)
Effective rank divided by the real dimension N. A fraction between 0 and 1. Near 0 = dangerous (risk concentrated in almost nothing). Near 1 = safe (risk spread across nearly everything). This is the number plotted on every scatter chart in this project labeled "danger score."
KL divergence (Kullback–Leibler divergence)
A real, standard statistic measuring how different two probability distributions are. Here: how different the model's real output becomes after a tensor is quantized, compared to the model's real, unmodified output. Zero means no change; bigger numbers mean the quantization damaged the model's real behavior more.
Isolated sensitivity measurement
The ORIGINAL, older measurement method: quantize one tensor, measure its KL divergence, then restore it to full precision before testing the next tensor. Every tensor is tested as if it were the ONLY thing ever quantized — which never matches the real deployed model, where every tensor is quantized at once.
Cascaded sensitivity measurement
This project's newer, corrected method: quantize a tensor, measure its KL divergence, then leave it quantized (locked at its real planned bit-width) before moving to the next tensor. Every tensor after the first is now measured against the network's real, actually-degraded upstream state — not an idealized untouched one.
Bit-width / bits
How many bits are used to store each number in a quantized tensor. Lower = smaller/faster but less precise (more real risk of damage). This project tests candidates 4, 5, 6, 8, and 16 (16 effectively means "don't shrink this tensor at all").
Calibration window / stratified calibration
A real chunk of real text run through the model to produce the real activations used in the measurements above. "Stratified" means the real text is deliberately drawn evenly from every real topic category (6 domains in this project), instead of randomly — so no single category's text can dominate or get silently dropped.
plan.json / "the plan"
A real file listing, for every real tensor in the model, which bit-width it has been assigned. The end product of the whole optimization process.
Baseline plan
Whatever existing plan.json is used as the STARTING point / assumed upstream context for a new round of cascaded measurement.
Floor (e.g. "early-QKV floor")
A hard rule forcing a whole named group of tensors to a minimum bit-width, regardless of what their own individual measurement says. Exists because this project already found real, documented regions where the measurement itself is known to be unreliable (see Section 5).
MILP (Mixed-Integer Linear Program)
The real optimization method used to pick each tensor's final bit-width, balancing "minimize real quality loss" against "stay within the real total size budget," subject to any floor rules above.
Round / fixed-point loop
Because a cascaded measurement's results depend on which baseline plan was used as context, this project re-measures repeatedly: measure → re-solve a new plan → check if the new plan agrees with the one just used as context → if not, use the new plan as the next round's context and repeat. Explained fully in Section 6.
Pearson correlation, r
A standard real statistic between −1 and +1 measuring how strongly two lists of numbers move together. +1 = they rise and fall in perfect lockstep. −1 = when one rises, the other falls, in perfect lockstep. 0 = no real relationship.
Saturation
When a real measured quantity stops being able to distinguish between cases because it has hit a practical ceiling or floor — explained with real evidence in Section 4.

1. The formula, derived from first principles

Source, verbatim, not paraphrased: yaqa_core.py:328.

def effective_rank(H):
    return trace(H)**2 / frobenius_norm(H)**2

Why this specific formula detects "concentrated vs. spread out"

Two toy 3-eigenvalue matrices, same total risk, different shape:

Caseeigenvaluestracetrace²Σ(eigenvalue²)effective rank
Spread evenly (safe)3, 3, 39812781/27 = 3.0
Concentrated (dangerous)9, 0, 09818181/81 = 1.0

Identical trace=9 in both cases (same total risk budget). The spread-out case scores the true dimension, 3, in full — nothing hidden. The concentrated case collapses to 1, correctly flagging that only one of the three directions actually matters, even though the matrix is technically 3-dimensional. This is the entire mechanism: comparing sum-then-square against square-then-sum tells the two shapes apart.

Turning it into the plotted "danger score"

danger_score = effective_rank / N

Real example, read directly from a real log line: H_I=15.6/6144 means effective_rank=15.6 at dimension N=6144, so danger_score = 15.6/6144 = 0.00254. Near 0 = dangerous. Near 1 = safe.

2. Is this real, corrected data, or the old flawed run?

You corrected a previous run that used N=8 calibration windows, which was found to be unreliable, and redid it with the equivalent of 24 stratified windows. Is the current data actually that corrected run?

1

Find the real, currently-running process

ps aux | grep 06_measure_sensitivity_cascaded_V4

Real result: PID 4722, running since 3:55pm, command line includes:

--n-calibration 4 --calibration-mode stratified --max-rounds 3 --output-dir 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE
2

Confirm what "--n-calibration 4 --calibration-mode stratified" actually means

Direct from the script's own real, dated code comment (06_measure_sensitivity_cascaded_V4.py, argparse help text for --n-calibration): under stratified mode, this number is PER DOMAIN, and there are 6 real domains. So 4 × 6 = 24 total calibration windows — explicitly documented in the code as chosen "matching the original N=24 flat baseline size for a fair, deliberate comparison."

Finding
Confirmed real and correct: the live process is using the corrected 24-window stratified protocol, not the old N=8 run. This is not assumed — it is read directly from the actual running process's own command line and the script's own documented reasoning for that default.

3. Round 1, complete (497 of 497 tensors): does the danger score predict real risk?

Step 1 — the raw, unfiltered correlation

Of the 497 real tensors round 1 measured, 360 also have a real danger score on record (the other 137 were never covered by the Hessian log — a real, separate gap, not an error in this calculation). Pearson correlation between danger_score and real cascaded KL (at bits=4), computed on those 360 real pairs:

Pearson r = +0.310

A positive number here is the wrong-direction sign: it means low danger_score (flagged as dangerous by the free signal) is weakly associated with LOWER real sensitivity, not higher. If the free signal were working as intended, this number should be negative.

Step 2 — hypothesis: is this a depth confound?

Splitting the same 360 real tensors into four depth bands (layers 0–15, 16–31, 32–47, 48–63) and computing the correlation separately within each band:

Layer bandnmean real sensitivitymean danger scorer within band
L0–15720.0370.00106−0.099
L16–31920.0740.00067+0.339
L32–47990.1460.00189+0.114
L48–63970.193−0.037−0.037

Real sensitivity climbs roughly 5× from the shallowest to the deepest band — expected, since cascaded error compounds through the network. But within any single band, the danger score's own predictive power is weak and inconsistent in sign (−0.10, +0.34, +0.11, −0.04). This raised the hypothesis that the global +0.31 was mostly depth leaking into the number, not real local signal.

Step 3 — testing that hypothesis directly (partial correlation)

Method: fit a straight line predicting danger_score from layer depth alone, and a separate straight line predicting real sensitivity from layer depth alone. Subtract each tensor's predicted value from its real value (the "residual" — what depth alone could NOT explain). Then correlate the two sets of residuals.

Raw r (depth not removed) = 0.310
Partial r (depth removed) = 0.282
Result of the test
The partial correlation barely moved (0.310→0.282). The depth-confound hypothesis was wrong — depth is not the dominant explanation. This is reported here, corrected, exactly as it was found to be wrong in real time; it is not hidden.

Step 4 — the real explanation: saturation, not depth

While testing the depth hypothesis, one number stood out: fitting real sensitivity from depth ALONE gave R²=0.968 — depth alone "explaining" 96.8% of the variance is far too clean to be a genuine linear relationship across 360 real, independently-measured tensors. This is the signature of the measurement running out of room, not a real clean trend. Checking directly, real per-band statistics:

Layer bandnat/near ceiling (≥0.195)minmaxstd deviation
L0–151240%0.0170.0740.010
L16–311240%0.0380.0950.014
L32–471240%0.0950.1820.026
L48–6312453%0.1800.2000.006
Real, verified finding
In the deepest quarter of the network, 53% of tensors (66 of 124) are pinned within 1% of a hard ceiling (≈0.1998), with almost no spread left (std collapses to 0.006 vs. 0.010–0.026 elsewhere). Past that point, the cascaded measurement can no longer tell which specific deep-layer tensor is riskier than another — more than half of them simply read "maxed out." This, not depth generally, is the real confound behind the misleading +0.31 global number.

Step 5 — direct proof the free signal misses real risk even where it CAN discriminate

L7.self_attn.q_proj was one of the 10 tensors this project's danger score flagged as MOST dangerous (danger_score=0.00018, among the lowest — i.e. most concentrated — in the whole dataset). Its real, measured cascaded sensitivity: 0.035 — near the bottom of the entire distribution. The free signal called it dangerous; the real, expensive measurement says it was one of the safer tensors in the network.

4. Does the current baseline plan actually agree with the Hessian data?

Are we automatically underprotecting tensors the Hessian flags as risky, or overprotecting tensors it says are safe, via a blanket floor rule?

Step 1 — find the exact real rule

Source, verbatim: 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py, function identify_early_qkv_layers. It selects every tensor whose name ends in one of a fixed real set of tokens — q_proj, k_proj, v_proj, in_proj_qkv, in_proj_a, in_proj_b, qkv_proj, wqkv, Wqkv, attn_qkv — that also sits in the first 25% of the network's real depth (layers 0–15 of 64), and forces each one to a minimum bit-width via a hard constraint in the real MILP solve, regardless of that specific tensor's own measured sensitivity.

Step 2 — a real mistake caught and corrected mid-analysis

The first attempt at this check used a different, wrong token set (borrowed from an unrelated chart-coloring function elsewhere in this project, which happens to also include in_proj_z). That produced a nonsensical result: two tensors appearing to violate the floor rule. Re-reading the optimizer's own real token list showed in_proj_z is deliberately NOT part of the early-QKV family — it is excluded on purpose. Redone with the exact, correct token set, the apparent contradiction disappeared. This correction is left visible here rather than hidden, per the standard this page is held to.

Step 3 — the real result, with the correct token set

Real, verified finding
48 real tensors are eligible for the early-QKV floor (layers 0–15, exact token match). Of those, all 48 already had ≥6 bits assigned by the baseline plan on their own — zero were below the floor to begin with. The floor rule had no actual effect on this baseline plan. It exists as a safety net for a scenario that, in this specific real plan, never occurred.

5. Baseline plan vs. round 1's own re-solved plan — do they agree?

Method

Direct comparison of two real files: the external baseline plan (V3_3_APPLES_TO_APPLES_0005/plan_Qwen38_V3_3_..._BPW.json) and round 1's own output (plan_v4_round1.json), tensor by tensor, comparing assigned bit-width.

Real, verified finding
145 of 497 tensors (29%) were assigned a different bit-width between the two plans — some swinging dramatically, e.g. 5→16 bits or 16→5 bits. This is real, substantial disagreement, not noise.

Why this matters — the fixed-point loop

Direct from the real code's own docstring (fixed_point_loop): "a single cascaded-measurement pass alone isn't enough, because the global MILP re-solve can assign DIFFERENT bits than the plan that was used as the measurement context." That is exactly what happened here. Because round 1's own re-solved plan disagrees with what it was measured against, the loop automatically starts round 2, using round 1's own plan as the new context.

The same docstring states plainly: max_rounds=3 "should converge in 1-2 rounds if the gap is small." The real, measured gap here is 29% — not small. Convergence by round 3 is genuinely uncertain, not guaranteed, by the method's own stated design assumption.

6. Round 1 vs. round 2, same tensors — is this a yo-yo, or a real shift?

Step 1 — why "n=107" (and then 111, then 112) is a moving, honest number

Round 2 is a live, still-running process. Every time its real output file is read, it may contain more tensors than the last read — the measurement has simply progressed further in the intervening minutes. "n=107" and "n=112" are not different answers to the same question; they are the honest answer to the same question asked at two different real moments.

1

Count what round 2 has measured so far

len(cascaded_checkpoint_round2.json) = 173 (grows over time)
2

Of those, count how many also have a real Hessian score

len([n for n in round2 if n in hessian_log]) = 107

The other 66 (all real layer-0 tensors, e.g. L0.mlp.down_proj) simply have no entry in the Hessian log at all — a real, separate data gap, not an error in the join.

Step 2 — apples-to-apples comparison

Restricting round 1's own values to that exact same 107-tensor set (not round 1's full 360):

MeasurementnPearson r
Round 1, same 107 tensors107−0.41
Round 2, those same 107 tensors107−0.40
Real, verified finding
Round 1 and round 2 agree on this identical tensor set — nearly the same correlation, same sign. The earlier apparent contradiction (round 1 full = +0.31 vs. round 2 partial = −0.40) was a same-tensor-population mismatch, not the same tensors disagreeing between rounds.

Step 3 — do the actual values move, and if so, how?

Direct per-tensor comparison, round 1 value vs. round 2 value, for the same real tensors. Every one of the largest movers shifted in the same direction: real sensitivity increased from round 1 to round 2 (mean 0.0415 → 0.0554 across this subset). Not oscillation — a consistent, one-directional shift.

Step 4 — why the shift is real and expected, not noise

Round 1's upstream context (layers before the one being measured) was locked at the external baseline plan's bits. Round 2's upstream context is locked at round 1's own re-solved plan's bits — and Section 5 already showed those two plans disagree on 29% of tensors. Since round 2 is currently measuring layers ~15–22, and the upstream layers below that range genuinely differ in real quantization state between the two contexts, the same downstream tensor legitimately measures differently. This is the cascaded method working as designed, not a flaw.

7. What is proven, and what is still genuinely open

QuestionAnswerEvidence
Is this real, corrected (24-window) data?YesLive process command line, Section 2
Does the danger score predict real risk, globally?Weakly, and the wrong sign globallyr=+0.31 raw; real cause is late-layer saturation, Section 3
Did the early-QKV floor over-protect anything?No — it had zero effect hereAll 48 eligible tensors already ≥6 bits, Section 4
Does the baseline plan agree with round 1's own re-solve?No — 29% disagreement145/497 tensors differ, Section 5
Is round 1 vs. round 2 a yo-yo?No — consistent, one-directionalSame-tensor comparison, Section 6
Does this run converge by round 3?Not yet knownRound 2 incomplete, no plan_v4_round2.json yet