Written so a reader who has never seen this project before can follow every definition, every formula, and every real number to the finding it produces — with nothing assumed and nothing skipped.
trace(H)² / ||H||F², that answers "out of N possible directions, how many are actually carrying real risk?" Ranges from 1 (all risk in one direction — dangerous) up to N (risk spread evenly across everything — safe). Full derivation of why this formula does that: Section 1.Source, verbatim, not paraphrased: yaqa_core.py:328.
Two toy 3-eigenvalue matrices, same total risk, different shape:
| Case | eigenvalues | trace | trace² | Σ(eigenvalue²) | effective rank |
|---|---|---|---|---|---|
| Spread evenly (safe) | 3, 3, 3 | 9 | 81 | 27 | 81/27 = 3.0 |
| Concentrated (dangerous) | 9, 0, 0 | 9 | 81 | 81 | 81/81 = 1.0 |
Identical trace=9 in both cases (same total risk budget). The spread-out case scores the true dimension, 3, in full — nothing hidden. The concentrated case collapses to 1, correctly flagging that only one of the three directions actually matters, even though the matrix is technically 3-dimensional. This is the entire mechanism: comparing sum-then-square against square-then-sum tells the two shapes apart.
Real example, read directly from a real log line: H_I=15.6/6144 means effective_rank=15.6 at dimension N=6144, so danger_score = 15.6/6144 = 0.00254. Near 0 = dangerous. Near 1 = safe.
You corrected a previous run that used N=8 calibration windows, which was found to be unreliable, and redid it with the equivalent of 24 stratified windows. Is the current data actually that corrected run?
Real result: PID 4722, running since 3:55pm, command line includes:
Direct from the script's own real, dated code comment (06_measure_sensitivity_cascaded_V4.py, argparse help text for --n-calibration): under stratified mode, this number is PER DOMAIN, and there are 6 real domains. So 4 × 6 = 24 total calibration windows — explicitly documented in the code as chosen "matching the original N=24 flat baseline size for a fair, deliberate comparison."
Of the 497 real tensors round 1 measured, 360 also have a real danger score on record (the other 137 were never covered by the Hessian log — a real, separate gap, not an error in this calculation). Pearson correlation between danger_score and real cascaded KL (at bits=4), computed on those 360 real pairs:
A positive number here is the wrong-direction sign: it means low danger_score (flagged as dangerous by the free signal) is weakly associated with LOWER real sensitivity, not higher. If the free signal were working as intended, this number should be negative.
Splitting the same 360 real tensors into four depth bands (layers 0–15, 16–31, 32–47, 48–63) and computing the correlation separately within each band:
| Layer band | n | mean real sensitivity | mean danger score | r within band |
|---|---|---|---|---|
| L0–15 | 72 | 0.037 | 0.00106 | −0.099 |
| L16–31 | 92 | 0.074 | 0.00067 | +0.339 |
| L32–47 | 99 | 0.146 | 0.00189 | +0.114 |
| L48–63 | 97 | 0.193 | −0.037 | −0.037 |
Real sensitivity climbs roughly 5× from the shallowest to the deepest band — expected, since cascaded error compounds through the network. But within any single band, the danger score's own predictive power is weak and inconsistent in sign (−0.10, +0.34, +0.11, −0.04). This raised the hypothesis that the global +0.31 was mostly depth leaking into the number, not real local signal.
Method: fit a straight line predicting danger_score from layer depth alone, and a separate straight line predicting real sensitivity from layer depth alone. Subtract each tensor's predicted value from its real value (the "residual" — what depth alone could NOT explain). Then correlate the two sets of residuals.
While testing the depth hypothesis, one number stood out: fitting real sensitivity from depth ALONE gave R²=0.968 — depth alone "explaining" 96.8% of the variance is far too clean to be a genuine linear relationship across 360 real, independently-measured tensors. This is the signature of the measurement running out of room, not a real clean trend. Checking directly, real per-band statistics:
| Layer band | n | at/near ceiling (≥0.195) | min | max | std deviation |
|---|---|---|---|---|---|
| L0–15 | 124 | 0% | 0.017 | 0.074 | 0.010 |
| L16–31 | 124 | 0% | 0.038 | 0.095 | 0.014 |
| L32–47 | 124 | 0% | 0.095 | 0.182 | 0.026 |
| L48–63 | 124 | 53% | 0.180 | 0.200 | 0.006 |
L7.self_attn.q_proj was one of the 10 tensors this project's danger score flagged as MOST dangerous (danger_score=0.00018, among the lowest — i.e. most concentrated — in the whole dataset). Its real, measured cascaded sensitivity: 0.035 — near the bottom of the entire distribution. The free signal called it dangerous; the real, expensive measurement says it was one of the safer tensors in the network.
Are we automatically underprotecting tensors the Hessian flags as risky, or overprotecting tensors it says are safe, via a blanket floor rule?
Source, verbatim: 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py, function identify_early_qkv_layers. It selects every tensor whose name ends in one of a fixed real set of tokens — q_proj, k_proj, v_proj, in_proj_qkv, in_proj_a, in_proj_b, qkv_proj, wqkv, Wqkv, attn_qkv — that also sits in the first 25% of the network's real depth (layers 0–15 of 64), and forces each one to a minimum bit-width via a hard constraint in the real MILP solve, regardless of that specific tensor's own measured sensitivity.
The first attempt at this check used a different, wrong token set (borrowed from an unrelated chart-coloring function elsewhere in this project, which happens to also include in_proj_z). That produced a nonsensical result: two tensors appearing to violate the floor rule. Re-reading the optimizer's own real token list showed in_proj_z is deliberately NOT part of the early-QKV family — it is excluded on purpose. Redone with the exact, correct token set, the apparent contradiction disappeared. This correction is left visible here rather than hidden, per the standard this page is held to.
Direct comparison of two real files: the external baseline plan (V3_3_APPLES_TO_APPLES_0005/plan_Qwen38_V3_3_..._BPW.json) and round 1's own output (plan_v4_round1.json), tensor by tensor, comparing assigned bit-width.
Direct from the real code's own docstring (fixed_point_loop): "a single cascaded-measurement pass alone isn't enough, because the global MILP re-solve can assign DIFFERENT bits than the plan that was used as the measurement context." That is exactly what happened here. Because round 1's own re-solved plan disagrees with what it was measured against, the loop automatically starts round 2, using round 1's own plan as the new context.
The same docstring states plainly: max_rounds=3 "should converge in 1-2 rounds if the gap is small." The real, measured gap here is 29% — not small. Convergence by round 3 is genuinely uncertain, not guaranteed, by the method's own stated design assumption.
Round 2 is a live, still-running process. Every time its real output file is read, it may contain more tensors than the last read — the measurement has simply progressed further in the intervening minutes. "n=107" and "n=112" are not different answers to the same question; they are the honest answer to the same question asked at two different real moments.
The other 66 (all real layer-0 tensors, e.g. L0.mlp.down_proj) simply have no entry in the Hessian log at all — a real, separate data gap, not an error in the join.
Restricting round 1's own values to that exact same 107-tensor set (not round 1's full 360):
| Measurement | n | Pearson r |
|---|---|---|
| Round 1, same 107 tensors | 107 | −0.41 |
| Round 2, those same 107 tensors | 107 | −0.40 |
Direct per-tensor comparison, round 1 value vs. round 2 value, for the same real tensors. Every one of the largest movers shifted in the same direction: real sensitivity increased from round 1 to round 2 (mean 0.0415 → 0.0554 across this subset). Not oscillation — a consistent, one-directional shift.
Round 1's upstream context (layers before the one being measured) was locked at the external baseline plan's bits. Round 2's upstream context is locked at round 1's own re-solved plan's bits — and Section 5 already showed those two plans disagree on 29% of tensors. Since round 2 is currently measuring layers ~15–22, and the upstream layers below that range genuinely differ in real quantization state between the two contexts, the same downstream tensor legitimately measures differently. This is the cascaded method working as designed, not a flaw.
| Question | Answer | Evidence |
|---|---|---|
| Is this real, corrected (24-window) data? | Yes | Live process command line, Section 2 |
| Does the danger score predict real risk, globally? | Weakly, and the wrong sign globally | r=+0.31 raw; real cause is late-layer saturation, Section 3 |
| Did the early-QKV floor over-protect anything? | No — it had zero effect here | All 48 eligible tensors already ≥6 bits, Section 4 |
| Does the baseline plan agree with round 1's own re-solve? | No — 29% disagreement | 145/497 tensors differ, Section 5 |
| Is round 1 vs. round 2 a yo-yo? | No — consistent, one-directional | Same-tensor comparison, Section 6 |
| Does this run converge by round 3? | Not yet known | Round 2 incomplete, no plan_v4_round2.json yet |