← Back to index
Real analysis · Qwen3.8-27B-heretic-ara · YAQA-MTP pipeline · live-refreshed

A number you already have, for free, predicts which tensors will hurt

Every YAQA build already computes a real curvature measurement (the Hessian) for every tensor, as a side effect of correcting the rounding. That data has been sitting on disk, unused for anything except rounding quality. This is what happens when you actually look at it.

Why this experiment exists

Three real, separate things feed this page — worth being precise about, since they're easy to blur together. Isolated KL is the old sensitivity-ranking method: quantize one tensor at a time, measure real output KL divergence, reset to a clean model before testing the next tensor. It's pure sensitivity estimation, nothing to do with rounding-error correction — and because it resets each time, it structurally can't see the noise that piles up once earlier layers are actually quantized, so it underrates risk, and the underrating gets dramatically worse deeper into the model: real data on this page shows the isolated-to-cascaded gap growing from about 1.2x at layer 0 to roughly 500x by layer 60. The Hessian (YAQA) is a separate, later addition: not a sensitivity-ranking tool at all, but a two-sided H_I/H_O correction applied when actually baking chosen bit-widths into real weights, to reduce rounding error. The cascaded measurement redoes the sensitivity ranking without resetting — each layer measured with everything before it already quantized, capturing the real compounding effect the isolated method can't. This page tests whether the two cheap, already-computed signals — isolated KL and the Hessian danger score — can together predict what the expensive cascaded measurement finds. If they can, that's real GPU-hours saved on top of a better bit-allocation plan.

Grounded in what's actually true

This entire analysis — the full 40+ hour cascaded measurement, the YAQA Hessian correction, every chart on this page — runs on a single Apple M3 Max, 128GB RAM (Mac15,9). As far as this project's own research has been able to establish, it is the first port of YAQA's rounding-error correction to Apple's MLX framework. That claim is a stated goal of the work, not an independently audited record — it is written here as motivation, not as a verified fact the way the numbers below are.

The setup, in one paragraph

Every tensor in the model has a real, already-computed "Hessian" — a measurement of how the model reacts when you nudge that tensor's weights. Some tensors react a little, spread evenly across many directions: safe, forgiving, no single mistake can hurt much. Others barely react at all — except in one or two very specific directions, where a small mistake gets amplified hard. We turned that into one number per tensor: a danger score. Lower means more dangerous — the tensor's risk is concentrated in a narrow weak spot instead of spread out.

496/497
Tensors with a real danger score
Pulled from real YAQA build logs — zero extra compute
497/497
Tensors with real ground-truth sensitivity
Round 3 of 3 rounds so far, 100% done · round(s) 1, 2 already complete
+0.17
Real correlation between them
n=496 · Spearman (rank, calibration-blind): +0.17

Finding 0 — the baseline itself was wrong, before any correction ran

Same real Hessian danger score, same real tensors, correlated against ground-truth risk at four stages: the isolated-only baseline this project used to build every plan before any cascade round existed, then each real cascade round measured so far on this live run. Both Pearson (straight-line agreement) and Spearman (agreement on ranking only, immune to calibration/scale) are shown, because the story only holds up if both tell it.

+0.50+0.25+0.00-0.25-0.50P0isolated baseline (n=496)P1round 1 (n=496)P2round 2 (n=496)P3round 3 (n=496)

Purple = Pearson, pink = Spearman. Both cross from negative to positive between P0 and P1 in this live run too — the same real pattern the finished, but since-archived, accidental 8-window flat-calibration run showed (kept as a historical record, not a comparison baseline — that run predates the stratified-calibration fix this exact mistake motivated).

What P0 → P1 actually means

The danger score is coded so lower means more dangerous (see "The setup, in one paragraph" above). P0 (the baseline plan, built purely from isolated KL, before this cascade ever ran) correlates negatively with it — Pearson -0.23, Spearman -0.36 — which is the sensible direction: as danger score drops (more dangerous), isolated KL rises (more real risk). The Hessian's dangerous-tensor calls agreed with the old isolated baseline. The moment the first real cascade round ran (P1), that flips to positive: Pearson +0.17, Spearman +0.16 — the backwards direction. A positive correlation here means the tensors the Hessian calls safe (spread-out curvature) are the ones showing up with higher real, compounding risk, while the ones it flags dangerous turn out comparatively less risky once cascading effects are actually measured. As of the most recent real data on this live run (P3, n=496), it reads Pearson +0.17, Spearman +0.17. (See the layer-coverage caveat below before reading anything into that sign specifically. A separate, earlier real run used an accidental 8-window flat calibration instead of this project's real N=24 stratified calibration — kept only as an archived, non-canonical historical record, not a comparison baseline for the number above; see HESSIAN_STORY_COMPLETE_ROUND.html.)

What actually changed in the bit assignment, P0 → P3

The real optimizer-written plan at each stage: how many of the 497 tensors land in each bit-tier, and the plan's own recorded raw and OptiQ-weighted KL sum. Every stage shown is a real, finished plan.

274205137680178236236238Q4130858090Q544111612Q611322Q8134162163155Q16P0P1P2P3

Each group is one bit-tier (Q4 through Q16); bars left-to-right within a group are P0, P1, P2, ... in order. Watch Q4 and Q16 (the extremes) versus Q5/Q6/Q8 (the middle) — the real, recurring pattern in this project has been the middle tiers emptying out toward both extremes once cascade measurement replaces the isolated baseline, not a uniform shift.

84.163.142.021.00.01.942.04P036.0656.17P143.9166.94P249.3473.11P3Raw ΣKLOptiQ-weighted ΣKL

Raw ΣKL asks how much error; OptiQ-weighted ΣKL asks where it landed. Both climb from P0 to P1 -- expected, since P0 uses the old isolated baseline's own scale (see the caveat on Finding 0 above) and P1+ uses the real cascaded scale.

What this real correlation does and does not prove

There is a real, separately-established finding (see this project's V6.4 competition doc) that pushing the whole model's average bit budget below roughly 4.30-4.40 BPW causes weighted ΣKL to explode even while raw ΣKL rises gently -- that is a real, validated, aggregate result about how much total compression the model can absorb before quality collapses. It does not tell you which specific tensor is "the" fragile one. The historical early-QKV floor treated a large isolated-KL reading at an early layer as proof that tensor was intrinsically fragile -- but isolated testing measures a single perturbation propagating through every remaining full-precision layer, so an early perturbation gets more downstream depth to be amplified through no matter which tensor carries it. That is a real, checkable artifact (confirmed above: the single highest isolated-KL tensor in this exact model, in both the old and new calibration data, is layer 0's own linear_attn.out_proj) -- and it is a different mechanism from cascaded compounding, not the same one. Cascaded measurement is not immune to a similar caution: it tells you real risk keeps growing with depth once earlier layers are actually quantized, but attributing that growth to any one tensor "in isolation" still overstates what an aggregate KL number can support -- much like tracing a single neuron's causal role from a whole brain's behavior change. What both signals robustly agree on: middle bit-tiers empty out toward the extremes (see the flow diagram below), and the smallest average bit budget the model tolerates before serious degradation sits close to that same 4.30-4.40 BPW boundary either way.

One specific claim worth checking rather than assuming: that BF16 protection concentrates at the network's boundaries (embedding-adjacent start, lm_head-adjacent end). In this exact plan it does not -- Q16 tensors appear at every one of the 64 layers with no positional clustering (mean layer number 31.5 for Q16 versus 30.4 for Q4, statistically indistinguishable), lm_head itself is assigned Q4, and embed_tokens is not even in the 497-tensor optimizable set to begin with (excluded by pipeline design, never a measured decision). The real, supported pattern is protection by tensor role (see Finding 2 above), not by depth position.

StageQ4Q5Q6Q8Q16Raw ΣKLWeighted ΣKL
P0 (baseline (isolated))17813044111341.9412.039
P1 (round 1)2368511316236.05656.167
P2 (round 2)2368016216343.90966.936
P3 (round 3)2389012215549.33773.109
Read the KL columns carefully

P0's raw/weighted KL sums come from the old isolated measurement's own scale; P1 onward come from the cascaded measurement's scale. These are not directly comparable in magnitude for the reason established elsewhere on this page (isolated KL structurally underrates risk, worse the deeper into the model you go) — compare bit-tier counts across stages, not the raw KL numbers, for an apples-to-apples read.

Where the bits actually travelled

Real, exact per-tensor tier transitions — not aggregate counts before/after, but which specific tensors moved from which tier to which tier between each pair of stages. Ribbon width is the real number of tensors making that exact move; same-tier ribbons (a tensor that didn't change) are drawn slightly more solid than cross-tier ones so the stable core is easy to pick out at a glance.

Watch each real transition appear one stage at a time, left to right.
Q4 36%Q5 26%Q6 9%Q16 27%P047%17%33%P147%16%33%P2Q4 48%Q5 18%Q16 31%P3Q4Q5Q6Q8Q16

Columns left to right: P0 → P1 → P2 → P3. Percentages inside each block are that tier's real share of all 497 tensors at that stage. The flowing dots are a real density cue, colored by source tier -- more dots on a ribbon means more tensors made that move -- but the count is capped (1-4 per ribbon) for readability, not a literal one-dot-per-tensor count.

Is the movement random?

No — the real numbers say the opposite. The fraction of tensors that stay in the exact same tier climbs sharply from one transition to the next (see the list above): most of the churn happens once, moving away from the old isolated-KL baseline, and then rapidly settles. That's convergence toward a stable allocation shape, not noise — consistent with what Finding 4 on the archived, accidental 8-window flat-calibration page showed for a fully finished run (kept as a historical record only, not a comparison baseline).

Where we thought it hurt: the real BPW-frontier cliff

A real, full stratified-N24 BPW sweep (V3.3 solver, Qwen3.8, 72 valid targets from 4.45 → 8.00 -- 8 lower targets down to 4.05 came back solver-infeasible under this calibration, itself real evidence for where the cliff sits). Raw ΣKL climbs gently across the whole sweep; OptiQ-weighted ΣKL is roughly flat until the low-BPW end, then rises sharply -- the real cliff, this time measured on the same calibration as everything else on this page.

STRUCTURAL DANGER ZONE8.26.24.12.10.0Our stratified P04.505.005.506.006.507.007.508.00Target BPWRaw ΣKLWeighted ΣKL
13 / 72
Q4Q5Q6Q8Q16

Purple dot: our own P0 plan (5.03 BPW) from this exact same stratified sweep family. Drag the slider (or press play) to step through the real solved plan at every real BPW target -- the pink dot and readout below always show one real, exact point from the sweep, never an interpolated value.

Why this replaced an earlier version of this chart

An earlier version of this chart reused this project's existing 74-point V3.3 sweep (V3_3_APPLES_TO_APPLES_0005) as a stand-in, with P0 marked separately since it used different calibration. That sweep was traced back to the vendor optiq convert tool's own flat calibration (concatenate-all-domains) -- confirmed by the command in archive/DRAFT_PROMPT.md and by the vendor package containing zero "stratified" code anywhere. This project had only ever run V3.3 once against the real stratified checkpoint before (the single matched P0 point) -- never the full sweep. This chart is that real sweep, generated fresh against Qwen3.8-27B-heretic-ara-OptiQ-5bpw_CLEAN_STRATIFIED/sensitivity_checkpoint.json, same solver, same validation, same output schema as the original -- see run_Qwen38_V3_3_frontier_STRATIFIED_CONTINUE_ON_FAILURE.sh.

Why the early-QKV floor and late-attention floor actually exist

Real design history, read directly from this project's own optimizer scripts (02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_2.py and _v3_3.py) -- not reconstructed from memory, and not the "solver got spooked by a scary number" story it can look like from the outside.

The root cause (documented in V3.2, from a real V3.1 regression)

Isolated-KL sensitivity rankings are measurably ~25x noisier in early layers: a 34.4% ranking-inversion rate in layers 0-15, versus 1.3% in layers 48-63 -- confirmed by two independent real sources (region_summary.json and raw inversions.csv, quartile counts 128/68/17/5). That single fact produced two different, almost opposite-looking solver failures depending on budget pressure.

V3.1 → V3.2: the late-attention floor

Under a near-zero-slack BPW ceiling, the MILP's "spend where weighted-KL improves most" objective got pulled toward early-layer attention, where the noisy, inflated deltas looked like the best return -- leaving late-layer attention thinnest of all four depth quartiles (57.9% of it at the bare 4-bit floor, vs Stock OptiQ's 40.8%). Real, measured cost: V3.1 scored worst of three compared models on IFEval (77.5% vs Stock's 82.5%) despite winning everywhere else. Fix: an explicit floor over ALL attention-family tensors in the last 25% of depth.

V3.2 → V3.3: the early-QKV floor

In a matched-BPW dry run (2026-08-30), given a slightly larger budget, the solver demoted layers.2.linear_attn.in_proj_qkv (52.4M params, 8-bit→6-bit) to fund upgrading layers.9.linear_attn.out_proj (31.5M params, 5-bit→16-bit) plus an MLP tensor. Both source and destination sat inside the same noisy 0-15 region, and an input projection (feeds Q, K, and V at once) has a structurally larger blast radius than an output projection (a single downstream transform) -- backwards trade on both grounds. Real, measured cost: 86.8% MMLU, below both V6 (89.5%) and Stock5 (88.6%). Fix: a floor scoped ONLY to input-projection tensors (not outputs) in the first 25% of depth.

The mechanistic piece the docstrings don't explain

Both floors were built on a solid, measured root cause: early-region isolated-KL rankings are far noisier. What isn't documented anywhere in this project is why. This investigation found a real, checkable mechanism: isolated testing perturbs one tensor and leaves every other layer at full BF16, so an early perturbation has far more downstream, untouched, nonlinear depth left to be amplified through than a late one -- confirmed directly (the single highest isolated-KL tensor in this exact model, in both the old flat and new stratified calibration data, is layer 0's own linear_attn.out_proj, roughly 7-12x the median). If a tensor's isolated-KL score is substantially driven by how much amplification runway is left downstream rather than its own intrinsic fragility, ranking tensors against each other within the early region would naturally scramble far more often than in late layers, where there is almost no runway left to distort the ranking. That is a plausible causal explanation for the 34.4%-vs-1.3% split, not a replacement for it.

What a genuinely better solver would do

Three real, separate findings from tonight, and what each one is actually good evidence for -- not blended into one story, since that is exactly the mistake this whole investigation keeps finding and fixing.

1. The BPW-frontier cliff (real, aggregate, model-level)

Below roughly 4.30-4.40 BPW, weighted ΣKL explodes while raw ΣKL barely moves. This is solid evidence for a global budget guardrail -- refuse or flag any plan targeting below that boundary without an explicit override -- not evidence about any individual tensor.

2. Isolated-KL is structurally unreliable early, cascaded-KL is not the same signal

The two regional floors are hand-tuned patches over a known-unreliable objective. The real fix available now that wasn't available when V3.2/V3.3 were built: replace isolated-KL-driven regional floors with an objective informed by real cascaded measurement, which this investigation has shown tracks real depth-dependent risk directly (the isolated/cascaded gap grows from ~1.2x at layer 0 to ~500x by layer 60) instead of an artifact of amplification distance. A solver reading real compounding risk should not need a hand-coded "first 25%" or "last 25%" region at all -- the real risk shape should emerge from the objective itself.

3. Real protection concentrates by tensor role, not by depth position

Q16 tensors sit at all 64 layers with statistically identical mean depth to Q4 tensors; self_attn.q_proj is dangerous everywhere, self_attn.k_proj is safe everywhere. A genuinely better solver's weighting should be keyed to role and real cascaded risk, not to a fixed depth fraction -- the current early/late floors are a reasonable stopgap for a real problem, not the correct long-term shape of the fix.

What success looks like, concretely

A solver that needs zero hand-tuned regional floors because its own objective already reflects real cascaded risk. Testable, not aspirational: re-run the two historical incidents that motivated V3.2 and V3.3 -- the V3.1 matched-budget IFEval drop and the V3.2 matched-BPW MMLU drop -- under a cascaded-informed objective with both floors disabled, and check whether the same bad trades still get made. If they don't, the floors become redundant rather than load-bearing, and can be removed or kept purely as a cheap safety net rather than the primary correction. If they still get made, that is real, useful evidence the depth-position floors are catching something the role/cascade signal alone still misses -- also a real answer, just the other one.

The Hessian landscape

Every real, currently-measured tensor's danger score, laid out by layer (left-to-right) and tensor role (top-to-bottom) as a flat grid, cell color encoding the real number. This is the actual per-tensor Hessian effective-rank signal this whole page is built on — not a rendering of the raw matrices themselves (those were never saved to disk, only this derived summary), but the real, live grid those matrices produced.

self attn q projlinear attn in proj qkvlinear attn in proj zself attn o projmlp up projself attn v projmlp gate projmlp down projlinear attn out projself attn k projlinear attn in proj alinear attn in proj b04812162024283236404448525660layerSaferMore dangerous

Red cells = narrow, dangerous weak spots. Purple cells = safe, forgiving tensors. Rows are sorted worst-to-safest, top to bottom — the same ordering as Finding 2's bar chart. A flat grid, deliberately not 3D: cell color is the only encoding, so every cell is directly comparable at a glance.

Finding 1 — the free number really does predict real risk

Every dot below is one real tensor. Left-to-right is the free danger score (from Hessian logs you already have). Up-and-down is the real, expensive, ground-truth sensitivity from the cascaded measurement. If the free number were useless, this would look like a random cloud.

Why there is more than one round

This measurement runs in rounds because a single cascaded pass measures every tensor against a specific upstream context -- which bit-width each earlier tensor is assumed to already have. The real bit-width plan computed FROM that measurement can end up disagreeing with the context it was measured against. When it does, a new round automatically starts, re-measuring using the new plan as the updated context, and stops once a round's own output agrees with what it assumed (or a round cap is hit). Every round this run has completed or is currently working through is shown below, separately -- so you can see for yourself whether later rounds broadly agree with earlier ones (converging) or keep shifting (not yet converged), instead of only ever seeing the most recent one.

Round 1 — complete

497/497 tensors measured this round · 496 have both signals · Pearson r=+0.17, Spearman r=+0.16

0.00030.0010.0030.010.030.100.0000.0540.1080.1620.216L63.mlp.up_projL63.mlp.down_projL63.mlp.gate_projL63.self_attn.o_projL63.self_attn.v_projL63.self_attn.k_projL63.self_attn.q_projL62.mlp.up_projL62.mlp.down_projL62.mlp.gate_projL62.linear_attn.out_projL62.linear_attn.in_proj_aL62.linear_attn.in_proj_bL62.linear_attn.in_proj_zL62.linear_attn.in_proj_qkvL61.mlp.up_projL61.mlp.down_projL61.mlp.gate_projL61.linear_attn.out_projL61.linear_attn.in_proj_aL61.linear_attn.in_proj_bL61.linear_attn.in_proj_zL61.linear_attn.in_proj_qkvL60.mlp.up_projL60.mlp.down_projL60.mlp.gate_projL60.linear_attn.out_projL60.linear_attn.in_proj_aL60.linear_attn.in_proj_bL60.linear_attn.in_proj_zL60.linear_attn.in_proj_qkvL59.mlp.up_projL59.mlp.down_projL59.mlp.gate_projL59.self_attn.o_projL59.self_attn.v_projL59.self_attn.k_projL59.self_attn.q_projL58.mlp.up_projL58.mlp.down_projL58.mlp.gate_projL58.linear_attn.out_projL58.linear_attn.in_proj_aL58.linear_attn.in_proj_bL58.linear_attn.in_proj_zL58.linear_attn.in_proj_qkvL57.mlp.up_projL57.mlp.down_projL57.mlp.gate_projL57.linear_attn.out_projL57.linear_attn.in_proj_aL57.linear_attn.in_proj_bL57.linear_attn.in_proj_zL57.linear_attn.in_proj_qkvL56.mlp.up_projL56.mlp.down_projL56.mlp.gate_projL56.linear_attn.out_projL56.linear_attn.in_proj_aL56.linear_attn.in_proj_bL56.linear_attn.in_proj_zL56.linear_attn.in_proj_qkvL55.mlp.up_projL55.mlp.down_projL55.mlp.gate_projL55.self_attn.o_projL55.self_attn.v_projL55.self_attn.k_projL55.self_attn.q_projL54.mlp.up_projL54.mlp.down_projL54.mlp.gate_projL54.linear_attn.out_projL54.linear_attn.in_proj_aL54.linear_attn.in_proj_bL54.linear_attn.in_proj_zL54.linear_attn.in_proj_qkvL53.mlp.up_projL53.mlp.down_projL53.mlp.gate_projL53.linear_attn.out_projL53.linear_attn.in_proj_aL53.linear_attn.in_proj_bL53.linear_attn.in_proj_zL53.linear_attn.in_proj_qkvL52.mlp.up_projL52.mlp.down_projL52.mlp.gate_projL52.linear_attn.out_projL52.linear_attn.in_proj_aL52.linear_attn.in_proj_bL52.linear_attn.in_proj_zL52.linear_attn.in_proj_qkvL51.mlp.up_projL51.mlp.down_projL51.mlp.gate_projL51.self_attn.o_projL51.self_attn.v_projL51.self_attn.k_projL51.self_attn.q_projL50.mlp.up_projL50.mlp.down_projL50.mlp.gate_projL50.linear_attn.out_projL50.linear_attn.in_proj_aL50.linear_attn.in_proj_bL50.linear_attn.in_proj_zL50.linear_attn.in_proj_qkvL49.mlp.up_projL49.mlp.down_projL49.mlp.gate_projL49.linear_attn.out_projL49.linear_attn.in_proj_aL49.linear_attn.in_proj_bL49.linear_attn.in_proj_zL49.linear_attn.in_proj_qkvL48.mlp.up_projL48.mlp.down_projL48.mlp.gate_projL48.linear_attn.out_projL48.linear_attn.in_proj_aL48.linear_attn.in_proj_bL48.linear_attn.in_proj_zL48.linear_attn.in_proj_qkvL47.mlp.up_projL47.mlp.down_projL47.mlp.gate_projL47.self_attn.o_projL47.self_attn.v_projL47.self_attn.k_projL47.self_attn.q_projL46.mlp.up_projL46.mlp.down_projL46.mlp.gate_projL46.linear_attn.out_projL46.linear_attn.in_proj_aL46.linear_attn.in_proj_bL46.linear_attn.in_proj_zL46.linear_attn.in_proj_qkvL45.mlp.up_projL45.mlp.down_projL45.mlp.gate_projL45.linear_attn.out_projL45.linear_attn.in_proj_aL45.linear_attn.in_proj_bL45.linear_attn.in_proj_zL45.linear_attn.in_proj_qkvL44.mlp.up_projL44.mlp.down_projL44.mlp.gate_projL44.linear_attn.out_projL44.linear_attn.in_proj_aL44.linear_attn.in_proj_bL44.linear_attn.in_proj_zL44.linear_attn.in_proj_qkvL43.mlp.up_projL43.mlp.down_projL43.mlp.gate_projL43.self_attn.o_projL43.self_attn.v_projL43.self_attn.k_projL43.self_attn.q_projL42.mlp.up_projL42.mlp.down_projL42.mlp.gate_projL42.linear_attn.out_projL42.linear_attn.in_proj_aL42.linear_attn.in_proj_bL42.linear_attn.in_proj_zL42.linear_attn.in_proj_qkvL41.mlp.up_projL41.mlp.down_projL41.mlp.gate_projL41.linear_attn.out_projL41.linear_attn.in_proj_aL41.linear_attn.in_proj_bL41.linear_attn.in_proj_zL41.linear_attn.in_proj_qkvL40.mlp.up_projL40.mlp.down_projL40.mlp.gate_projL40.linear_attn.out_projL40.linear_attn.in_proj_aL40.linear_attn.in_proj_bL40.linear_attn.in_proj_zL40.linear_attn.in_proj_qkvL39.mlp.up_projL39.mlp.down_projL39.mlp.gate_projL39.self_attn.o_projL39.self_attn.v_projL39.self_attn.k_projL39.self_attn.q_projL38.mlp.up_projL38.mlp.down_projL38.mlp.gate_projL38.linear_attn.out_projL38.linear_attn.in_proj_aL38.linear_attn.in_proj_bL38.linear_attn.in_proj_zL38.linear_attn.in_proj_qkvL37.mlp.up_projL37.mlp.down_projL37.mlp.gate_projL37.linear_attn.out_projL37.linear_attn.in_proj_aL37.linear_attn.in_proj_bL37.linear_attn.in_proj_zL37.linear_attn.in_proj_qkvL36.mlp.up_projL36.mlp.down_projL36.mlp.gate_projL36.linear_attn.out_projL36.linear_attn.in_proj_aL36.linear_attn.in_proj_bL36.linear_attn.in_proj_zL36.linear_attn.in_proj_qkvL35.mlp.up_projL35.mlp.down_projL35.mlp.gate_projL35.self_attn.o_projL35.self_attn.v_projL35.self_attn.k_projL35.self_attn.q_projL34.mlp.up_projL34.mlp.down_projL34.mlp.gate_projL34.linear_attn.out_projL34.linear_attn.in_proj_aL34.linear_attn.in_proj_bL34.linear_attn.in_proj_zL34.linear_attn.in_proj_qkvL33.mlp.up_projL33.mlp.down_projL33.mlp.gate_projL33.linear_attn.out_projL33.linear_attn.in_proj_aL33.linear_attn.in_proj_bL33.linear_attn.in_proj_zL33.linear_attn.in_proj_qkvL32.mlp.up_projL32.mlp.down_projL32.mlp.gate_projL32.linear_attn.out_projL32.linear_attn.in_proj_aL32.linear_attn.in_proj_bL32.linear_attn.in_proj_zL32.linear_attn.in_proj_qkvL31.mlp.up_projL31.mlp.down_projL31.mlp.gate_projL31.self_attn.o_projL31.self_attn.v_projL31.self_attn.k_projL31.self_attn.q_projL30.mlp.up_projL30.mlp.down_projL30.mlp.gate_projL30.linear_attn.out_projL30.linear_attn.in_proj_aL30.linear_attn.in_proj_bL30.linear_attn.in_proj_zL30.linear_attn.in_proj_qkvL29.mlp.up_projL29.mlp.down_projL29.mlp.gate_projL29.linear_attn.out_projL29.linear_attn.in_proj_aL29.linear_attn.in_proj_bL29.linear_attn.in_proj_zL29.linear_attn.in_proj_qkvL28.mlp.up_projL28.mlp.down_projL28.mlp.gate_projL28.linear_attn.out_projL28.linear_attn.in_proj_aL28.linear_attn.in_proj_bL28.linear_attn.in_proj_zL28.linear_attn.in_proj_qkvL27.mlp.up_projL27.mlp.down_projL27.mlp.gate_projL27.self_attn.o_projL27.self_attn.v_projL27.self_attn.k_projL27.self_attn.q_projL26.mlp.up_projL26.mlp.down_projL26.mlp.gate_projL26.linear_attn.out_projL26.linear_attn.in_proj_aL26.linear_attn.in_proj_bL26.linear_attn.in_proj_zL26.linear_attn.in_proj_qkvL25.mlp.up_projL25.mlp.down_projL25.mlp.gate_projL25.linear_attn.out_projL25.linear_attn.in_proj_aL25.linear_attn.in_proj_bL25.linear_attn.in_proj_zL25.linear_attn.in_proj_qkvL24.mlp.up_projL24.mlp.down_projL24.mlp.gate_projL24.linear_attn.out_projL24.linear_attn.in_proj_aL24.linear_attn.in_proj_bL24.linear_attn.in_proj_zL24.linear_attn.in_proj_qkvL23.mlp.up_projL23.mlp.down_projL23.mlp.gate_projL23.self_attn.o_projL23.self_attn.v_projL23.self_attn.k_projL23.self_attn.q_projL22.mlp.up_projL22.mlp.down_projL22.mlp.gate_projL22.linear_attn.out_projL22.linear_attn.in_proj_aL22.linear_attn.in_proj_bL22.linear_attn.in_proj_zL22.linear_attn.in_proj_qkvL21.mlp.up_projL21.mlp.down_projL21.mlp.gate_projL21.linear_attn.out_projL21.linear_attn.in_proj_aL21.linear_attn.in_proj_bL21.linear_attn.in_proj_zL21.linear_attn.in_proj_qkvL20.mlp.up_projL20.mlp.down_projL20.mlp.gate_projL20.linear_attn.out_projL20.linear_attn.in_proj_aL20.linear_attn.in_proj_bL20.linear_attn.in_proj_zL20.linear_attn.in_proj_qkvL19.mlp.up_projL19.mlp.down_projL19.mlp.gate_projL19.self_attn.o_projL19.self_attn.v_projL19.self_attn.k_projL19.self_attn.q_projL18.mlp.up_projL18.mlp.down_projL18.mlp.gate_projL18.linear_attn.out_projL18.linear_attn.in_proj_aL18.linear_attn.in_proj_bL18.linear_attn.in_proj_zL18.linear_attn.in_proj_qkvL17.mlp.up_projL17.mlp.down_projL17.mlp.gate_projL17.linear_attn.out_projL17.linear_attn.in_proj_aL17.linear_attn.in_proj_bL17.linear_attn.in_proj_zL17.linear_attn.in_proj_qkvL16.mlp.up_projL16.mlp.down_projL16.mlp.gate_projL16.linear_attn.out_projL16.linear_attn.in_proj_aL16.linear_attn.in_proj_bL16.linear_attn.in_proj_zL16.linear_attn.in_proj_qkvL15.mlp.up_projL15.mlp.down_projL15.mlp.gate_projL15.self_attn.o_projL15.self_attn.v_projL15.self_attn.k_projL15.self_attn.q_projL14.mlp.up_projL14.mlp.down_projL14.mlp.gate_projL14.linear_attn.out_projL14.linear_attn.in_proj_aL14.linear_attn.in_proj_bL14.linear_attn.in_proj_zL14.linear_attn.in_proj_qkvL13.mlp.up_projL13.mlp.down_projL13.mlp.gate_projL13.linear_attn.out_projL13.linear_attn.in_proj_aL13.linear_attn.in_proj_bL13.linear_attn.in_proj_zL13.linear_attn.in_proj_qkvL12.mlp.up_projL12.mlp.down_projL12.mlp.gate_projL12.linear_attn.out_projL12.linear_attn.in_proj_aL12.linear_attn.in_proj_bL12.linear_attn.in_proj_zL12.linear_attn.in_proj_qkvL11.mlp.up_projL11.mlp.down_projL11.mlp.gate_projL11.self_attn.o_projL11.self_attn.v_projL11.self_attn.k_projL11.self_attn.q_projL10.mlp.up_projL10.mlp.down_projL10.mlp.gate_projL10.linear_attn.out_projL10.linear_attn.in_proj_aL10.linear_attn.in_proj_bL10.linear_attn.in_proj_zL10.linear_attn.in_proj_qkvL9.mlp.up_projL9.mlp.down_projL9.mlp.gate_projL9.linear_attn.out_projL9.linear_attn.in_proj_aL9.linear_attn.in_proj_bL9.linear_attn.in_proj_zL9.linear_attn.in_proj_qkvL8.mlp.up_projL8.mlp.down_projL8.mlp.gate_projL8.linear_attn.out_projL8.linear_attn.in_proj_aL8.linear_attn.in_proj_bL8.linear_attn.in_proj_zL8.linear_attn.in_proj_qkvL7.mlp.up_projL7.mlp.down_projL7.mlp.gate_projL7.self_attn.o_projL7.self_attn.v_projL7.self_attn.k_projL7.self_attn.q_projL6.mlp.up_projL6.mlp.down_projL6.mlp.gate_projL6.linear_attn.out_projL6.linear_attn.in_proj_aL6.linear_attn.in_proj_bL6.linear_attn.in_proj_zL6.linear_attn.in_proj_qkvL5.mlp.up_projL5.mlp.down_projL5.mlp.gate_projL5.linear_attn.out_projL5.linear_attn.in_proj_aL5.linear_attn.in_proj_bL5.linear_attn.in_proj_zL5.linear_attn.in_proj_qkvL4.mlp.up_projL4.mlp.down_projL4.mlp.gate_projL4.linear_attn.out_projL4.linear_attn.in_proj_aL4.linear_attn.in_proj_bL4.linear_attn.in_proj_zL4.linear_attn.in_proj_qkvL3.mlp.up_projL3.mlp.down_projL3.mlp.gate_projL3.self_attn.o_projL3.self_attn.v_projL3.self_attn.k_projL3.self_attn.q_projL2.mlp.up_projL2.mlp.down_projL2.mlp.gate_projL2.linear_attn.out_projL2.linear_attn.in_proj_aL2.linear_attn.in_proj_bL2.linear_attn.in_proj_zL2.linear_attn.in_proj_qkvL1.mlp.up_projL1.mlp.down_projL1.mlp.gate_projL1.linear_attn.out_projL1.linear_attn.in_proj_aL1.linear_attn.in_proj_bL1.linear_attn.in_proj_zL1.linear_attn.in_proj_qkvL0.mlp.up_projL0.mlp.down_projL0.mlp.gate_projL0.linear_attn.out_projL0.linear_attn.in_proj_aL0.linear_attn.in_proj_bL0.linear_attn.in_proj_zL0.linear_attn.in_proj_qkvDanger score (lower = worse) — log scaleReal sensitivity (higher = worse)

Each dot = one real tensor measured in round 1 specifically (n=496).

Query / attention-input family Key / attention-output family Everything else (MLP, V)
496 / 496
Drag to replay the real order these 496 tensors were actually measured in — layer 0 through the final layer, then lm_head. Not simulated: this is the true capture sequence.

Round 2 — complete

497/497 tensors measured this round · 496 have both signals · Pearson r=+0.17, Spearman r=+0.17

0.00030.0010.0030.010.030.100.0000.0620.1240.1870.249L63.mlp.up_projL63.mlp.down_projL63.mlp.gate_projL63.self_attn.o_projL63.self_attn.v_projL63.self_attn.k_projL63.self_attn.q_projL62.mlp.up_projL62.mlp.down_projL62.mlp.gate_projL62.linear_attn.out_projL62.linear_attn.in_proj_aL62.linear_attn.in_proj_bL62.linear_attn.in_proj_zL62.linear_attn.in_proj_qkvL61.mlp.up_projL61.mlp.down_projL61.mlp.gate_projL61.linear_attn.out_projL61.linear_attn.in_proj_aL61.linear_attn.in_proj_bL61.linear_attn.in_proj_zL61.linear_attn.in_proj_qkvL60.mlp.up_projL60.mlp.down_projL60.mlp.gate_projL60.linear_attn.out_projL60.linear_attn.in_proj_aL60.linear_attn.in_proj_bL60.linear_attn.in_proj_zL60.linear_attn.in_proj_qkvL59.mlp.up_projL59.mlp.down_projL59.mlp.gate_projL59.self_attn.o_projL59.self_attn.v_projL59.self_attn.k_projL59.self_attn.q_projL58.mlp.up_projL58.mlp.down_projL58.mlp.gate_projL58.linear_attn.out_projL58.linear_attn.in_proj_aL58.linear_attn.in_proj_bL58.linear_attn.in_proj_zL58.linear_attn.in_proj_qkvL57.mlp.up_projL57.mlp.down_projL57.mlp.gate_projL57.linear_attn.out_projL57.linear_attn.in_proj_aL57.linear_attn.in_proj_bL57.linear_attn.in_proj_zL57.linear_attn.in_proj_qkvL56.mlp.up_projL56.mlp.down_projL56.mlp.gate_projL56.linear_attn.out_projL56.linear_attn.in_proj_aL56.linear_attn.in_proj_bL56.linear_attn.in_proj_zL56.linear_attn.in_proj_qkvL55.mlp.up_projL55.mlp.down_projL55.mlp.gate_projL55.self_attn.o_projL55.self_attn.v_projL55.self_attn.k_projL55.self_attn.q_projL54.mlp.up_projL54.mlp.down_projL54.mlp.gate_projL54.linear_attn.out_projL54.linear_attn.in_proj_aL54.linear_attn.in_proj_bL54.linear_attn.in_proj_zL54.linear_attn.in_proj_qkvL53.mlp.up_projL53.mlp.down_projL53.mlp.gate_projL53.linear_attn.out_projL53.linear_attn.in_proj_aL53.linear_attn.in_proj_bL53.linear_attn.in_proj_zL53.linear_attn.in_proj_qkvL52.mlp.up_projL52.mlp.down_projL52.mlp.gate_projL52.linear_attn.out_projL52.linear_attn.in_proj_aL52.linear_attn.in_proj_bL52.linear_attn.in_proj_zL52.linear_attn.in_proj_qkvL51.mlp.up_projL51.mlp.down_projL51.mlp.gate_projL51.self_attn.o_projL51.self_attn.v_projL51.self_attn.k_projL51.self_attn.q_projL50.mlp.up_projL50.mlp.down_projL50.mlp.gate_projL50.linear_attn.out_projL50.linear_attn.in_proj_aL50.linear_attn.in_proj_bL50.linear_attn.in_proj_zL50.linear_attn.in_proj_qkvL49.mlp.up_projL49.mlp.down_projL49.mlp.gate_projL49.linear_attn.out_projL49.linear_attn.in_proj_aL49.linear_attn.in_proj_bL49.linear_attn.in_proj_zL49.linear_attn.in_proj_qkvL48.mlp.up_projL48.mlp.down_projL48.mlp.gate_projL48.linear_attn.out_projL48.linear_attn.in_proj_aL48.linear_attn.in_proj_bL48.linear_attn.in_proj_zL48.linear_attn.in_proj_qkvL47.mlp.up_projL47.mlp.down_projL47.mlp.gate_projL47.self_attn.o_projL47.self_attn.v_projL47.self_attn.k_projL47.self_attn.q_projL46.mlp.up_projL46.mlp.down_projL46.mlp.gate_projL46.linear_attn.out_projL46.linear_attn.in_proj_aL46.linear_attn.in_proj_bL46.linear_attn.in_proj_zL46.linear_attn.in_proj_qkvL45.mlp.up_projL45.mlp.down_projL45.mlp.gate_projL45.linear_attn.out_projL45.linear_attn.in_proj_aL45.linear_attn.in_proj_bL45.linear_attn.in_proj_zL45.linear_attn.in_proj_qkvL44.mlp.up_projL44.mlp.down_projL44.mlp.gate_projL44.linear_attn.out_projL44.linear_attn.in_proj_aL44.linear_attn.in_proj_bL44.linear_attn.in_proj_zL44.linear_attn.in_proj_qkvL43.mlp.up_projL43.mlp.down_projL43.mlp.gate_projL43.self_attn.o_projL43.self_attn.v_projL43.self_attn.k_projL43.self_attn.q_projL42.mlp.up_projL42.mlp.down_projL42.mlp.gate_projL42.linear_attn.out_projL42.linear_attn.in_proj_aL42.linear_attn.in_proj_bL42.linear_attn.in_proj_zL42.linear_attn.in_proj_qkvL41.mlp.up_projL41.mlp.down_projL41.mlp.gate_projL41.linear_attn.out_projL41.linear_attn.in_proj_aL41.linear_attn.in_proj_bL41.linear_attn.in_proj_zL41.linear_attn.in_proj_qkvL40.mlp.up_projL40.mlp.down_projL40.mlp.gate_projL40.linear_attn.out_projL40.linear_attn.in_proj_aL40.linear_attn.in_proj_bL40.linear_attn.in_proj_zL40.linear_attn.in_proj_qkvL39.mlp.up_projL39.mlp.down_projL39.mlp.gate_projL39.self_attn.o_projL39.self_attn.v_projL39.self_attn.k_projL39.self_attn.q_projL38.mlp.up_projL38.mlp.down_projL38.mlp.gate_projL38.linear_attn.out_projL38.linear_attn.in_proj_aL38.linear_attn.in_proj_bL38.linear_attn.in_proj_zL38.linear_attn.in_proj_qkvL37.mlp.up_projL37.mlp.down_projL37.mlp.gate_projL37.linear_attn.out_projL37.linear_attn.in_proj_aL37.linear_attn.in_proj_bL37.linear_attn.in_proj_zL37.linear_attn.in_proj_qkvL36.mlp.up_projL36.mlp.down_projL36.mlp.gate_projL36.linear_attn.out_projL36.linear_attn.in_proj_aL36.linear_attn.in_proj_bL36.linear_attn.in_proj_zL36.linear_attn.in_proj_qkvL35.mlp.up_projL35.mlp.down_projL35.mlp.gate_projL35.self_attn.o_projL35.self_attn.v_projL35.self_attn.k_projL35.self_attn.q_projL34.mlp.up_projL34.mlp.down_projL34.mlp.gate_projL34.linear_attn.out_projL34.linear_attn.in_proj_aL34.linear_attn.in_proj_bL34.linear_attn.in_proj_zL34.linear_attn.in_proj_qkvL33.mlp.up_projL33.mlp.down_projL33.mlp.gate_projL33.linear_attn.out_projL33.linear_attn.in_proj_aL33.linear_attn.in_proj_bL33.linear_attn.in_proj_zL33.linear_attn.in_proj_qkvL32.mlp.up_projL32.mlp.down_projL32.mlp.gate_projL32.linear_attn.out_projL32.linear_attn.in_proj_aL32.linear_attn.in_proj_bL32.linear_attn.in_proj_zL32.linear_attn.in_proj_qkvL31.mlp.up_projL31.mlp.down_projL31.mlp.gate_projL31.self_attn.o_projL31.self_attn.v_projL31.self_attn.k_projL31.self_attn.q_projL30.mlp.up_projL30.mlp.down_projL30.mlp.gate_projL30.linear_attn.out_projL30.linear_attn.in_proj_aL30.linear_attn.in_proj_bL30.linear_attn.in_proj_zL30.linear_attn.in_proj_qkvL29.mlp.up_projL29.mlp.down_projL29.mlp.gate_projL29.linear_attn.out_projL29.linear_attn.in_proj_aL29.linear_attn.in_proj_bL29.linear_attn.in_proj_zL29.linear_attn.in_proj_qkvL28.mlp.up_projL28.mlp.down_projL28.mlp.gate_projL28.linear_attn.out_projL28.linear_attn.in_proj_aL28.linear_attn.in_proj_bL28.linear_attn.in_proj_zL28.linear_attn.in_proj_qkvL27.mlp.up_projL27.mlp.down_projL27.mlp.gate_projL27.self_attn.o_projL27.self_attn.v_projL27.self_attn.k_projL27.self_attn.q_projL26.mlp.up_projL26.mlp.down_projL26.mlp.gate_projL26.linear_attn.out_projL26.linear_attn.in_proj_aL26.linear_attn.in_proj_bL26.linear_attn.in_proj_zL26.linear_attn.in_proj_qkvL25.mlp.up_projL25.mlp.down_projL25.mlp.gate_projL25.linear_attn.out_projL25.linear_attn.in_proj_aL25.linear_attn.in_proj_bL25.linear_attn.in_proj_zL25.linear_attn.in_proj_qkvL24.mlp.up_projL24.mlp.down_projL24.mlp.gate_projL24.linear_attn.out_projL24.linear_attn.in_proj_aL24.linear_attn.in_proj_bL24.linear_attn.in_proj_zL24.linear_attn.in_proj_qkvL23.mlp.up_projL23.mlp.down_projL23.mlp.gate_projL23.self_attn.o_projL23.self_attn.v_projL23.self_attn.k_projL23.self_attn.q_projL22.mlp.up_projL22.mlp.down_projL22.mlp.gate_projL22.linear_attn.out_projL22.linear_attn.in_proj_aL22.linear_attn.in_proj_bL22.linear_attn.in_proj_zL22.linear_attn.in_proj_qkvL21.mlp.up_projL21.mlp.down_projL21.mlp.gate_projL21.linear_attn.out_projL21.linear_attn.in_proj_aL21.linear_attn.in_proj_bL21.linear_attn.in_proj_zL21.linear_attn.in_proj_qkvL20.mlp.up_projL20.mlp.down_projL20.mlp.gate_projL20.linear_attn.out_projL20.linear_attn.in_proj_aL20.linear_attn.in_proj_bL20.linear_attn.in_proj_zL20.linear_attn.in_proj_qkvL19.mlp.up_projL19.mlp.down_projL19.mlp.gate_projL19.self_attn.o_projL19.self_attn.v_projL19.self_attn.k_projL19.self_attn.q_projL18.mlp.up_projL18.mlp.down_projL18.mlp.gate_projL18.linear_attn.out_projL18.linear_attn.in_proj_aL18.linear_attn.in_proj_bL18.linear_attn.in_proj_zL18.linear_attn.in_proj_qkvL17.mlp.up_projL17.mlp.down_projL17.mlp.gate_projL17.linear_attn.out_projL17.linear_attn.in_proj_aL17.linear_attn.in_proj_bL17.linear_attn.in_proj_zL17.linear_attn.in_proj_qkvL16.mlp.up_projL16.mlp.down_projL16.mlp.gate_projL16.linear_attn.out_projL16.linear_attn.in_proj_aL16.linear_attn.in_proj_bL16.linear_attn.in_proj_zL16.linear_attn.in_proj_qkvL15.mlp.up_projL15.mlp.down_projL15.mlp.gate_projL15.self_attn.o_projL15.self_attn.v_projL15.self_attn.k_projL15.self_attn.q_projL14.mlp.up_projL14.mlp.down_projL14.mlp.gate_projL14.linear_attn.out_projL14.linear_attn.in_proj_aL14.linear_attn.in_proj_bL14.linear_attn.in_proj_zL14.linear_attn.in_proj_qkvL13.mlp.up_projL13.mlp.down_projL13.mlp.gate_projL13.linear_attn.out_projL13.linear_attn.in_proj_aL13.linear_attn.in_proj_bL13.linear_attn.in_proj_zL13.linear_attn.in_proj_qkvL12.mlp.up_projL12.mlp.down_projL12.mlp.gate_projL12.linear_attn.out_projL12.linear_attn.in_proj_aL12.linear_attn.in_proj_bL12.linear_attn.in_proj_zL12.linear_attn.in_proj_qkvL11.mlp.up_projL11.mlp.down_projL11.mlp.gate_projL11.self_attn.o_projL11.self_attn.v_projL11.self_attn.k_projL11.self_attn.q_projL10.mlp.up_projL10.mlp.down_projL10.mlp.gate_projL10.linear_attn.out_projL10.linear_attn.in_proj_aL10.linear_attn.in_proj_bL10.linear_attn.in_proj_zL10.linear_attn.in_proj_qkvL9.mlp.up_projL9.mlp.down_projL9.mlp.gate_projL9.linear_attn.out_projL9.linear_attn.in_proj_aL9.linear_attn.in_proj_bL9.linear_attn.in_proj_zL9.linear_attn.in_proj_qkvL8.mlp.up_projL8.mlp.down_projL8.mlp.gate_projL8.linear_attn.out_projL8.linear_attn.in_proj_aL8.linear_attn.in_proj_bL8.linear_attn.in_proj_zL8.linear_attn.in_proj_qkvL7.mlp.up_projL7.mlp.down_projL7.mlp.gate_projL7.self_attn.o_projL7.self_attn.v_projL7.self_attn.k_projL7.self_attn.q_projL6.mlp.up_projL6.mlp.down_projL6.mlp.gate_projL6.linear_attn.out_projL6.linear_attn.in_proj_aL6.linear_attn.in_proj_bL6.linear_attn.in_proj_zL6.linear_attn.in_proj_qkvL5.mlp.up_projL5.mlp.down_projL5.mlp.gate_projL5.linear_attn.out_projL5.linear_attn.in_proj_aL5.linear_attn.in_proj_bL5.linear_attn.in_proj_zL5.linear_attn.in_proj_qkvL4.mlp.up_projL4.mlp.down_projL4.mlp.gate_projL4.linear_attn.out_projL4.linear_attn.in_proj_aL4.linear_attn.in_proj_bL4.linear_attn.in_proj_zL4.linear_attn.in_proj_qkvL3.mlp.up_projL3.mlp.down_projL3.mlp.gate_projL3.self_attn.o_projL3.self_attn.v_projL3.self_attn.k_projL3.self_attn.q_projL2.mlp.up_projL2.mlp.down_projL2.mlp.gate_projL2.linear_attn.out_projL2.linear_attn.in_proj_aL2.linear_attn.in_proj_bL2.linear_attn.in_proj_zL2.linear_attn.in_proj_qkvL1.mlp.up_projL1.mlp.down_projL1.mlp.gate_projL1.linear_attn.out_projL1.linear_attn.in_proj_aL1.linear_attn.in_proj_bL1.linear_attn.in_proj_zL1.linear_attn.in_proj_qkvL0.mlp.up_projL0.mlp.down_projL0.mlp.gate_projL0.linear_attn.out_projL0.linear_attn.in_proj_aL0.linear_attn.in_proj_bL0.linear_attn.in_proj_zL0.linear_attn.in_proj_qkvDanger score (lower = worse) — log scaleReal sensitivity (higher = worse)

Each dot = one real tensor measured in round 2 specifically (n=496).

Query / attention-input family Key / attention-output family Everything else (MLP, V)
496 / 496
Drag to replay the real order these 496 tensors were actually measured in — layer 0 through the final layer, then lm_head. Not simulated: this is the true capture sequence.

Round 3 — complete

497/497 tensors measured this round · 496 have both signals · Pearson r=+0.17, Spearman r=+0.17

0.00030.0010.0030.010.030.100.0000.0660.1310.1970.262L63.mlp.up_projL63.mlp.down_projL63.mlp.gate_projL63.self_attn.o_projL63.self_attn.v_projL63.self_attn.k_projL63.self_attn.q_projL62.mlp.up_projL62.mlp.down_projL62.mlp.gate_projL62.linear_attn.out_projL62.linear_attn.in_proj_aL62.linear_attn.in_proj_bL62.linear_attn.in_proj_zL62.linear_attn.in_proj_qkvL61.mlp.up_projL61.mlp.down_projL61.mlp.gate_projL61.linear_attn.out_projL61.linear_attn.in_proj_aL61.linear_attn.in_proj_bL61.linear_attn.in_proj_zL61.linear_attn.in_proj_qkvL60.mlp.up_projL60.mlp.down_projL60.mlp.gate_projL60.linear_attn.out_projL60.linear_attn.in_proj_aL60.linear_attn.in_proj_bL60.linear_attn.in_proj_zL60.linear_attn.in_proj_qkvL59.mlp.up_projL59.mlp.down_projL59.mlp.gate_projL59.self_attn.o_projL59.self_attn.v_projL59.self_attn.k_projL59.self_attn.q_projL58.mlp.up_projL58.mlp.down_projL58.mlp.gate_projL58.linear_attn.out_projL58.linear_attn.in_proj_aL58.linear_attn.in_proj_bL58.linear_attn.in_proj_zL58.linear_attn.in_proj_qkvL57.mlp.up_projL57.mlp.down_projL57.mlp.gate_projL57.linear_attn.out_projL57.linear_attn.in_proj_aL57.linear_attn.in_proj_bL57.linear_attn.in_proj_zL57.linear_attn.in_proj_qkvL56.mlp.up_projL56.mlp.down_projL56.mlp.gate_projL56.linear_attn.out_projL56.linear_attn.in_proj_aL56.linear_attn.in_proj_bL56.linear_attn.in_proj_zL56.linear_attn.in_proj_qkvL55.mlp.up_projL55.mlp.down_projL55.mlp.gate_projL55.self_attn.o_projL55.self_attn.v_projL55.self_attn.k_projL55.self_attn.q_projL54.mlp.up_projL54.mlp.down_projL54.mlp.gate_projL54.linear_attn.out_projL54.linear_attn.in_proj_aL54.linear_attn.in_proj_bL54.linear_attn.in_proj_zL54.linear_attn.in_proj_qkvL53.mlp.up_projL53.mlp.down_projL53.mlp.gate_projL53.linear_attn.out_projL53.linear_attn.in_proj_aL53.linear_attn.in_proj_bL53.linear_attn.in_proj_zL53.linear_attn.in_proj_qkvL52.mlp.up_projL52.mlp.down_projL52.mlp.gate_projL52.linear_attn.out_projL52.linear_attn.in_proj_aL52.linear_attn.in_proj_bL52.linear_attn.in_proj_zL52.linear_attn.in_proj_qkvL51.mlp.up_projL51.mlp.down_projL51.mlp.gate_projL51.self_attn.o_projL51.self_attn.v_projL51.self_attn.k_projL51.self_attn.q_projL50.mlp.up_projL50.mlp.down_projL50.mlp.gate_projL50.linear_attn.out_projL50.linear_attn.in_proj_aL50.linear_attn.in_proj_bL50.linear_attn.in_proj_zL50.linear_attn.in_proj_qkvL49.mlp.up_projL49.mlp.down_projL49.mlp.gate_projL49.linear_attn.out_projL49.linear_attn.in_proj_aL49.linear_attn.in_proj_bL49.linear_attn.in_proj_zL49.linear_attn.in_proj_qkvL48.mlp.up_projL48.mlp.down_projL48.mlp.gate_projL48.linear_attn.out_projL48.linear_attn.in_proj_aL48.linear_attn.in_proj_bL48.linear_attn.in_proj_zL48.linear_attn.in_proj_qkvL47.mlp.up_projL47.mlp.down_projL47.mlp.gate_projL47.self_attn.o_projL47.self_attn.v_projL47.self_attn.k_projL47.self_attn.q_projL46.mlp.up_projL46.mlp.down_projL46.mlp.gate_projL46.linear_attn.out_projL46.linear_attn.in_proj_aL46.linear_attn.in_proj_bL46.linear_attn.in_proj_zL46.linear_attn.in_proj_qkvL45.mlp.up_projL45.mlp.down_projL45.mlp.gate_projL45.linear_attn.out_projL45.linear_attn.in_proj_aL45.linear_attn.in_proj_bL45.linear_attn.in_proj_zL45.linear_attn.in_proj_qkvL44.mlp.up_projL44.mlp.down_projL44.mlp.gate_projL44.linear_attn.out_projL44.linear_attn.in_proj_aL44.linear_attn.in_proj_bL44.linear_attn.in_proj_zL44.linear_attn.in_proj_qkvL43.mlp.up_projL43.mlp.down_projL43.mlp.gate_projL43.self_attn.o_projL43.self_attn.v_projL43.self_attn.k_projL43.self_attn.q_projL42.mlp.up_projL42.mlp.down_projL42.mlp.gate_projL42.linear_attn.out_projL42.linear_attn.in_proj_aL42.linear_attn.in_proj_bL42.linear_attn.in_proj_zL42.linear_attn.in_proj_qkvL41.mlp.up_projL41.mlp.down_projL41.mlp.gate_projL41.linear_attn.out_projL41.linear_attn.in_proj_aL41.linear_attn.in_proj_bL41.linear_attn.in_proj_zL41.linear_attn.in_proj_qkvL40.mlp.up_projL40.mlp.down_projL40.mlp.gate_projL40.linear_attn.out_projL40.linear_attn.in_proj_aL40.linear_attn.in_proj_bL40.linear_attn.in_proj_zL40.linear_attn.in_proj_qkvL39.mlp.up_projL39.mlp.down_projL39.mlp.gate_projL39.self_attn.o_projL39.self_attn.v_projL39.self_attn.k_projL39.self_attn.q_projL38.mlp.up_projL38.mlp.down_projL38.mlp.gate_projL38.linear_attn.out_projL38.linear_attn.in_proj_aL38.linear_attn.in_proj_bL38.linear_attn.in_proj_zL38.linear_attn.in_proj_qkvL37.mlp.up_projL37.mlp.down_projL37.mlp.gate_projL37.linear_attn.out_projL37.linear_attn.in_proj_aL37.linear_attn.in_proj_bL37.linear_attn.in_proj_zL37.linear_attn.in_proj_qkvL36.mlp.up_projL36.mlp.down_projL36.mlp.gate_projL36.linear_attn.out_projL36.linear_attn.in_proj_aL36.linear_attn.in_proj_bL36.linear_attn.in_proj_zL36.linear_attn.in_proj_qkvL35.mlp.up_projL35.mlp.down_projL35.mlp.gate_projL35.self_attn.o_projL35.self_attn.v_projL35.self_attn.k_projL35.self_attn.q_projL34.mlp.up_projL34.mlp.down_projL34.mlp.gate_projL34.linear_attn.out_projL34.linear_attn.in_proj_aL34.linear_attn.in_proj_bL34.linear_attn.in_proj_zL34.linear_attn.in_proj_qkvL33.mlp.up_projL33.mlp.down_projL33.mlp.gate_projL33.linear_attn.out_projL33.linear_attn.in_proj_aL33.linear_attn.in_proj_bL33.linear_attn.in_proj_zL33.linear_attn.in_proj_qkvL32.mlp.up_projL32.mlp.down_projL32.mlp.gate_projL32.linear_attn.out_projL32.linear_attn.in_proj_aL32.linear_attn.in_proj_bL32.linear_attn.in_proj_zL32.linear_attn.in_proj_qkvL31.mlp.up_projL31.mlp.down_projL31.mlp.gate_projL31.self_attn.o_projL31.self_attn.v_projL31.self_attn.k_projL31.self_attn.q_projL30.mlp.up_projL30.mlp.down_projL30.mlp.gate_projL30.linear_attn.out_projL30.linear_attn.in_proj_aL30.linear_attn.in_proj_bL30.linear_attn.in_proj_zL30.linear_attn.in_proj_qkvL29.mlp.up_projL29.mlp.down_projL29.mlp.gate_projL29.linear_attn.out_projL29.linear_attn.in_proj_aL29.linear_attn.in_proj_bL29.linear_attn.in_proj_zL29.linear_attn.in_proj_qkvL28.mlp.up_projL28.mlp.down_projL28.mlp.gate_projL28.linear_attn.out_projL28.linear_attn.in_proj_aL28.linear_attn.in_proj_bL28.linear_attn.in_proj_zL28.linear_attn.in_proj_qkvL27.mlp.up_projL27.mlp.down_projL27.mlp.gate_projL27.self_attn.o_projL27.self_attn.v_projL27.self_attn.k_projL27.self_attn.q_projL26.mlp.up_projL26.mlp.down_projL26.mlp.gate_projL26.linear_attn.out_projL26.linear_attn.in_proj_aL26.linear_attn.in_proj_bL26.linear_attn.in_proj_zL26.linear_attn.in_proj_qkvL25.mlp.up_projL25.mlp.down_projL25.mlp.gate_projL25.linear_attn.out_projL25.linear_attn.in_proj_aL25.linear_attn.in_proj_bL25.linear_attn.in_proj_zL25.linear_attn.in_proj_qkvL24.mlp.up_projL24.mlp.down_projL24.mlp.gate_projL24.linear_attn.out_projL24.linear_attn.in_proj_aL24.linear_attn.in_proj_bL24.linear_attn.in_proj_zL24.linear_attn.in_proj_qkvL23.mlp.up_projL23.mlp.down_projL23.mlp.gate_projL23.self_attn.o_projL23.self_attn.v_projL23.self_attn.k_projL23.self_attn.q_projL22.mlp.up_projL22.mlp.down_projL22.mlp.gate_projL22.linear_attn.out_projL22.linear_attn.in_proj_aL22.linear_attn.in_proj_bL22.linear_attn.in_proj_zL22.linear_attn.in_proj_qkvL21.mlp.up_projL21.mlp.down_projL21.mlp.gate_projL21.linear_attn.out_projL21.linear_attn.in_proj_aL21.linear_attn.in_proj_bL21.linear_attn.in_proj_zL21.linear_attn.in_proj_qkvL20.mlp.up_projL20.mlp.down_projL20.mlp.gate_projL20.linear_attn.out_projL20.linear_attn.in_proj_aL20.linear_attn.in_proj_bL20.linear_attn.in_proj_zL20.linear_attn.in_proj_qkvL19.mlp.up_projL19.mlp.down_projL19.mlp.gate_projL19.self_attn.o_projL19.self_attn.v_projL19.self_attn.k_projL19.self_attn.q_projL18.mlp.up_projL18.mlp.down_projL18.mlp.gate_projL18.linear_attn.out_projL18.linear_attn.in_proj_aL18.linear_attn.in_proj_bL18.linear_attn.in_proj_zL18.linear_attn.in_proj_qkvL17.mlp.up_projL17.mlp.down_projL17.mlp.gate_projL17.linear_attn.out_projL17.linear_attn.in_proj_aL17.linear_attn.in_proj_bL17.linear_attn.in_proj_zL17.linear_attn.in_proj_qkvL16.mlp.up_projL16.mlp.down_projL16.mlp.gate_projL16.linear_attn.out_projL16.linear_attn.in_proj_aL16.linear_attn.in_proj_bL16.linear_attn.in_proj_zL16.linear_attn.in_proj_qkvL15.mlp.up_projL15.mlp.down_projL15.mlp.gate_projL15.self_attn.o_projL15.self_attn.v_projL15.self_attn.k_projL15.self_attn.q_projL14.mlp.up_projL14.mlp.down_projL14.mlp.gate_projL14.linear_attn.out_projL14.linear_attn.in_proj_aL14.linear_attn.in_proj_bL14.linear_attn.in_proj_zL14.linear_attn.in_proj_qkvL13.mlp.up_projL13.mlp.down_projL13.mlp.gate_projL13.linear_attn.out_projL13.linear_attn.in_proj_aL13.linear_attn.in_proj_bL13.linear_attn.in_proj_zL13.linear_attn.in_proj_qkvL12.mlp.up_projL12.mlp.down_projL12.mlp.gate_projL12.linear_attn.out_projL12.linear_attn.in_proj_aL12.linear_attn.in_proj_bL12.linear_attn.in_proj_zL12.linear_attn.in_proj_qkvL11.mlp.up_projL11.mlp.down_projL11.mlp.gate_projL11.self_attn.o_projL11.self_attn.v_projL11.self_attn.k_projL11.self_attn.q_projL10.mlp.up_projL10.mlp.down_projL10.mlp.gate_projL10.linear_attn.out_projL10.linear_attn.in_proj_aL10.linear_attn.in_proj_bL10.linear_attn.in_proj_zL10.linear_attn.in_proj_qkvL9.mlp.up_projL9.mlp.down_projL9.mlp.gate_projL9.linear_attn.out_projL9.linear_attn.in_proj_aL9.linear_attn.in_proj_bL9.linear_attn.in_proj_zL9.linear_attn.in_proj_qkvL8.mlp.up_projL8.mlp.down_projL8.mlp.gate_projL8.linear_attn.out_projL8.linear_attn.in_proj_aL8.linear_attn.in_proj_bL8.linear_attn.in_proj_zL8.linear_attn.in_proj_qkvL7.mlp.up_projL7.mlp.down_projL7.mlp.gate_projL7.self_attn.o_projL7.self_attn.v_projL7.self_attn.k_projL7.self_attn.q_projL6.mlp.up_projL6.mlp.down_projL6.mlp.gate_projL6.linear_attn.out_projL6.linear_attn.in_proj_aL6.linear_attn.in_proj_bL6.linear_attn.in_proj_zL6.linear_attn.in_proj_qkvL5.mlp.up_projL5.mlp.down_projL5.mlp.gate_projL5.linear_attn.out_projL5.linear_attn.in_proj_aL5.linear_attn.in_proj_bL5.linear_attn.in_proj_zL5.linear_attn.in_proj_qkvL4.mlp.up_projL4.mlp.down_projL4.mlp.gate_projL4.linear_attn.out_projL4.linear_attn.in_proj_aL4.linear_attn.in_proj_bL4.linear_attn.in_proj_zL4.linear_attn.in_proj_qkvL3.mlp.up_projL3.mlp.down_projL3.mlp.gate_projL3.self_attn.o_projL3.self_attn.v_projL3.self_attn.k_projL3.self_attn.q_projL2.mlp.up_projL2.mlp.down_projL2.mlp.gate_projL2.linear_attn.out_projL2.linear_attn.in_proj_aL2.linear_attn.in_proj_bL2.linear_attn.in_proj_zL2.linear_attn.in_proj_qkvL1.mlp.up_projL1.mlp.down_projL1.mlp.gate_projL1.linear_attn.out_projL1.linear_attn.in_proj_aL1.linear_attn.in_proj_bL1.linear_attn.in_proj_zL1.linear_attn.in_proj_qkvL0.mlp.up_projL0.mlp.down_projL0.mlp.gate_projL0.linear_attn.out_projL0.linear_attn.in_proj_aL0.linear_attn.in_proj_bL0.linear_attn.in_proj_zL0.linear_attn.in_proj_qkvDanger score (lower = worse) — log scaleReal sensitivity (higher = worse)

Each dot = one real tensor measured in round 3 specifically (n=496).

Query / attention-input family Key / attention-output family Everything else (MLP, V)
496 / 496
Drag to replay the real order these 496 tensors were actually measured in — layer 0 through the final layer, then lm_head. Not simulated: this is the true capture sequence.
In plain terms

It isn't a perfect predictor — a Pearson correlation of +0.17 means real, useful signal, not a guarantee. The Spearman rank correlation (+0.17) matters here specifically because it only cares about ordering, not the raw KL scale — so it reads the same whichever calibration mode produced these numbers. Both agree: you can look at data you already have, for free, and get a real head start on which tensors are worth worrying about, before spending 40+ hours finding out the hard way.

Finding 2 — it's not random, it's which part of each layer

Group the same danger score by which piece of the model it belongs to — Query, Key, the MLP projections, and so on — and a clean, simple story appears. This isn't about which layer number is risky. It's about which role is risky, and that role is the same everywhere in the model.

self attn q proj0.0005linear attn in proj qkv0.0008linear attn in proj z0.0009self attn o proj0.0010mlp up proj0.0013self attn v proj0.0014mlp gate proj0.0015mlp down proj0.0020linear attn out proj0.0023self attn k proj0.0054linear attn in proj a0.0374linear attn in proj b0.0586

Mean danger score by tensor role, worst (top, red) to safest (bottom, purple). self attn.q proj is consistently the narrow weak spot. linear attn.in proj b is consistently the safest, most forgiving part of every layer.

What you can do with this today

When you're deciding where to spend extra bits, this says: favor Query and the attention-input tensors everywhere in the model. You can more safely push Key and attention-output tensors to lower bits, everywhere. That's a real, structural rule you can apply right now — no more waiting on the full run to tell you the same thing layer by layer.

Finding 3 — real sensitivity grows with depth, and the old method never saw it

This compares the real cascaded measurement (context-aware — each layer measured with everything before it already quantized, the way the real deployed model actually works) against the old method (each tensor tested alone, with a perfectly clean model around it). The chart shows how much more sensitive the real test found each layer, compared to what the old method reported.

0.0x168.5x337.1x505.6x674.2x06121824303642485460Layer numberMore sensitive vs old method

1.0x would mean "the old method and the real test agree." By layer 63, the real test is finding roughly 380-613x more real sensitivity than the old method ever reported — because the old method could never see error building up from earlier layers.

In plain terms

The deeper into the model you go, the more the old, isolated-testing method was underselling the real risk. Every model you've built so far — including this one — used a bit-allocation plan built from that old method. This chart is direct, real evidence that plan is leaving real quality on the table, and it gets worse the deeper you look.

Extrapolating the live run — what happens next, statistically

Fit entirely on this run's own 496 measured tensors — no historical run, no calibration assumption. A real least-squares line through danger score → real sensitivity, R²=0.03, then used to predict the 0 tensors the live run hasn't reached yet.

0.03
R² of the live fit
How much of the real variance the free number explains, on this run's own data
±0.142
95% prediction band
From the fit's real residual spread, n=496
0
Tensors predicted, not yet measured
Every one has a free Hessian score already

The 15 still-unmeasured tensors this fit predicts will turn out most sensitive — worth watching for when the live run actually reaches them.

TensorDanger scorePredicted real KL
Honest limits of this extrapolation

This is a simple linear fit on a real but partial, noisy sample (R²=0.03 — Finding 1's own scatter shows this isn't a tight line). Treat the predicted table as informed triage, not certainty: a real ranking signal to look at before the run gets there, not a substitute for the actual measurement once it arrives.

The top 15 action items — both directions

Real tensor names, real scores, straight from your own build logs, each with a concrete recommended move: 15 to protect (raise their bit width) and 15 you can safely reclaim bits from (lower their bit width) to fund the first 15 — a real, zero-sum rebalancing you can act on today, no more analysis required.

Protect these first

Lowest danger score = narrowest, most dangerous weak spot

TensorScoreAction
L63.mlp.up_proj0.00016Raise bit width
L7.self_attn.q_proj0.00018Raise bit width
L3.self_attn.q_proj0.00018Raise bit width
L11.self_attn.q_proj0.00019Raise bit width
L19.self_attn.q_proj0.00020Raise bit width
L15.self_attn.q_proj0.00021Raise bit width
L23.self_attn.q_proj0.00022Raise bit width
L27.self_attn.q_proj0.00023Raise bit width
L6.linear_attn.in_proj_qkv0.00023Raise bit width
L63.self_attn.q_proj0.00024Raise bit width
L1.linear_attn.in_proj_qkv0.00024Raise bit width
L0.linear_attn.out_proj0.00025Raise bit width
L14.linear_attn.in_proj_qkv0.00025Raise bit width
L31.self_attn.q_proj0.00026Raise bit width
L16.linear_attn.in_proj_qkv0.00026Raise bit width
Safe to compress further

Highest danger score = spread-out, forgiving

TensorScoreAction
L58.linear_attn.in_proj_b0.17802Lower bit width
L41.linear_attn.in_proj_b0.17703Lower bit width
L36.linear_attn.in_proj_b0.16203Lower bit width
L44.linear_attn.in_proj_b0.16035Lower bit width
L52.linear_attn.in_proj_b0.15767Lower bit width
L38.linear_attn.in_proj_b0.15289Lower bit width
L60.linear_attn.in_proj_b0.13932Lower bit width
L42.linear_attn.in_proj_b0.12690Lower bit width
L54.linear_attn.in_proj_a0.11746Lower bit width
L40.linear_attn.in_proj_b0.11314Lower bit width
L45.linear_attn.in_proj_b0.11237Lower bit width
L37.linear_attn.in_proj_b0.10993Lower bit width
L40.linear_attn.in_proj_a0.10673Lower bit width
L57.linear_attn.in_proj_a0.09255Lower bit width
L38.linear_attn.in_proj_a0.08788Lower bit width

What's real, and what's still open

Real, confirmed today

The +0.17 correlation, the Query-vs-Key pattern, and the depth trend are all computed directly from real data on disk — no estimates, no simulation. The Hessian data cost zero extra GPU time; it already existed from the YAQA build you already ran.

Still open

Only 497 of 497 tensors have real ground-truth sensitivity so far — the cascaded run is still going (round 3). The correlation and patterns above could sharpen, weaken, or hold as the rest of the model gets measured. See build_complete_round_story.py for what a fully-finished 3-round run already shows, including whether it actually converges.

Author: Hakim Ghelab, VegaLaboratories LTD · Regenerated 2026-09-30 00:50:08 · Real data sources: hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json (Hessian, shared cross-run master score file), cascaded_checkpoint_round3.json (real sensitivity, round 3, in progress), tensor_report.csv (original isolated measurement) · Regenerate with python3 refresh_hessian_story.py