← Back to index
Hessian danger score · isolated KL · 394 real tensors · written 2026-09-16, live-refreshed since

Does the free Hessian score agree with the expensive KL measurement, tensor by tensor?

Every one of the 488 real tensors this model has a Hessian curvature score for, compared directly against its own real, measured isolated-KL sensitivity — both real rankings shown side by side, no averaging, no rounding into buckets. Upload a different Hessian score file below to recompute this live against the same real KL baseline.

Where this sits in the real story (added 2026-09-17): deep-dive companion, not the story's spine — for the current, continuing account, start at Improvement Ledger §07 and read forward through §10. Everything on this page is still accurate as of 2026-09-16.

The real numbers, up front

Computed directly from tensor_report.csv (real isolated KL) and this model's own real Hessian logs — not estimated.

What each tensor role actually does

Ten real roles appear in this data, across two real attention mechanisms this model mixes layer-by-layer (standard full attention, and a linear/recurrent Gated DeltaNet block — confirmed directly from this model's own real source, mlx_lm/models/qwen3_5.py) plus the MLP block every layer has. Each card's verdict is computed live from whichever dataset is currently loaded.

Every tensor, both rankings, one picture

X = Hessian danger rank (1 = most dangerous). Y = isolated-KL danger rank (1 = most dangerous). The diagonal line is what perfect agreement would look like — a tensor sitting exactly on it means both signals rank it identically. Color shows how far each tensor sits from that line.

Close agreement Wide disagreement

Hover any point for its exact name, both ranks, and both real values. The line y=x is the "if the two signals agreed perfectly" reference — not a fit, not a trend, the literal identity line.

The real pattern in the top-left corner (added 2026-09-17)

The top-left region of the chart above — Hessian says dangerous (low rank), isolated KL says safe (high rank) — is not scattered randomly across the model. Checked directly against the real, current data:

Every tensor where the Hessian ranks it in the top 40 most dangerous (of 488) while isolated KL ranks it in roughly the safest quarter is a late layer — 59, 62, and 63 only, out of a 64-layer model (0–63). No early or middle layer appears in this disagreement zone at any threshold checked.

LayerComponentHessian rankKL rank
63mlp.up_proj1395/488
63self_attn.q_proj10420/488
62mlp.up_proj20429/488
63mlp.gate_proj24414/488
62linear_attn.in_proj_z38458/488
59self_attn.o_proj40423/488

Table and summary sentence above are recomputed by refresh_hessian_vs_kl_explorer.py on every refresh, same as the chart. The interpretation below is not — that's a standing hypothesis, re-checked by hand, not regenerated.

Why this isn't a coincidence: this project already suspected late layers need protection isolated KL doesn't guarantee — that's the entire real reason the --late-attn-min-bits floor exists, and it's exactly the failure mode that produced the real lm_head Q4 incident (§04). This chart is a third, independent, per-tensor confirmation of the same structural blind spot — not proof of a mechanism, but real, quantified, repeatable evidence pointing at one.

A plausible mechanism, explicitly unconfirmed: isolated KL measures one tensor's perturbation with nothing downstream left to compound or reveal the damage — a late-layer tensor's rounding error has no later layer left to amplify it before the output. Hessian curvature measures local rounding-sensitivity directly, independent of what's downstream, so it wouldn't share that blind spot. This is a real, testable hypothesis, not a settled explanation.

What this suggests doing next, and why: in active-learning terms, this is a query-by-committee situation — when two independent measurement principles disagree sharply on a specific tensor, that disagreement itself is the highest-value place to spend real measurement next, more valuable than confirming tensors the two signals already agree on. The live godmode sweep (still running) should prioritize exactly these 8 tensors, and the layer 48–63 region generally, ahead of arbitrary layer order — resolving whether Hessian or KL is right here has real, concrete stakes given §04 already burned this project once in this exact region.

The mirror pattern in the bottom-right corner (added 2026-09-17)

The exact opposite corner of the chart — KL says dangerous (low rank), Hessian says safe (high rank) — shows the same kind of structural clustering, in the opposite direction.

Every tensor where isolated KL ranks it in the top 50 most dangerous while the Hessian ranks it in roughly the safest quarter is an early layer — 2, 6, and 8 only, out of a 64-layer model. No late or middle layer appears in this mirror disagreement zone at this threshold.

LayerComponentKL rankHessian rank
6linear_attn.in_proj_a36407/488
8linear_attn.in_proj_b48457/488
2linear_attn.in_proj_a50429/488

Table and summary sentence above are recomputed by refresh_hessian_vs_kl_explorer.py on every refresh, same as the top-left table. The interpretation below is not.

A secondary, weaker pattern worth naming honestly: 8 of these 9 tensors are linear_attn components (the Gated DeltaNet block), not self_attn or mlp. Checked directly against this model's real source (mlx_lm/models/qwen3_5.py, full_attention_interval=4): linear_attn already makes up 75% of all layers uniformly across the whole network, so an 8-of-9 (~89%) concentration here is only a modest lift over the base rate at this small a sample size (n=9) — real, but not a strong claim on its own.

The pattern that ties both corners together: Hessian curvature is a local measurement — how sensitive a tensor's own rounding is, independent of what's downstream. Isolated KL is a global measurement — how much a tensor's perturbation actually changes the final output, which depends entirely on how many layers are left to propagate and compound that error. Late layers have almost nothing downstream left to propagate through, so KL's global view goes blind exactly where Hessian's local view still sees real danger (the top-left corner). Early layers have the entire rest of the network left to propagate through, so KL's global view correctly amplifies real cumulative risk that Hessian's local, position-blind view has no way to see (this corner). Both corners are the same underlying asymmetry, seen from opposite ends.

What this suggests doing next: beyond simply prioritizing these 9 tensors in the live sweep (the same query-by-committee logic as the top-left corner), this is a real, concrete, testable design for the "super solver" idea from earlier — not just trust-agreement/flag-disagreement, but position-dependent signal weighting: weight isolated KL more heavily near the start of the network, weight Hessian more heavily near the end, rather than treating both signals as equally trustworthy everywhere. This is a hypothesis suggested by combining both corners, not yet a design decision.

Recompute this live against a different Hessian score file

The chart, role cards, and table below are all driven by one live dataset. By default that's tonight's real computed comparison. Upload a different hess_scores_*.json file (the same flat {"tensor_name": score} format this project's own 08_extract_hessian_scores.py produces) to see how a different Hessian run compares against the same real isolated-KL baseline — recomputed entirely in your browser, nothing uploaded anywhere.

Smart recompute
Currently showing: the real, live-refreshed dataset (394 tensors)

Complete list, every tensor identified

Sortable, searchable, filterable — every row real, nothing aggregated away. Boundary tensors (layer 0 / layer 63, always hard-protected regardless of either signal) are flagged, not hidden.

Hess rank ▾ KL rank ▾ Gap ▾ Layer ▾ Role ▾ Tensor name ▾ Hess score ▾ Isolated KL ▾
Author: Hakim Ghelab, VegaLaboratories LTD · Real data sources: tensor_report.csv (isolated KL, fixed baseline), this model's own real Hessian logs (default Hessian side, swappable via upload) · Role definitions verified directly against mlx_lm/models/qwen3_5.py (Gated DeltaNet linear-attention block) · Every number on this page is computed live in your browser from real data, nothing pre-rendered or illustrative.