← Back to index
Improvement ledger 07 · the Hessian signal · 2026-09-13, updated 2026-09-15

Four real questions, asked in order, that made the solver scientific instead of guessing

Every number on this page comes from a real solver run this session, or from a real measurement already sitting in this project's own build logs. Nothing here is illustrative — the 3D chart below plots the real, measured mean danger score for every tensor role in this model, exactly as measured.

01 · Trust KL? 02 · Cascade it 03 · Floors: a patch 04 · The Hessian score 05 · Feed it to the solver 06 · Where we really are
Full derivation (added 2026-09-17)

This chapter compresses ground first covered, in much fuller detail, by four earlier real pages: The Free Signal Hiding in Your Hessians, What a Complete Round Actually Shows, Cascade Investigation, Full Reasoning, and Danger Score Geometry. Read those for the full hand-verifiable derivation this page's four questions summarize.

Real data · mean Hessian danger score by tensor role
Drag to rotate · red = narrow, dangerous Hessian · purple/mint = diffuse, forgiving Hessian · height = real measured score
STEP 01

Can we trust KL?

"Is the number we're measuring actually telling us the truth?"

Isolated KL sensitivity — quantize one tensor, everything else stays BF16, measure the output distance — is cheap and was this project's original signal. But checked directly against real cascaded ground truth, it inverts: tensors it ranks as safe are sometimes the ones that hurt most once every OTHER tensor is also quantized around them. The unreliability isn't uniform — it's worst in the first 25% of the network's depth (a real, measured 34.4% ranking-inversion rate) and much cleaner in the last 25% (1.3%). Isolated KL was never wrong about magnitude everywhere — it was wrong specifically where cross-layer effects compound the most, and it can't see those effects by construction: it only ever quantizes one tensor at a time.

STEP 02

Cascaded KL — a real improvement, still an imperfect one

"Can we measure cross-layer interaction directly, instead of assuming it away?"

The real fix: measure each tensor's sensitivity with every already-decided tensor genuinely quantized around it, in causal depth order, then re-run the whole pass again from a fresh reload — a real fixed-point loop, not a one-shot estimate. This is strictly better: it can actually see compounding effects instead of assuming they don't exist. But it inherits a real, structural problem of its own — it's a chain. The tensor measured last in any round (in this model, lm_head) is measured on top of whatever degradation the round's own earlier choices already introduced. Its own measured sensitivity curve came back almost flat across candidate bit-widths (0.2326 at 4-bit vs. 0.2289 at 8-bit — a 1.6% range) — not because lm_head is safe, but because by the time measurement reaches it, the signal is already blurred by everything decided before it. Think of it as a snowball rolling downhill, picking up every prior layer's noise along the way.

Round 1Fresh measurementEvery tensor measured causally, each one locked in before the next is measured.
Round 2Reload, re-measureSame causal pass, but now every tensor sees round 1's real, quantized context — not BF16 anymore.
Round 3Convergence checkReal data: 62 of 497 tensors still oscillate between round 2 and round 3's own choices — the chain hasn't fully settled.
The costCompounding noiseWhatever was noisy upstream doesn't cancel out — it accumulates into every later measurement in the same round.
STEP 03

Floors — a real fix for a real problem, built from best practice, not measurement

"If the number can't be trusted here, what do we do instead?"

The honest answer, at the time: hand-code a minimum bit-width for tensors in the regions and categories known, from outside evidence, to be risky — early Q/K/V input projections, late-layer attention, output/boundary layers. Each floor is real and each one was motivated by a real, observed regression it fixed. But every one of them is the same shape of fix: a single flat multiplier (100×) or a single minimum-bits rule, applied to an entire category — a depth range, a name pattern — never to an individual tensor's own measured properties. lm_head carried that exact 100× weight and still landed at Q4 in two real, independent cascaded rounds, because a rigorous MILP solver will always find a cheaper trade elsewhere once a tensor's own measured marginal KL is small enough — no matter how large a price you attach to it. A soft weight is a price. Prices can always be outbid. Only a hard floor — a real bits ≥ N constraint — cannot be traded away, which is exactly why lm_head needed one.

The real gap this leaves open

A floor answers "which category is this tensor in?" It has no opinion on this specific tensor's own measured curvature. Two tensors in the same category, one genuinely narrow and dangerous, one genuinely diffuse and forgiving, get the exact same flat treatment. That is a real, remaining source of imprecision — not because best practice was wrong, but because a category flag was always going to be a coarser instrument than a real per-tensor measurement.

STEP 04

The Hessian score — a second signal, immune to the snowball

"Is there a real, per-tensor measurement of risk that doesn't compound the way KL does?"

Every tensor already has a real Hessian computed for it during YAQA's own two-sided correction — H_I (input curvature) and H_O (output curvature), the exact matrices used to compute the optimal rounding correction. Buried inside that Hessian is a second, free, genuinely independent signal: effective rank, a real spectral diagnostic of how concentrated that Hessian's energy is.

The real formula, exactly as implemented
effective_rank(H) = trace(H)² / ∑(H²)   # yaqa_core.py
hess_score(tensor) = mean( effective_rank(H_I)/dim_in , effective_rank(H_O)/dim_out )

Low score = the Hessian's energy is concentrated in very few directions — a small rounding error landing in exactly the wrong direction does real damage, because there's nowhere for the error to hide. High score = energy spread broadly across many directions — error in any one direction barely moves the total, because hundreds of others are absorbing it too. This is pure linear algebra on that one tensor's own real curvature. It never sees another tensor's bit-width, never compounds across layers, and cannot inherit upstream noise the way a cascaded measurement structurally must.

And it is measurable for every real tensor except one. lm_head's own two-sided Hessian would need 246.7GB to even compute — the same real memory wall this whole investigation started from. There is no effective rank for a Hessian that was never built. That single fact turns out to matter a great deal for what comes next.

STEP 05

What happens when you actually feed it to the solver

"Does a genuinely richer signal change the solver's real decisions — and does it agree with what the field already knows?"

Two real experiments, run against the exact same real cascaded checkpoint this build already had, same real BPW budget, nothing simulated.

Experiment 1 — no hand-coded floors at all, pure cascaded KL only

Exact real config: --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 — --no-protect-first-last is not used, so the flat 100× boundary category (lm_head + layer 0 + layer 63) stays on exactly as it does in production. Only the two much larger regional floors are disabled. 69 of 497 real tensors move relative to the floored production plan — and on its own initiative, with zero Hessian signal involved at all, the solver pushes lm_head from its 6-bit floor up to 16-bit, full precision. Nothing told it to. The moment hand-coded regional floors stop competing for the same budget, the solver's own real cascaded-weighted objective decides, entirely on its own, that this tensor deserves more protection than the floor even requires.

Experiment 2 — replace the flat boundary category with the real Hessian score

Exact real config: --late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 --no-protect-first-last --hessian-score-file ... --hessian-weight-strength 2.0. "Regional floors exactly as shipped" means the real production values, 5 and 6 — those stay on. What changes is the flat 100× "first/last layer" category flag (lm_head + layer 0 + layer 63) — removed via --no-protect-first-last and replaced by the real, continuous, percentile-ranked Hessian score instead. 34 of 497 real tensors move:

TensorFlat-category plan (shipped)Hessian-informed planReal reason
lm_head6-bit (floor)16-bitSelf-protects again — even more emphatically this time
L0 / L63 mlp.*_proj16-bit4-bitNot a safe call. Layer 0's real curvature is unmeasured (0 real score — never corrected in the build the score file came from). Layer 63's is measured, and it says the opposite of safe: these are the single most dangerous tensors of all 360 real scored tensors (percentile r≈0.00–0.07). Real mechanism: --no-protect-first-last removes the flat 100× entirely, leaving only the Hessian multiplier — at strength=2.0 and r≈0, that's roughly 3×, far too weak on its own to stop the MILP from cutting them when the real BPW budget needs the room elsewhere. This is a real argument against Hessian-alone replacing the boundary category at this strength, not evidence it's safe.
L11 self_attn.v_proj16-bit6-bitFreed budget lands where measured risk, not position, says it belongs
L19 self_attn.k_proj16-bit5-bitSettles at the real late-attention floor instead of being over-protected by accident
L1 linear_attn.in_proj_b6-bit16-bitReal curvature flags it as genuinely narrow — promoted on the evidence
Corrected 2026-09-14 — this experiment argues against removing boundary protection

An earlier version of this page described the L0/L63 mlp.*_proj demotion as curvature-justified. It wasn't checked against the real score at the time, and it was wrong for layer 63 specifically (real score says maximally dangerous, not safe). The real, current recommendation is the opposite of what this experiment tried: keep the flat 100× boundary category on, always — --no-protect-first-last is never used. See 08_HESSIAN_HYBRID_SOLVER_DEBRIEF.html section 7 for the real, current final setting.

Same real BPW budget throughout, start to finish — this is a reallocation, not "spend more." Real, remaining caveat: not every move in this table holds up under direct inspection (see the correction above) — this experiment is kept here for the record, not as a template to follow.

The correlation that makes this genuinely beautiful, not just convenient

Grouped by role instead of by layer number, the real mean Hessian danger score draws a clean, consistent picture — the same 3D chart at the top of this page: self_attn.q_proj is the narrowest, most dangerous role in every layer (mean score 0.0005). self_attn.k_proj is the safest, most forgiving role in every layer (mean score 0.0091 — 18× more forgiving). This is not something we told the measurement to find. It falls directly out of real curvature data, and it lines up with what attention mechanics already predicts on structural grounds: a query projection directly shapes which tokens get attended to — perturb it and the whole attention pattern for that position can shift. A key projection's role is more redundant across heads and positions, so the same size of perturbation has far more places to get absorbed. The Hessian score didn't need to be told this. It measured it.

STEP 06

Where the real solver actually stands right now

Precisely, so nothing here is overstated in either direction.

shipped, default-on
Cascaded KL measurement

The real, causal-order, fixed-point measurement loop. This has been the production signal since before this investigation began.

shipped, default-on
lm_head / boundary hard floors (≥6-bit)

Genuine MILP constraints, not prices. As of 2026-09-14, L0/L63 get the same floor+weight combination lm_head always had (--boundary-min-bits) — real ablation proved the flat weight alone was not enough for them either.

shipped, default-on
Early-QKV / late-attention regional floors

Off when running on cascade KL alone (confirmed, Experiment 1). Current working answer for when Hessian is primary: keep them on too, as an extra safety net, not yet benchmark-validated either way.

built, verified, opt-in only
Per-tensor hard floor + lexicographic priority (new, 2026-09-15)

--hessian-floor-tiers / --hessian-primary — real MILP mechanisms, not a weight. Closes a real gap the soft weighting below could not: even at 500× the strength, it changed zero outcomes for the top 20 most dangerous tensors. Full story: 08_HESSIAN_HYBRID_SOLVER_DEBRIEF.html §8.

The one step not yet taken

Everything on this page is real solver output and real measured data — but it is not yet a proven quality result. Two historical incidents (the V3.1 IFEval drop, the V3.2 MMLU drop) exist as real regression tests for exactly this kind of change. The honest next building block, before this becomes anyone's default, is running those same two checks against a model actually built from the Hessian-informed plan — not another plan diff.

Author: Hakim Ghelab, VegaLaboratories LTD · Companion to IMPROVEMENT_LEDGER/04, 05, 06 · Every number on this page is either read directly from a real solver run this session or from this project's own real build logs — none estimated, none illustrative.