Four real questions, asked in order, that made the solver scientific instead of guessing
Every number on this page comes from a real solver run this session, or from a real measurement already sitting in this project's own build logs. Nothing here is illustrative — the 3D chart below plots the real, measured mean danger score for every tensor role in this model, exactly as measured.
This chapter compresses ground first covered, in much fuller detail, by four earlier real pages: The Free Signal Hiding in Your Hessians, What a Complete Round Actually Shows, Cascade Investigation, Full Reasoning, and Danger Score Geometry. Read those for the full hand-verifiable derivation this page's four questions summarize.
Can we trust KL?
"Is the number we're measuring actually telling us the truth?"
Isolated KL sensitivity — quantize one tensor, everything else stays BF16, measure the output distance — is cheap and was this project's original signal. But checked directly against real cascaded ground truth, it inverts: tensors it ranks as safe are sometimes the ones that hurt most once every OTHER tensor is also quantized around them. The unreliability isn't uniform — it's worst in the first 25% of the network's depth (a real, measured 34.4% ranking-inversion rate) and much cleaner in the last 25% (1.3%). Isolated KL was never wrong about magnitude everywhere — it was wrong specifically where cross-layer effects compound the most, and it can't see those effects by construction: it only ever quantizes one tensor at a time.
Cascaded KL — a real improvement, still an imperfect one
"Can we measure cross-layer interaction directly, instead of assuming it away?"
The real fix: measure each tensor's sensitivity with every already-decided tensor genuinely quantized
around it, in causal depth order, then re-run the whole pass again from a fresh reload — a real fixed-point
loop, not a one-shot estimate. This is strictly better: it can actually see compounding effects instead of
assuming they don't exist. But it inherits a real, structural problem of its own — it's a chain. The tensor
measured last in any round (in this model, lm_head) is measured on top of whatever degradation the
round's own earlier choices already introduced. Its own measured sensitivity curve came back almost flat across
candidate bit-widths (0.2326 at 4-bit vs. 0.2289 at 8-bit — a 1.6% range) — not because lm_head is
safe, but because by the time measurement reaches it, the signal is already blurred by everything decided before
it. Think of it as a snowball rolling downhill, picking up every prior layer's noise along the way.
Floors — a real fix for a real problem, built from best practice, not measurement
"If the number can't be trusted here, what do we do instead?"
The honest answer, at the time: hand-code a minimum bit-width for tensors in the regions and categories known,
from outside evidence, to be risky — early Q/K/V input projections, late-layer attention, output/boundary layers.
Each floor is real and each one was motivated by a real, observed regression it fixed. But every one of them is
the same shape of fix: a single flat multiplier (100×) or a single minimum-bits rule, applied to an entire
category — a depth range, a name pattern — never to an individual tensor's own measured properties.
lm_head carried that exact 100× weight and still landed at Q4 in two real, independent
cascaded rounds, because a rigorous MILP solver will always find a cheaper trade elsewhere once a tensor's own
measured marginal KL is small enough — no matter how large a price you attach to it. A soft weight is a
price. Prices can always be outbid. Only a hard floor — a real bits ≥ N constraint — cannot be
traded away, which is exactly why lm_head needed one.
A floor answers "which category is this tensor in?" It has no opinion on this specific tensor's own measured curvature. Two tensors in the same category, one genuinely narrow and dangerous, one genuinely diffuse and forgiving, get the exact same flat treatment. That is a real, remaining source of imprecision — not because best practice was wrong, but because a category flag was always going to be a coarser instrument than a real per-tensor measurement.
The Hessian score — a second signal, immune to the snowball
"Is there a real, per-tensor measurement of risk that doesn't compound the way KL does?"
Every tensor already has a real Hessian computed for it during YAQA's own two-sided correction — H_I (input curvature) and H_O (output curvature), the exact matrices used to compute the optimal rounding correction. Buried inside that Hessian is a second, free, genuinely independent signal: effective rank, a real spectral diagnostic of how concentrated that Hessian's energy is.
hess_score(tensor) = mean( effective_rank(H_I)/dim_in , effective_rank(H_O)/dim_out )
Low score = the Hessian's energy is concentrated in very few directions — a small rounding error landing in exactly the wrong direction does real damage, because there's nowhere for the error to hide. High score = energy spread broadly across many directions — error in any one direction barely moves the total, because hundreds of others are absorbing it too. This is pure linear algebra on that one tensor's own real curvature. It never sees another tensor's bit-width, never compounds across layers, and cannot inherit upstream noise the way a cascaded measurement structurally must.
And it is measurable for every real tensor except one. lm_head's own two-sided Hessian
would need 246.7GB to even compute — the same real memory wall this whole investigation started from. There is
no effective rank for a Hessian that was never built. That single fact turns out to matter a great deal for what
comes next.
What happens when you actually feed it to the solver
"Does a genuinely richer signal change the solver's real decisions — and does it agree with what the field already knows?"
Two real experiments, run against the exact same real cascaded checkpoint this build already had, same real BPW budget, nothing simulated.
Exact real config: --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 —
--no-protect-first-last is not used, so the flat 100× boundary category
(lm_head + layer 0 + layer 63) stays on exactly as it does in production. Only the two much larger
regional floors are disabled. 69 of 497 real tensors move relative to the floored production plan
— and on its own initiative, with zero Hessian signal involved at all, the solver pushes lm_head
from its 6-bit floor up to 16-bit, full precision. Nothing told it to. The moment hand-coded regional
floors stop competing for the same budget, the solver's own real cascaded-weighted objective decides, entirely
on its own, that this tensor deserves more protection than the floor even requires.
Exact real config: --late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6
--no-protect-first-last --hessian-score-file ... --hessian-weight-strength 2.0. "Regional floors exactly
as shipped" means the real production values, 5 and 6 — those stay on. What changes is the flat 100×
"first/last layer" category flag (lm_head + layer 0 + layer 63) — removed via
--no-protect-first-last and replaced by the real, continuous, percentile-ranked Hessian score
instead. 34 of 497 real tensors move:
| Tensor | Flat-category plan (shipped) | Hessian-informed plan | Real reason |
|---|---|---|---|
| lm_head | 6-bit (floor) | 16-bit | Self-protects again — even more emphatically this time |
| L0 / L63 mlp.*_proj | 16-bit | 4-bit | Not a safe call. Layer 0's real curvature is unmeasured (0 real score — never corrected in the build the score file came from). Layer 63's is measured, and it says the opposite of safe: these are the single most dangerous tensors of all 360 real scored tensors (percentile r≈0.00–0.07). Real mechanism: --no-protect-first-last removes the flat 100× entirely, leaving only the Hessian multiplier — at strength=2.0 and r≈0, that's roughly 3×, far too weak on its own to stop the MILP from cutting them when the real BPW budget needs the room elsewhere. This is a real argument against Hessian-alone replacing the boundary category at this strength, not evidence it's safe. |
| L11 self_attn.v_proj | 16-bit | 6-bit | Freed budget lands where measured risk, not position, says it belongs |
| L19 self_attn.k_proj | 16-bit | 5-bit | Settles at the real late-attention floor instead of being over-protected by accident |
| L1 linear_attn.in_proj_b | 6-bit | 16-bit | Real curvature flags it as genuinely narrow — promoted on the evidence |
An earlier version of this page described the L0/L63 mlp.*_proj demotion as curvature-justified.
It wasn't checked against the real score at the time, and it was wrong for layer 63 specifically (real score
says maximally dangerous, not safe). The real, current recommendation is the opposite of what this experiment
tried: keep the flat 100× boundary category on, always — --no-protect-first-last is never
used. See 08_HESSIAN_HYBRID_SOLVER_DEBRIEF.html section 7 for the real, current final setting.
Same real BPW budget throughout, start to finish — this is a reallocation, not "spend more." Real, remaining caveat: not every move in this table holds up under direct inspection (see the correction above) — this experiment is kept here for the record, not as a template to follow.
Grouped by role instead of by layer number, the real mean Hessian danger score draws a clean, consistent
picture — the same 3D chart at the top of this page: self_attn.q_proj is the narrowest, most
dangerous role in every layer (mean score 0.0005). self_attn.k_proj is the safest, most
forgiving role in every layer (mean score 0.0091 — 18× more forgiving). This is not something we
told the measurement to find. It falls directly out of real curvature data, and it lines up with what attention
mechanics already predicts on structural grounds: a query projection directly shapes which tokens get
attended to — perturb it and the whole attention pattern for that position can shift. A key projection's role is
more redundant across heads and positions, so the same size of perturbation has far more places to get
absorbed. The Hessian score didn't need to be told this. It measured it.
Where the real solver actually stands right now
Precisely, so nothing here is overstated in either direction.
The real, causal-order, fixed-point measurement loop. This has been the production signal since before this investigation began.
lm_head / boundary hard floors (≥6-bit)
Genuine MILP constraints, not prices. As of 2026-09-14, L0/L63 get the same floor+weight combination lm_head always had (--boundary-min-bits) — real ablation proved the flat weight alone was not enough for them either.
Off when running on cascade KL alone (confirmed, Experiment 1). Current working answer for when Hessian is primary: keep them on too, as an extra safety net, not yet benchmark-validated either way.
--hessian-floor-tiers / --hessian-primary — real MILP mechanisms, not a weight. Closes a real gap the soft weighting below could not: even at 500× the strength, it changed zero outcomes for the top 20 most dangerous tensors. Full story: 08_HESSIAN_HYBRID_SOLVER_DEBRIEF.html §8.
Everything on this page is real solver output and real measured data — but it is not yet a proven quality result. Two historical incidents (the V3.1 IFEval drop, the V3.2 MMLU drop) exist as real regression tests for exactly this kind of change. The honest next building block, before this becomes anyone's default, is running those same two checks against a model actually built from the Hessian-informed plan — not another plan diff.