The Hessian-aware hybrid solver — every flag, every command, every real result, in order
This page exists so nothing from this investigation gets lost. Every number here is either read directly from a real solver run this session, or computed live from the real 360-tensor score file that run produced. No strength sweep was performed — see 4.1 for the exact, honest scope of what was tested.
The two real flags, exactly as implemented
Both live in 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py. Both are off by
default — omit --hessian-score-file and the solver's output is byte-identical to before this
feature existed (verified directly).
Where the real score comes from
A real spectral participation-ratio diagnostic of each tensor's own already-computed H_I/H_O (the same Hessians YAQA's own two-sided correction already builds — nothing new is measured). Low score = energy concentrated in few directions (narrow, dangerous). High score = spread broadly (forgiving).
The raw per-tensor data was always free — every real YAQA correction pass already prints it. The first
automation attempt hooked run_full_yaqa.sh only — a direct user question ("are we 100% sure
this covers every real build?") caught that MAIN_RESUME_BUILD.sh, a real, existing script that
calls 05_full_model_quantize.py --assemble-from-resume directly, bypasses
run_full_yaqa.sh entirely and would have silently skipped it. The hook now lives inside
05_full_model_quantize.py's own main() instead — the one real code path every
real assembly runs through, regardless of which script invokes it — writing
<resume_dir>/hessian_scores.json as a non-blocking side effect of the real save. No
manual step is needed for any future real build, from either entry point.
Exact command (what 05_full_model_quantize.py's own assembly step now runs automatically,
and what you'd run by hand against any older, already-completed build)
Real, standalone, tested script — not a one-off snippet anymore (08_extract_hessian_scores.py);
reuses the same, already-tested hessian_story_lib.load_hessian() extraction logic. Verified
byte-for-byte identical to the original manual extraction. Real, permanent output (not a scratch file):
research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json —
360 real tensors scored. lm_head is absent by real construction (its own two-sided
Hessian needs 246.7GB and is never built) — not an omission, see section 5. Real scope note: the score is a
property of the model's own real weights and calibration data, not of which plan built it — one real
extraction is reusable for any future plan against the same model.
The exact granularity mechanism — real numbers, not the abstract version
This is the direct answer to "how is this actually more granular than a flat 100× score" — three real steps, then five real tensors run through them.
Rank all 360 real scored tensors by their real
hess_score, most dangerous (lowest) to safest (highest).
Convert each tensor's rank position into a
percentile r from 0 (most dangerous) to 1 (safest).
Compute its real weight:
factor = 1.0 + strength × (1 − r).
| Real tensor | Real hess_score | Percentile rank r | Real weight factor (strength=2.0) |
|---|---|---|---|
| layers.63.mlp.up_proj (most dangerous of 360) | 0.000161 | 0.000 | 3.000× |
| layers.25.mlp.up_proj | 0.000499 | 0.251 | 2.499× |
| layers.15.self_attn.o_proj (dead middle) | 0.001130 | 0.501 | 1.997× |
| layers.36.mlp.up_proj | 0.002000 | 0.752 | 1.496× |
| layers.55.self_attn.k_proj (safest of 360) | 0.012979 | 1.000 | 1.000× |
Five genuinely different real weights, out of 360 total — every scored tensor gets its own precise
number. The old flat 100× system would have given all five of these tensors 1.0 (none are
named lm_head or a boundary layer) — completely blind to the fact that
layers.63.mlp.up_proj is objectively far more dangerous than layers.36.mlp.up_proj
by real measured curvature. This weight multiplies straight into the MILP's real objective
(weighted_kl = kl × structural_weight) — same slot the flat 100× used to occupy,
continuous input instead of binary.
All four real experiments — exact commands, exact results
2.0 (the flag's own coded default) and 8.0 (one deliberate second point,
chosen to see how the effect scales). This is not a systematic sweep across 1–8, and no
optimization was performed to justify 2.0 as correct — it is a sensible default pending the real eval in
section 6, nothing more.| # | Real config (target-bpw/group-size/candidates held constant) | Strength | Compared against | Real result |
|---|---|---|---|---|
| A | late-attn=0, early-qkv=0, lm-head=6 (no Hessian file) | — | real shipped production plan | 69/497 differ. lm_head self-promotes 6→16, zero Hessian involved. |
| B | same floors as A + Hessian | 2.0 | plan A | Only 3/497 differ from A. lm_head unchanged (16). Real move: layers.55.self_attn.k_proj: 16→4 — matches the original role-based recommendation. |
| C | same floors as A + Hessian | 8.0 | plan A | 40/497 differ from A. Broad promotion of layers 34–62. lm_head drops 16→6 — too much signal dilutes its own specialness. |
| D | late-attn=5, early-qkv=6 (production), lm-head=6, --no-protect-first-last + Hessian | 2.0 | real shipped production plan | 34/497 differ. lm_head self-promotes again, 6→16, floors still active. |
Experiment A — pure cascaded KL, no floors, no Hessian (the zero-Hessian baseline)
Experiment B — Hessian at the default strength (2.0), floors still off
Experiment C — same, strength cranked to 8.0 (the real cautionary result)
Experiment D — the actual "hybrid" configuration: real production floors kept ON
Direct answer: should the regional floors stay ON or come OFF?
--late-attn-min-bits and --early-qkv-min-bits stay at their real production
values (5 and 6). Experiments A/B/C prove the mechanism works and produces defensible reallocation — they do
not prove removing the floors is safe. Nothing here ran the real IFEval/MMLU eval that would actually
answer that. Each floor was added because of a real, measured historical regression.
--no-protect-first-last is never used, full stop. An earlier draft of this section
recommended it as a substitution for the flat 100× boundary category — that recommendation is retracted.
It only touches 16 tensors (lm_head + layer 0's 8 + layer 63's 7) that should stay protected no
matter what, for no real benefit. The real, current recommendation instead removes the two much larger
region floors (124 tensors each) once real cascade measurement (P1) is used — see section 7 for the
full real reasoning and the real, final command.
lm_head is never touched by any of thisStructurally exempt: 0 real entries for language_model.lm_head in the real score file, because
its own two-sided Hessian would need 246.7GB and is never built. Its hard floor
(--lm-head-min-bits 6) is the only thing that has ever reliably protected it, and stays on
regardless of any Hessian setting.
Where this actually stands right now
Unchanged, production signal since before this investigation.
lm_head hard floorA genuine MILP constraint, unaffected by Hessian settings.
A real, second protection layer for the boundary category — the existing flat 100× weight alone was proven insufficient (real ablation: L0's mlp.up_proj/gate_proj still dropped to Q4/Q5 with the weight on and no competing floor). --boundary-min-bits gives L0/L63 the same floor+weight combination lm_head already had.
--hessian-weight-strength)Real, verified, zero-regression when omitted. Superseded as the primary mechanism — see section 8: even at strength=1000 it changed zero outcomes in the top 20 most dangerous real tensors.
--hessian-floor-tiers, new, 2026-09-15)Real MILP minimum-bits constraint from each tensor's own danger rank. Real, confirmed: 0 of the top 54 most dangerous tensors cut to ≤5-bit, vs. 24-44 under every weight-only approach.
--hessian-primary, new, 2026-09-15)Real two-phase solve — Hessian decides first, KL only breaks real leftover ties. Real, measured: KL ends up deciding 0 of 360 scored tensors, 3 of 497 total.
Everything on this page is real solver output — not yet a proven quality result. The honest next building block: build a real model from the real recommended plan (section 8) and run it through the real IFEval/MMLU regression checks that motivated the original floors, before this becomes anyone's default.
Hard floor + lexicographic priority — closing the gap this whole debrief left open
Section 3/4 above showed the soft percentile weight is a real, continuous, per-tensor
signal — a genuine improvement over the flat 100× category weight. This section shows it is still not a
guarantee, and what actually closes that gap. Full real page, every diagram and every number:
research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/MILP_AWARE_HESSIAN_GUIDE.html. Full
command reference: RUNNING_GUIDE.md §11.8.
Pushed --hessian-weight-strength from the tested default (2.0) to 1000 — a 500× jump.
Result: zero outcomes changed in the top 20 most dangerous real tensors in the model. Even the single
most dangerous real tensor measured anywhere (layers.63.mlp.up_proj) only reached 8-bit either
way. The real cause: these tensors' raw KL cost is close enough to flat across candidate bit-widths that
multiplying an already-small number by 1000 is still small next to tensors with genuinely large KL cost,
under the same fixed BPW budget. No multiplier fixes a signal too small to matter in the objective it's
added to.
| Top 54 most dangerous | Production | Soft weight (2.0) | Hard floor only | Lexicographic only | Combined |
|---|---|---|---|---|---|
| Kept ≥16-bit | 1 | 9 | 22 | 20 | 29 |
| Still cut ≤5-bit | 24 | 44 | 0 | 30 | 0 |
Hard floor (--hessian-floor-tiers "0.05:8,0.15:6") is a real MILP minimum, not a
preference — the solver cannot trade it away. Zero of the top 54 most dangerous tensors land at ≤5-bit,
where the soft weight left 44. Real feasibility ceiling found by bisection: this exact 5.028 BPW budget allows
up to ~18.7% real tensor coverage (67 of 360) before the solve goes genuinely infeasible.
Lexicographic (--hessian-primary) solves Hessian danger first, with KL only allowed to
break real leftover ties afterward — the same two-phase mechanism this project already uses for
--raw-tiebreak, just with Hessian promoted to primary. Alone, it beats the soft weight broadly
(108 most-dangerous tensors reaching full precision: 32 vs. 16) but does not guarantee the single worst
tensors specifically — its objective sums danger across all tensors at once, so it can still trade the literal
worst one for a better aggregate. That's exactly what the floor is for; the two are complementary.
Floor for the guarantee, lexicographic for everything past it: 29 of the top 54 most dangerous tensors reach full 16-bit (vs. 22 for floor alone), with the same zero-cut guarantee for all 54. Real, measured: KL ends up deciding 0 of the 360 real Hessian-scored tensors, and only 3 of 497 total (both among the 137 tensors with no Hessian data at all).
How a raw Hessian score becomes a percentile rank, a danger weight and finally a bit-width, why
--hessian-primary gives 16-bit to small tensors rather than to the most dangerous ones, and why
unscored tensors are pushed down, is written up with diagrams and checked numbers in
How the Solver Turns a Hessian Score into a Bit-Width.
Nothing above is changed; this is the mechanism reference for it.
Independent code review (2026-09-15) confirmed the core design correct and properly gated (every new flag
off by default, zero output change when omitted — re-verified after every fix). It found and fixed two real
issues: the floor's own audit field originally reported requested tiers rather than actually-shipped bits
(now reports real shipped bits plus an explicit violations list — [] on the tested run), and
Phase 1's gap tolerance was originally looser than the pin's own framing implied (fixed via
--hessian-primary-gap, default 1e-7, independent of the general
--quality-mip-rel-gap — the tested run now proves gap=0.0, not an accepted
approximation).
L0 / L63 / lm_head: always both floor and weight, in every configuration — this never
changes based on whether Hessian data is available.
Regional floors (early-QKV, late-attn): off when running on cascade KL alone, same as section 5's existing recommendation.
When full Hessian data is available: Hessian takes precedence via --hessian-primary +
--hessian-floor-tiers, KL becomes the secondary, tie-breaking signal only.
Open question, current working answer YES: should the regional floors stay on even when Hessian is primary, as an extra safety net until this is benchmark-validated? Not yet re-tested with both mechanisms active together at the same time — the real next step, alongside the actual model build and the same IFEval/MMLU regression checks section 6 already calls for.
--pareto none and --hessian-weight-strength 0 above are both real corrections, not
part of this section's original recommendation. The real default (--pareto meaningful) filters
candidate bits using isolated-KL dominance, which can silently remove the exact bit-width a Hessian floor
needs — real, checked: 55 of 497 tensors affected at 405/497 coverage, sometimes producing genuine MILP
infeasibility together with --hessian-primary's own re-pinning. The soft weight was separately,
directly isolated-tested and confirmed to change zero real output once the floor/primary are active. Full
evidence: design doc and
MILP-aware Hessian guide. The
open regional-floor question above this command is unaffected — still genuinely open.
The real cascade validation — P0 → P1 → P2 → P3, and the final setting
Isolated KL (P0) measures every tensor in a vacuum, assuming every other tensor stays bf16 — never true in the real deployed model. Cascaded KL (P1, P2, P3) exists to fix exactly that: each round measures every tensor against the real, already-quantized state the model actually has. The real question: does the cascade's own correction agree with the independent Hessian curvature signal, and where does it settle?
Real correlation: Hessian danger score vs. each round's own plan, all four rounds
| Round | Real Pearson r | Real Spearman r | Reading |
|---|---|---|---|
| P0 (isolated) | −0.1173 | −0.3597 | negative — agrees with Hessian |
| P1 (cascade, round 1) | +0.0505 | −0.1596 | negative — agrees with Hessian |
| P2 (cascade, round 2) | −0.0841 | −0.1389 | negative — agrees with Hessian |
| P3 (cascade, round 3) | −0.0579 | −0.1211 | negative — agrees with Hessian |
Negative correlation (a real, dangerous tensor — low score — getting more real bits) holds at every single round. The independent Hessian signal agrees with the cascade's own correction direction the whole way through, not just at P0.
Real plan stability, round to round
| Transition | Real tensors that change bits | % of the 497 real targets |
|---|---|---|
| P0 → P1 | 167 | 34% |
| P1 → P2 | 83 | 17% |
| P2 → P3 | 68 | 14% |
The big real correction happens at P0→P1. What follows shrinks each round — real, genuine stabilization, not noise. P1 is the real, correct stopping point: further cascade rounds (P2, P3) are real but diminishing, and not worth the real GPU cost of chasing further.
The real, final recommendation
Real workflow, confirmed: P0 (isolated, confirms the plan) → P1 (cascade, corrects and stabilizes it)
→ Hessian overlaid on P1, boundary protection (lm_head + layer 0 + layer 63) always on,
--no-protect-first-last never used, --hessian-weight-strength 2.0 — the minimal real
intervention that does genuine, Hessian-consistent work without falling into the strength≥3 cliff. Real output
plans: research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/plans_P1_both_off/plan_strength_2.json
(region floors off) and .../plans_P1_both_on/plan_strength_2.json (region floors on — the real
production-shape candidate). Full real commands, logs, and every supporting number: SWEEP_SUMMARY.md
in that same folder.
--help output and the real
360-tensor score file — none estimated.