← Back to index
Improvement ledger 08 · full technical debrief · 2026-09-14, updated 2026-09-15

The Hessian-aware hybrid solver — every flag, every command, every real result, in order

This page exists so nothing from this investigation gets lost. Every number here is either read directly from a real solver run this session, or computed live from the real 360-tensor score file that run produced. No strength sweep was performed — see 4.1 for the exact, honest scope of what was tested.

What this page is not: a claim that this feature is validated or ready to become anyone's default. It documents real mechanism and real solver behavior. The real quality question — does any of this make a shipped model better or worse — has not been tested yet (see section 6).
SECTION 1

The two real flags, exactly as implemented

Both live in 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py. Both are off by default — omit --hessian-score-file and the solver's output is byte-identical to before this feature existed (verified directly).

--hessian-score-file HESSIAN_SCORE_FILE Optional path to a real JSON {tensor_name: hess_score} map (e.g. exported from hessian_story_lib.load_hessian()). When given, multiplies each named tensor's structural_weight by a real, bounded factor derived from its score's percentile rank among all scored tensors -- the most dangerous tensor (lowest score) gets the largest boost, the safest gets none. Tensors with no real score in the file are left at factor 1.0 (unchanged). Default: no file, no effect. --hessian-weight-strength HESSIAN_WEIGHT_STRENGTH Real multiplier range applied via --hessian-score-file: the single most dangerous tensor's structural_weight is scaled by (1 + this), the safest by 1.0, linearly interpolated by real percentile rank in between. Only read when --hessian-score-file is given. (default: 2.0)
SECTION 2

Where the real score comes from

effective_rank(H) = trace(H)² / ∑(H²)   # yaqa_core.py hess_score(tensor) = mean( effective_rank(H_I)/dim_in , effective_rank(H_O)/dim_out )

A real spectral participation-ratio diagnostic of each tensor's own already-computed H_I/H_O (the same Hessians YAQA's own two-sided correction already builds — nothing new is measured). Low score = energy concentrated in few directions (narrow, dangerous). High score = spread broadly (forgiving).

Now automatic — and fixed to cover every real entry point (2026-09-14)

The raw per-tensor data was always free — every real YAQA correction pass already prints it. The first automation attempt hooked run_full_yaqa.sh only — a direct user question ("are we 100% sure this covers every real build?") caught that MAIN_RESUME_BUILD.sh, a real, existing script that calls 05_full_model_quantize.py --assemble-from-resume directly, bypasses run_full_yaqa.sh entirely and would have silently skipped it. The hook now lives inside 05_full_model_quantize.py's own main() instead — the one real code path every real assembly runs through, regardless of which script invokes it — writing <resume_dir>/hessian_scores.json as a non-blocking side effect of the real save. No manual step is needed for any future real build, from either entry point.

Exact command (what 05_full_model_quantize.py's own assembly step now runs automatically, and what you'd run by hand against any older, already-completed build)

cd .../scripts/yaqa_port python3 08_extract_hessian_scores.py \ ~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume \ --output research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json

Real, standalone, tested script — not a one-off snippet anymore (08_extract_hessian_scores.py); reuses the same, already-tested hessian_story_lib.load_hessian() extraction logic. Verified byte-for-byte identical to the original manual extraction. Real, permanent output (not a scratch file): research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json — 360 real tensors scored. lm_head is absent by real construction (its own two-sided Hessian needs 246.7GB and is never built) — not an omission, see section 5. Real scope note: the score is a property of the model's own real weights and calibration data, not of which plan built it — one real extraction is reusable for any future plan against the same model.

SECTION 3

The exact granularity mechanism — real numbers, not the abstract version

This is the direct answer to "how is this actually more granular than a flat 100× score" — three real steps, then five real tensors run through them.

Step 1

Rank all 360 real scored tensors by their real hess_score, most dangerous (lowest) to safest (highest).

Step 2

Convert each tensor's rank position into a percentile r from 0 (most dangerous) to 1 (safest).

Step 3

Compute its real weight: factor = 1.0 + strength × (1 − r).

Real tensorReal hess_scorePercentile rank rReal weight factor (strength=2.0)
layers.63.mlp.up_proj (most dangerous of 360)0.0001610.0003.000×
layers.25.mlp.up_proj0.0004990.2512.499×
layers.15.self_attn.o_proj (dead middle)0.0011300.5011.997×
layers.36.mlp.up_proj0.0020000.7521.496×
layers.55.self_attn.k_proj (safest of 360)0.0129791.0001.000×
The direct contrast

Five genuinely different real weights, out of 360 total — every scored tensor gets its own precise number. The old flat 100× system would have given all five of these tensors 1.0 (none are named lm_head or a boundary layer) — completely blind to the fact that layers.63.mlp.up_proj is objectively far more dangerous than layers.36.mlp.up_proj by real measured curvature. This weight multiplies straight into the MILP's real objective (weighted_kl = kl × structural_weight) — same slot the flat 100× used to occupy, continuous input instead of binary.

SECTION 4

All four real experiments — exact commands, exact results

4.1 — exact scope, stated plainly: only two real strength values were ever run: 2.0 (the flag's own coded default) and 8.0 (one deliberate second point, chosen to see how the effect scales). This is not a systematic sweep across 1–8, and no optimization was performed to justify 2.0 as correct — it is a sensible default pending the real eval in section 6, nothing more.
#Real config (target-bpw/group-size/candidates held constant)StrengthCompared againstReal result
Alate-attn=0, early-qkv=0, lm-head=6 (no Hessian file)—real shipped production plan69/497 differ. lm_head self-promotes 6→16, zero Hessian involved.
Bsame floors as A + Hessian2.0plan AOnly 3/497 differ from A. lm_head unchanged (16). Real move: layers.55.self_attn.k_proj: 16→4 — matches the original role-based recommendation.
Csame floors as A + Hessian8.0plan A40/497 differ from A. Broad promotion of layers 34–62. lm_head drops 16→6 — too much signal dilutes its own specialness.
Dlate-attn=5, early-qkv=6 (production), lm-head=6, --no-protect-first-last + Hessian2.0real shipped production plan34/497 differ. lm_head self-promotes again, 6→16, floors still active.

Experiment A — pure cascaded KL, no floors, no Hessian (the zero-Hessian baseline)

cd .../qwen38-27b-heretic-ara python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \ 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \ --target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \ --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \ --output .../plan_v4_round1_NOFLOORS.json

Experiment B — Hessian at the default strength (2.0), floors still off

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \ 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \ --target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \ --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \ --hessian-score-file research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \ --hessian-weight-strength 2.0 \ --output /tmp/plan_v4_round1_NOFLOORS_HYBRID.json

Experiment C — same, strength cranked to 8.0 (the real cautionary result)

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \ 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \ --target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \ --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \ --hessian-score-file research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \ --hessian-weight-strength 8.0 \ --output /tmp/plan_v4_round1_NOFLOORS_HYBRID_STRONG.json

Experiment D — the actual "hybrid" configuration: real production floors kept ON

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \ 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \ --target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \ --late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 \ --no-protect-first-last \ --hessian-score-file research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \ --hessian-weight-strength 2.0 \ --output .../plan_v4_round1_HESSIAN_HYBRID.json
SECTION 5

Direct answer: should the regional floors stay ON or come OFF?

Keep them ON — do not disable

--late-attn-min-bits and --early-qkv-min-bits stay at their real production values (5 and 6). Experiments A/B/C prove the mechanism works and produces defensible reallocation — they do not prove removing the floors is safe. Nothing here ran the real IFEval/MMLU eval that would actually answer that. Each floor was added because of a real, measured historical regression.

Superseded — see section 7

--no-protect-first-last is never used, full stop. An earlier draft of this section recommended it as a substitution for the flat 100× boundary category — that recommendation is retracted. It only touches 16 tensors (lm_head + layer 0's 8 + layer 63's 7) that should stay protected no matter what, for no real benefit. The real, current recommendation instead removes the two much larger region floors (124 tensors each) once real cascade measurement (P1) is used — see section 7 for the full real reasoning and the real, final command.

lm_head is never touched by any of this

Structurally exempt: 0 real entries for language_model.lm_head in the real score file, because its own two-sided Hessian would need 246.7GB and is never built. Its hard floor (--lm-head-min-bits 6) is the only thing that has ever reliably protected it, and stays on regardless of any Hessian setting.

SECTION 6

Where this actually stands right now

shipped, default-onCascaded KL measurement

Unchanged, production signal since before this investigation.

shipped, default-onlm_head hard floor

A genuine MILP constraint, unaffected by Hessian settings.

shipped, default-onBoundary hard floor, layer 0 + layer 63 (new, 2026-09-14)

A real, second protection layer for the boundary category — the existing flat 100× weight alone was proven insufficient (real ablation: L0's mlp.up_proj/gate_proj still dropped to Q4/Q5 with the weight on and no competing floor). --boundary-min-bits gives L0/L63 the same floor+weight combination lm_head already had.

built, tested, opt-in onlySoft Hessian weighting (--hessian-weight-strength)

Real, verified, zero-regression when omitted. Superseded as the primary mechanism — see section 8: even at strength=1000 it changed zero outcomes in the top 20 most dangerous real tensors.

built, tested, opt-in onlyPer-tensor Hessian hard floor (--hessian-floor-tiers, new, 2026-09-15)

Real MILP minimum-bits constraint from each tensor's own danger rank. Real, confirmed: 0 of the top 54 most dangerous tensors cut to ≤5-bit, vs. 24-44 under every weight-only approach.

built, tested, opt-in onlyLexicographic Hessian-primary solve (--hessian-primary, new, 2026-09-15)

Real two-phase solve — Hessian decides first, KL only breaks real leftover ties. Real, measured: KL ends up deciding 0 of 360 scored tensors, 3 of 497 total.

The one step not yet taken

Everything on this page is real solver output — not yet a proven quality result. The honest next building block: build a real model from the real recommended plan (section 8) and run it through the real IFEval/MMLU regression checks that motivated the original floors, before this becomes anyone's default.

SECTION 8 · 2026-09-15

Hard floor + lexicographic priority — closing the gap this whole debrief left open

Section 3/4 above showed the soft percentile weight is a real, continuous, per-tensor signal — a genuine improvement over the flat 100× category weight. This section shows it is still not a guarantee, and what actually closes that gap. Full real page, every diagram and every number: research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/MILP_AWARE_HESSIAN_GUIDE.html. Full command reference: RUNNING_GUIDE.md §11.8.

The soft weight saturates — real, tested proof

Pushed --hessian-weight-strength from the tested default (2.0) to 1000 — a 500× jump. Result: zero outcomes changed in the top 20 most dangerous real tensors in the model. Even the single most dangerous real tensor measured anywhere (layers.63.mlp.up_proj) only reached 8-bit either way. The real cause: these tensors' raw KL cost is close enough to flat across candidate bit-widths that multiplying an already-small number by 1000 is still small next to tensors with genuinely large KL cost, under the same fixed BPW budget. No multiplier fixes a signal too small to matter in the objective it's added to.

Top 54 most dangerousProductionSoft weight (2.0)Hard floor onlyLexicographic onlyCombined
Kept ≥16-bit19222029
Still cut ≤5-bit24440300

Hard floor (--hessian-floor-tiers "0.05:8,0.15:6") is a real MILP minimum, not a preference — the solver cannot trade it away. Zero of the top 54 most dangerous tensors land at ≤5-bit, where the soft weight left 44. Real feasibility ceiling found by bisection: this exact 5.028 BPW budget allows up to ~18.7% real tensor coverage (67 of 360) before the solve goes genuinely infeasible.

Lexicographic (--hessian-primary) solves Hessian danger first, with KL only allowed to break real leftover ties afterward — the same two-phase mechanism this project already uses for --raw-tiebreak, just with Hessian promoted to primary. Alone, it beats the soft weight broadly (108 most-dangerous tensors reaching full precision: 32 vs. 16) but does not guarantee the single worst tensors specifically — its objective sums danger across all tensors at once, so it can still trade the literal worst one for a better aggregate. That's exactly what the floor is for; the two are complementary.

Combined — the real optimum, both properties at once

Floor for the guarantee, lexicographic for everything past it: 29 of the top 54 most dangerous tensors reach full 16-bit (vs. 22 for floor alone), with the same zero-cut guarantee for all 54. Real, measured: KL ends up deciding 0 of the 360 real Hessian-scored tensors, and only 3 of 497 total (both among the 137 tensors with no Hessian data at all).

Added 2026-09-19 — the mechanism, explained step by step

How a raw Hessian score becomes a percentile rank, a danger weight and finally a bit-width, why --hessian-primary gives 16-bit to small tensors rather than to the most dangerous ones, and why unscored tensors are pushed down, is written up with diagrams and checked numbers in How the Solver Turns a Hessian Score into a Bit-Width. Nothing above is changed; this is the mechanism reference for it.

Independent code review (2026-09-15) confirmed the core design correct and properly gated (every new flag off by default, zero output change when omitted — re-verified after every fix). It found and fixed two real issues: the floor's own audit field originally reported requested tiers rather than actually-shipped bits (now reports real shipped bits plus an explicit violations list — [] on the tested run), and Phase 1's gap tolerance was originally looser than the pin's own framing implied (fixed via --hessian-primary-gap, default 1e-7, independent of the general --quality-mip-rel-gap — the tested run now proves gap=0.0, not an accepted approximation).

The real, current strategy going into the next build

L0 / L63 / lm_head: always both floor and weight, in every configuration — this never changes based on whether Hessian data is available.

Regional floors (early-QKV, late-attn): off when running on cascade KL alone, same as section 5's existing recommendation.

When full Hessian data is available: Hessian takes precedence via --hessian-primary + --hessian-floor-tiers, KL becomes the secondary, tie-breaking signal only.

Open question, current working answer YES: should the regional floors stay on even when Hessian is primary, as an extra safety net until this is benchmark-validated? Not yet re-tested with both mechanisms active together at the same time — the real next step, alongside the actual model build and the same IFEval/MMLU regression checks section 6 already calls for.

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \ 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \ --target-bpw 5.028075122774708 --candidates 4,5,6,8,16 --pareto none --group-size 64 \ --late-attn-min-bits 0 --early-qkv-min-bits 0 \ --lm-head-min-bits 6 --boundary-min-bits 6 \ --hessian-score-file research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \ --hessian-weight-strength 0 \ --hessian-floor-tiers "0.05:8,0.15:6" \ --hessian-primary \ --max-low-bit-run 999 \ --output <your_output_path>.json
Real, confirmed finding (2026-09-18) — added since this section was written

--pareto none and --hessian-weight-strength 0 above are both real corrections, not part of this section's original recommendation. The real default (--pareto meaningful) filters candidate bits using isolated-KL dominance, which can silently remove the exact bit-width a Hessian floor needs — real, checked: 55 of 497 tensors affected at 405/497 coverage, sometimes producing genuine MILP infeasibility together with --hessian-primary's own re-pinning. The soft weight was separately, directly isolated-tested and confirmed to change zero real output once the floor/primary are active. Full evidence: design doc and MILP-aware Hessian guide. The open regional-floor question above this command is unaffected — still genuinely open.

SECTION 7

The real cascade validation — P0 → P1 → P2 → P3, and the final setting

Isolated KL (P0) measures every tensor in a vacuum, assuming every other tensor stays bf16 — never true in the real deployed model. Cascaded KL (P1, P2, P3) exists to fix exactly that: each round measures every tensor against the real, already-quantized state the model actually has. The real question: does the cascade's own correction agree with the independent Hessian curvature signal, and where does it settle?

Real correlation: Hessian danger score vs. each round's own plan, all four rounds

RoundReal Pearson rReal Spearman rReading
P0 (isolated)−0.1173−0.3597negative — agrees with Hessian
P1 (cascade, round 1)+0.0505−0.1596negative — agrees with Hessian
P2 (cascade, round 2)−0.0841−0.1389negative — agrees with Hessian
P3 (cascade, round 3)−0.0579−0.1211negative — agrees with Hessian

Negative correlation (a real, dangerous tensor — low score — getting more real bits) holds at every single round. The independent Hessian signal agrees with the cascade's own correction direction the whole way through, not just at P0.

Real plan stability, round to round

TransitionReal tensors that change bits% of the 497 real targets
P0 → P116734%
P1 → P28317%
P2 → P36814%
Real conclusion

The big real correction happens at P0→P1. What follows shrinks each round — real, genuine stabilization, not noise. P1 is the real, correct stopping point: further cascade rounds (P2, P3) are real but diminishing, and not worth the real GPU cost of chasing further.

The real, final recommendation

Real workflow, confirmed: P0 (isolated, confirms the plan) → P1 (cascade, corrects and stabilizes it) → Hessian overlaid on P1, boundary protection (lm_head + layer 0 + layer 63) always on, --no-protect-first-last never used, --hessian-weight-strength 2.0 — the minimal real intervention that does genuine, Hessian-consistent work without falling into the strength≥3 cliff. Real output plans: research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/plans_P1_both_off/plan_strength_2.json (region floors off) and .../plans_P1_both_on/plan_strength_2.json (region floors on — the real production-shape candidate). Full real commands, logs, and every supporting number: SWEEP_SUMMARY.md in that same folder.

Author: Hakim Ghelab, VegaLaboratories LTD · Companion to IMPROVEMENT_LEDGER/04, 05, 06, 07 and RUNNING_GUIDE.md section 11 · Every command and number on this page is real, run this session, cross-checked against the real script's own --help output and the real 360-tensor score file — none estimated.