← Back to index
Real, standalone guide · 2026-09-15

Making the solver listen to curvature, not just KL

Plain-language record of a real, tested fix: the model-quantization solver was leaving its own most-dangerous tensors under-protected even when real risk data said not to. This explains what was found, what was tried, what actually worked, and exactly how to turn it on.

THE GOAL

What this whole thing is actually for

One sentence: the solver had real data saying which tensors were dangerous to compress, but a real check showed it was still cutting the most dangerous ones anyway. This is the fix, and the proof it works.

The real problem, in plain terms

Every tensor in the model gets a real "Hessian danger score" (where data exists — 360 of 497 tensors have one). Low score = genuinely dangerous to compress. The solver was already told to care about this score by multiplying it into its normal cost signal — but a real, direct check found that even the single most dangerous tensor in the entire model still only got 8 of a possible 16 bits. Pushing that multiplier all the way up to 1000× barely changed anything for the worst tensors. That's the real problem this page fixes.

BEFORE THE DETAILS

Glossary — every term used below

If a word further down is unfamiliar, it's defined here first.

MILP (Mixed-Integer Linear Program)
The real math solver that decides, for every tensor, exactly how many bits it gets. It's given a budget and a cost signal, and it finds the one assignment that spends the budget for the lowest total cost.
KL (KL divergence)
The solver's main real cost signal — a measured number saying how much a tensor's output changes if you compress it. This is what the solver has always tried to minimize.
Hessian score
A second, independent real signal — how narrow and fragile a tensor's own internal math is, regardless of what KL says. Low score = dangerous. Comes free as a byproduct of a real correction pass, but only for tensors that pass actually measured.
Percentile rank
A tensor's position in the real, sorted list of all scored tensors, from 0 (most dangerous) to 1 (safest). Used instead of the raw score because the raw numbers span two orders of magnitude and would let one outlier dominate.
Soft weight
The solver's existing way of caring about Hessian danger: multiply a tensor's KL cost by a real number bigger than 1. The solver can still choose to ignore it if the underlying KL number is small enough — proven directly this session, even at a 1000× multiplier.
Hard floor
A real, absolute rule: "this tensor may not go below N bits," full stop. Unlike a soft weight, the solver has no way around it — it isn't a preference, it's a constraint.
Feasible / infeasible
Whether a real solve is even possible under the current rules and budget. Add too many hard floors and the real math simply has nowhere left to fit them — the solver reports "infeasible" and produces nothing.
Lexicographic solve
Solving in two real, separate steps instead of one blended one. Step 1 solves using only the Hessian signal, ignoring KL completely. Step 2 then only picks among the results that were already exactly as good on step 1 — KL can refine, never override.
BPW (bits per weight)
The real, fixed average budget across the whole model — e.g. 5.028 means "on average, just over 5 bits per number," even though individual tensors range from 4 to 16.
THE REAL EVIDENCE

How bad the problem actually was

Ranking every one of the 360 real Hessian-scored tensors from most to least dangerous, and checking what the real, already-benchmarked production plan actually did with the worst of them:

Danger tierKept at full 16-bitStill cut to ≤5-bit
Top 10 most dangerous (real)01
Top 20 most dangerous (real)06
Top 54 most dangerous (real)124

The single most dangerous real tensor measured anywhere in the model — layers.63.mlp.up_proj — only reached 8-bit. Not one of the top 20 reached full precision.

WHY A BIGGER MULTIPLIER DIDN'T FIX IT

The soft weight saturates — real, tested proof

Before building anything new, the obvious fix was tried first: just make the existing multiplier much bigger.

--hessian-weight-strength 2.0 (existing default) --hessian-weight-strength 1000 (500× stronger — real test run)
Danger tierKept ≥16-bit @2.0Kept ≥16-bit @1000
Top 10 most dangerous22 (identical)
Top 20 most dangerous22 (identical)
Real finding

A 500× jump in the multiplier changed zero outcomes in the top 10 or top 20 most dangerous tensors. The reason: these tensors' real, raw KL cost is so close to flat across bit-widths that multiplying an almost-zero number by 1000 is still almost zero, next to tensors with genuinely large KL cost competing for the same fixed budget. No multiplier fixes that — the mechanism itself needed to change.

FIX #1

A real, per-tensor hard floor

Instead of a preference the solver can out-vote, give the most dangerous tensors a real, absolute minimum — the same mechanism already proven to work for lm_head and the first/last layers, just computed per-tensor from each one's own real danger rank instead of one flat rule for one named group.

--hessian-floor-tiers "0.05:8,0.15:6"

Meaning: any tensor in the real most-dangerous 5% gets a floor of 8-bit; the next tier out to 15% gets a floor of 6-bit. This is a genuine MILP constraint, not a weight — the solver cannot trade it away no matter how the rest of the budget competes.

Danger tierStill cut ≤5-bit, soft weightStill cut ≤5-bit, hard floor
Top 20 most dangerous170
Top 54 most dangerous440
Real feasibility ceiling

How far can this be pushed at a fixed budget before the real math runs out of room? Bisected it directly against the solver's own feasibility check: the real ceiling for this exact 5.028 BPW budget is ~18.7% real tensor coverage (67 of 360 scored tensors) — one step further (30% / 108 tensors) comes back genuinely infeasible.

FIX #2

Lexicographic priority — like a dictionary, not a blend

A floor guarantees the worst tensors specifically, but says nothing about everything just below the floor line. This second, real mechanism gives Hessian a genuine first say over the whole model, with KL only ever breaking ties.

--hessian-primary
In plain terms

Think of how a dictionary sorts words: "car" comes before "cat" only because of the third letter — but only because the first two letters are already identical. If the first letters differed at all, the third letter would never even be looked at. A lexicographic solve works the same way: it solves once using only real Hessian danger, completely ignoring KL, and locks that result in. Only then does it look at KL — and only to pick between whichever specific allocations still tie for the best possible Hessian outcome. KL can refine; it can never override.

PHASE 1 Hessian decides for all 497 tensors PHASE 2 KL breaks real ties only where Phase 1 tied RESULT The final plan 497 tensors, bits set

Purple = Hessian deciding alone. Grey = KL's one real chance to act, only on tensors Phase 1 didn't already settle. Mint = the finished plan handed to the solver's normal budget check.

Tested alone, this real mechanism does noticeably better than the soft weight across the whole model (108 most-dangerous tensors reaching full precision: 32 vs. 16) — but a real, direct check found it does not, by itself, guarantee the single worst tensors specifically, because its own objective adds danger up across all tensors at once, and can still trade the very worst one for a better total. That's exactly what the floor above is for — the two are complementary, not interchangeable.

THE REAL RESULT

Floor + lexicographic together — the actual optimum

Every real combination was tested and compared against the exact same real, already-benchmarked production plan. This is the full, real picture:

Top 54 most dangerousProductionSoft weightFloor onlyLexicographic onlyCombined
Kept ≥16-bit19222029
Still cut ≤5-bit24440300
Why combined wins outright

The floor guarantees the individual worst tensors can't be sacrificed — that's the one thing neither the soft weight nor lexicographic-alone could do. Lexicographic then spends whatever real budget is left, past the floor, still Hessian-danger-first instead of handing it to KL. Neither piece alone gets both properties; combined gets both, at the same real, fixed 5.028 BPW budget the production model already used.

HOW MUCH DOES KL STILL MATTER?

Measured directly, not guessed

Two different signals decide how the model gets compressed: KL (the original signal, used for years) and Hessian (the new one, added in this work). The mechanism above makes Hessian go first, with KL only allowed to help afterward. That raises an obvious question: does KL still do anything at all now, or did Hessian just override it completely?

Here is exactly how the two versions of the plan can end up different. For any tensor where Hessian has a clear opinion, both versions pick the same bits — there's nothing left to disagree about. The only place they can differ is a tensor where Hessian has no opinion at all. There, the Hessian-only version just falls back on whatever the solver picks by default, with no real reasoning behind it — while the real version lets KL make that specific choice on purpose, using its own real measurement. Comparing the two plans, tensor by tensor, shows exactly how often that difference actually happens.

Real, measured result

Only 3 of 497 tensors (0.6%) differ between the two. Zero of the 360 real Hessian-scored tensors were affected by KL at all — every one of them was fully decided by Hessian danger alone. The only 3 tensors KL still had a say over are among the 137 tensors that have no real Hessian data in the first place.

This is a direct, structural consequence of using a real, high-precision danger rank as the first priority: because almost no two tensors ever tie exactly on that rank, the tie-break step (KL) is almost never actually needed. If more room for KL is wanted later, the real lever is the solve's own tie-break tolerance — loosening it would let KL decide among a wider set of "close enough" Hessian outcomes instead of the single exact best one.

NO REGRESSION — VERIFIED

Every existing build stays exactly the same

Both new mechanisms are new, optional command-line flags. Neither is on by default, and neither was inserted into the solver's existing logic — they only activate if explicitly requested.

FlagDefaultEffect when omitted
--hessian-floor-tiersoff (no floors)zero behavior change
--hessian-primaryoff (single-phase KL solve)zero behavior change
Real, direct proof

Ran the exact same real command twice — once before these changes existed, once after, with the new flags left out — and diffed the two output plans directly. 0 of 497 real tensor assignments differed. This was re-verified after every further edit made to the file.

This also means the real, practical use-case that prompted this is already covered: on any future build where collecting a real Hessian score is too GPU- or time-expensive to justify, simply don't pass --hessian-score-file (or either new flag) — the solver falls back to exactly its normal, existing, KL-only behavior. Nothing about these mechanisms requires Hessian data to exist; they simply have nothing to do without it.

INDEPENDENT VERIFICATION

Code review — checking the reasoning holds

Before treating any of this as production-ready, a fresh, independent review (no shared context with the work above) was run against the real code — the MILP constraint logic, the two-phase lexicographic solve, interaction with the existing soft weight, and the backward-compatibility claim above.

Verdict: core design sound, two real bugs found and fixed

The review confirmed the underlying math is correctly wired and genuinely opt-in (independently re-derived the same zero-regression result). It also found two real, concrete issues, both now fixed and re-verified:

1. Audit-trail bug — the floor's own report field showed the requested tier targets, not what the solver actually shipped, so a silent floor failure wouldn't have been visible. Fixed to report real shipped bits plus an explicit violations list — re-verified on the real combined plan: 0 violations, every one of the 54 floored tensors genuinely met its floor.

2. Loose tolerance — the "exact pin" step was allowed to lock in whatever incumbent the solver stopped at, which could have been measurably short of the true best answer. Fixed by giving that step its own, much tighter accuracy requirement — re-verified on the real combined plan: gap=0.0, a proven exact answer, not an accepted approximation. The KL-relevance measurement (3 of 497) was independently re-run under this tightened check and came back identical.

Real output was unaffected by either fix — both closed a gap in how trustworthy the proof was, not a wrong result.

WHAT'S NEXT

The real path from here

This page covers the solver only — nothing here has touched a real model yet. The full, exact, real command, everything else matching the already-benchmarked production checkpoint:

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \ 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \ --target-bpw 5.028075122774708 \ --candidates 4,5,6,8,16 \ --pareto none \ --group-size 64 \ --late-attn-min-bits 0 \ --early-qkv-min-bits 0 \ --lm-head-min-bits 6 \ --boundary-min-bits 6 \ --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \ --hessian-weight-strength 0 \ --hessian-floor-tiers "0.05:8,0.15:6" \ --hessian-primary \ --max-low-bit-run 999 \ --output <your_output_path>.json
Real, confirmed finding (2026-09-18) — --pareto none is required whenever a Hessian file drives allocation

--pareto defaults to meaningful, which removes any candidate bit-width that a tensor's own isolated-KL data says isn't a meaningful improvement over a cheaper one — a real, useful filter when KL is the only signal deciding allocation. But it has no idea a Hessian floor is about to demand a specific bit-width for a completely different, KL-independent reason. Real, checked data at 405/497 coverage: 55 real tensors had their Hessian-floor-required bit-width filtered out by this exact mechanism, forcing several of them to jump all the way to 16-bit instead of the tier they actually needed — and in some real runs, produced genuine MILP infeasibility once combined with --hessian-primary's own tight re-pinning. The failure rate gets worse as real Hessian coverage grows, not better, since more tensors acquire a real floor requirement over time. --pareto none keeps the full candidate range available for every tensor, so the Hessian mechanism can actually request the bit-width it wants. Confirmed real effect on this exact command: raw ΣKL dropped from 49.25 to 38.89 (production's own baseline is 37.45) once both this fix and --hessian-weight-strength 0 (see below) were applied together.

Also changed: --hessian-weight-strength 0, not the real default (2.0) — confirmed to make zero difference

The soft Hessian weight's entire original purpose (2026-09-13) was nudging the most dangerous tensors upward before the hard Hessian floor existed — and it was already proven insufficient for that job even at strength 1000. Now that the floor (and --hessian-primary) do that job directly and unconditionally, the soft weight is redundant. Directly tested, 2026-09-18: re-running the identical real command with --pareto none held constant and only --hessian-weight-strength toggled between its default (2.0) and 0 produced identical results — the soft weight changes nothing once the floor and --hessian-primary are active. It was only ever silently inflating the reported OptiQ-weighted ΣKL statistic, not affecting the actual allocation. Pass --hessian-weight-strength 0 explicitly to keep that statistic honest, but know that --pareto none is the real, confirmed lever — the weight was never doing anything real to begin with.

Why the weighted-KL number swings so hard: two real weights compound multiplicatively on a handful of tensors

Corrects an earlier, understated "up to 3×" claim here. Real, traced (2026-09-18): the reported OptiQ-weighted ΣKL is dominated by a tiny number of tensors — the ones that are both boundary-protected (--protection-multiplier, real default 100×, applies only to lm_head + layer 0 + layer 63) and Hessian-critical. Their real combined weight is protection_multiplier × (1+strength), not either alone: 300× at strength=2.0, 900× at strength=8.0, confirmed directly from real per-tensor data. Just the top 5 of 497 tensors accounted for 62% of the total at strength=2.0 and 73% at strength=8.0 — the statistic is barely measuring the other ~492 tensors at all. Raw ΣKL is unaffected by any of this (it's never weighted) and remains the one to trust for any real comparison.

Is --protection-multiplier itself also dead weight, like the Hessian soft weight? Checked directly (2026-09-18) — no, left as-is

Real, code-verified scope: --protection-multiplier only ever touches 16 of 497 real tensors on this model — lm_head (1) plus layer 0 and layer 63 (15, via identify_boundary_layers). A real identify_moe_protected_layers() function also exists in this same script for Mixture-of-Experts tensors (mlp.gate/router/shared_experts patterns) — checked directly against this model's real tensor list: zero matches, confirming this model is dense and that code path is currently dormant, not broken or misconfigured. Real, direct evidence that the multiplier still does something (unlike the Hessian weight): 5 of these 16 tensors sit measurably above their real floor minimum in a weighted run (one as much as 10 bits above, 16 vs. its 6-bit floor) — a fixed, shared BPW budget means pushing a few tensors above their floor necessarily redistributes bits away from the rest of the model, the same mechanism already proven for the Hessian floor's own 122-126 "tensors demoted to fund it" finding. Decision: leave it unchanged. It isn't broken, it preserves real MoE-readiness for a future model, and the only cost of not "cleaning it up" is a weighted-KL statistic nobody trusts anyway — not worth the risk of touching working, real infrastructure for a cosmetic gain.

Added 2026-09-19 — how the flags actually pick bit-widths

Why --hessian-primary allocates by danger per million parameters (so a tiny, nearly-safe tensor gets 16-bit while the single most dangerous tensor stays at its floor), how the percentile rank and danger weight are computed line by line, and what happens to tensors with no score: see How the Solver Turns a Hessian Score into a Bit-Width.

Once GPU capacity is free: close the real Hessian coverage gap (136 still-unscored tensors, see the companion probe guide), rebuild the plan with these real flags on, build the real model, and run the same real benchmark suite already used for every other candidate — checked directly against the current production model for regression, not just assumed. Only if that real result beats every prior model (KL-only, cascaded-KL, YAQA-corrected) with no regression does it become the new real baseline.

Author: Hakim Ghelab, VegaLaboratories LTD