Mechanism reference · built 2026-09-30 01:14 · every number recomputed from live data
From a Hessian score to a bit-width: how the solver decides
The Hessian sweep hands the solver one number per tensor. This page follows that number step by step until it becomes a choice between 4, 5, 6, 8 or 16 bits, and shows what that choice does on the real plans.
How to read this page. Every term is explained where it first appears, and every calculation is shown so you can redo it by hand. Every number and diagram below is recomputed from the live Hessian score file (496 of 497 tensors scored at the last build, 2026-09-30 01:14) and from real solver runs, so it moves as the sweep progresses. The table “What has changed since the plans were first run” in section 7 shows the difference against the moment the plans were first run (407 scored), and the box below lists what changed since the previous build.
Where this sits
This is a mechanism reference, not part of the story. The story lives in the ledger: start at §08. The fixed commands are in the MILP-aware guide and RUNNING_GUIDE §11.8. This page explains why those flags behave the way they do.
Start here: what this page is about
Written so that an engineer can check every step and a non-specialist can follow the logic. On a first read you can skip the equations: the sentence under each one says the same thing in words.
A language model is mostly numbers. This one has about 25.6 billion of them (its weights), stored in 497 large tables called tensors. Storing every number with 16 bits is accurate but heavy. Storing them with 4 to 8 bits (quantization) makes the model several times smaller (and usually cheaper to run), at some cost in accuracy.
That cost is not spread evenly. Some tensors hardly notice being squeezed. Others are fragile: squeeze them and the model’s answers get worse. The model also has a fixed size limit (this project aims for an average of about 5.03 bits per number). The solver’s job is to decide, for each of the 497 tensors, whether it gets 4, 5, 6, 8 or 16 bits, so that the limit is respected and the damage is as small as possible.
Think of packing 497 items into luggage with a fixed weight allowance. Fragile items get heavy padding, sturdy ones almost none, and the allowance forces you to choose.
The pipeline in one picture. This page is about the solver box: how it turns the fragility score into a choice of bit-width. The last stage, building the compressed model and measuring its real quality, has not been done for any of the plans discussed here.
Fragility score
Also called the Hessian score
A measure, computed from calibration data, of how sharply the model’s error rises when a tensor’s numbers are disturbed. A low score means fragile.
KL damage measurement
Kullback–Leibler divergence
A measured number: how much the model’s output changes when only this tensor is squeezed to a given bit-width. Larger means more damage.
Solver
A mixed-integer linear program (MILP)
An exact optimizer. It is given rules and a target and finds the best whole-number choices. Here it uses the open-source HiGHS engine.
Where things stand
Verified The mechanism on this page. An independent re-solve reproduces a real plan exactly, tensor for tensor.
Not measured Whether any of these plans gives a better model. No plan compared here has been built and benchmarked yet.
Incomplete Fragility scores exist for 496 of 497 tensors (as of 2026-09-30; 407 when the plans were first run). Some behaviour below depends on that gap; an opt-in fix for the unscored tensors exists (section 5.6).
What would settle it Finish the score sweep, build the floor-plus-KL plan and the floor-plus-primary plan, and benchmark both against the production model.
Live status · built 2026-09-30 01:14
Coverage:496 of 497 tensors have a fragility score (99.8%); 1 do not: 0 are the tiny in_proj_a/in_proj_b gates and 1 are others, together 5.0% of all parameters. Scored count over time: 479 → 481 → 483 → 485 → 488 → 496.
Sweep process: running (PID 98682, up 22:24).
Checks on this build
✗ Every tensor left at its floor has lower danger per million parameters (0.0411 at most) than every tensor that reached 16-bit (0.0277 at least): 146 at 16-bit, 347 at their floor, 4 in between.
✗ Plain primary squeezes most unscored tensors to 4-bit (0 of 1).
✗ The KL fallback accounts for 0% of the raw-KL gap between plain primary and production (0.0 of 2.8 points).
✓ An independent re-solve of phase 1 differs from the real solver on 0 tensors (C2) and 0 tensors (C3) out of 497.
Every number and diagram on this page is recomputed from the live score file and real solver runs by hessian_primary_page.py whenever the score file, the solver or the sweep changes. Plan B2 is the one exception: its exact command is not known, so it stays frozen at the 407-tensor run and is labelled so.
What changed since the previous build
Nothing measurable changed since the previous build.
The exact command, first
Most people learn a tool from one complete example. This is the recommended command, exactly as you would type it, then what each line means, then what it prints for real. Everything after this section explains why it behaves the way it does.
Copy and paste this. It only reads files and writes one new plan file; it does not build the model.
The solver: the program that chooses a bit-width for every tensor.
…/cascaded_checkpoint_round1.json
Its input: the measured KL damage of every tensor at every candidate bit-width (497 tensors × 5 bit-widths).
--target-bpw 5.0280…
The size budget: the average number of bits per weight to aim for. This is the value of the production build, so plans compare like for like.
--candidates 4,5,6,8,16
The bit-widths a tensor may be given.
--pareto none
Keep every candidate bit-width. Required whenever a Hessian score file is used; the default silently removes bit-widths a floor needs and can make the problem unsolvable.
--group-size 64
Recorded in the plan for the later build step: how many numbers share one scale when a tensor is stored. It does not change which bit-width the solver picks.
--late-attn-min-bits 0 --early-qkv-min-bits 0
Two regional floors (a minimum number of bits for attention tensors in the last quarter of the model, and for the query/key/value tensors in the first quarter). 0 switches each one off, as in plan C2.
--lm-head-min-bits 6
The output layer, lm_head, is never given fewer than 6 bits.
--boundary-min-bits 6
The tensors in the first and last layer (0 and 63) are never given fewer than 6 bits.
--hessian-score-file ….json
The fragility score of every tensor the Hessian sweep has scored so far.
--hessian-weight-strength 0
Switches off the older soft weight. It does nothing once a floor or primary is active (section 1).
--hessian-floor-tiers "0.05:8,0.15:6"
The floor: the most fragile 5% of scored tensors get at least 8 bits, the next 10% at least 6 (section 4).
--hessian-primary
Decide the spare bits by fragility first and KL second (sections 5 and 6).
--hessian-unscored-fallback kl
Tensors with no fragility score borrow their own KL ranking instead of counting as “safest” (section 5.6).
--max-low-bit-run 999
Switches off a separate safety rule about long runs of low-bit blocks. Required alongside --hessian-primary.
--output ….json
Where the plan is written. A new file name, so no existing plan is overwritten.
What it prints, for real
This is the actual output of that command, from 2026-09-30 01:15 at 496 scored tensors (long lines shortened with …, the output path replaced by a placeholder). The two solves took about a second and a half.
[lm-head-floor] 1 lm_head tensor(s) floored at >= 6-bit (real fix, 2026-09-12 -- see IMPROVEMENT_LEDGER/04_LMHEAD_PROTECTION_INCIDENT.html)
[boundary-floor] 15 boundary tensor(s) (layer 0 + layer 63) floored at >= 6-bit (real fix, 2026-09-14 -- the flat 100x weight alone was not enough, see the flag's own help text)
[hessian-floor] real MILP-aware Hessian floor (2026-09-15): 0.05:8,0.15:6 -- 75 real tensor(s) floored, by tier bits: {6: 50, 8: 25}
[hessian-unscored-fallback] kl: 1 of 1 unscored tensor(s) given the percentile rank of their own isolated KL at 4-bit as danger weight (instead of 0.0 = safest); scored tensors and floors ...
[hessian-primary] real lexicographic solve ENABLED -- Phase 1 optimizes danger-weighted bit allocation for 496 real Hessian-scored tensor(s) plus 1 KL-fallback tensor(s) with zero KL infl ...
[P0_hessian_primary] optimal_or_gap_target_reached | 0.78s | gap=0.0 | nodes=893
[Q_global_weighted_KL] optimal_or_gap_target_reached | 7.68s | gap=1.512286573675527e-05 | nodes=684
Tensors: 497
Requested BPW: 5.028075123
Hard quality ceiling BPW: 5.028575123
Selected / achieved BPW: 5.028573895
Raw Σ isolated KL: 40.272079778
Q4 : 235 tensors | 69.17% param mass
Q5 : 59 tensors | 8.55% param mass
Q6 : 35 tensors | 13.13% param mass
Q8 : 22 tensors | 5.22% param mass
Q16: 146 tensors | 3.93% param mass
Wrote: <the --output path you gave>
How to read the output
[hessian-floor] … 75 tensors floored: 25 tensors got an 8-bit floor and 50 a 6-bit floor.
[hessian-unscored-fallback] kl: 1 of 1: every tensor still missing a score borrowed its KL ranking.
[P0_hessian_primary] … gap=0.0: phase 1, the fragility solve, finished with the exact optimum. [Q_global_weighted_KL] is phase 2, the KL tie-break.
Selected / achieved BPW 5.028573895 must be at or under the ceiling 5.028575123: the size budget was respected.
Raw Σ isolated KL 40.27 is the damage estimate compared in section 7 (production: 37.45). Q16: 146 tensors is how many tensors were kept at full 16-bit precision.
The complete commands for the other two variants, floors plus KL and plain primary without the fallback, are in section 8. Nothing here builds or benchmarks a model; that is a separate step.
1 · The short answer
The solver receives one fragility (Hessian) score per tensor and can use it in three separate ways, selected by three command-line flags. They differ in when they act and in what the solver may still trade away.
--hessian-floor-tiers
A floor: a rule
Acts first, before any optimizing. It says “this tensor may never go below N bits.” The solver cannot trade it away for anything.
It only sets a minimum. It does not say how far above the minimum a tensor should go. In the current plans it covers 75 of 496 scored tensors (25 at 8-bit, 50 at 6-bit).
--hessian-weight-strength
A weight: a preference
Makes a tensor’s KL cost look bigger, so the solver is more reluctant to damage it. It is a preference, so it can still be traded away.
Measured: changes nothing once a floor or primary is active, and it inflates the reported weighted KL. Keep it at 0.
--hessian-primary
A different question
Changes what the solver optimizes first. Instead of “which allocation has the smallest KL?”, it asks “which allocation leaves the least Hessian danger unprotected?” KL only breaks ties.
It acts on the spare bits above the floors: who gets 16, who stays low.
How does primary choose between 4, 5, 6, 8 and 16?
Not tensor by tensor. Every tensor, judged alone, would take 16-bit, so its own candidates never make the choice. The choice comes from a shared budget:
Every tensor starts at its floor. That uses 4.6308 bits per weight; the ceiling is 5.0286, leaving 0.3978 to spend.
Upgrading a tensor to 16-bit removes danger in proportion to its danger weight, and costs budget in proportion to its size. The solver spends the spare budget on the tensors with the most danger per million parameters.
In plan C2 the split between 16-bit tensors and floor tensors no longer follows danger per million parameters cleanly; section 5.4 shows the check.
Two surprises worth knowing before you read on
The single most dangerous tensor in the model (layers.63.mlp.up_proj) does not get 16-bit under primary. It stays at 8-bit. A nearly-safe tiny tensor (layers.1.linear_attn.in_proj_b, 245,760 parameters) does get 16-bit.
Tensors with no score get danger weight 0, and on that scale 0 means safest, so primary pushes them down to their floors. 0 of the 1 unscored tensors end at 4-bit, where floor-plus-KL keeps 0 of them at 16-bit. This is fixable with an opt-in flag (section 5.6).
2 · Three numbers that all get called “danger”
“Danger” here just means fragility: a dangerous tensor is one that is easily damaged by squeezing. The code and the older docs use that one word for three different numbers, and they do not point the same way. That is the main source of confusion, so they get separate names here.
Raw score
Low means dangerous
The number from the Hessian sweep (hess_score): how spread out a tensor’s curvature is. A low score means the curvature is concentrated in a few directions, which is fragile.
Real range: 0.000161 to 0.1780, a 1103-fold span. How it is computed: Hessian Score — Origin.
Rank r
Low means dangerous
Sort all scored tensors from lowest raw score to highest. A tensor’s position i is its place in that list, counting from 0. Then r = i ÷ (n − 1), a number from 0 to 1.
r = 0 is the most dangerous tensor, r = 1 the safest. The floor uses r.
Danger weight w
High means dangerous
Simply w = 1 − r. The direction flips: w = 1 is the most dangerous tensor, w = 0 the safest. Primary uses w.
The solver calls this hessian_danger_weight.
The same 496 tensors, sorted from lowest to highest raw score, drawn three ways. The top curve is the raw score on a log scale. It climbs slowly and then steeply, because the values span 312-fold. The middle line is the rank, and the bottom line is the danger weight: both are straight lines, which is why the solver uses them instead of the raw score. Four real tensors are marked:
1layers.63.mlp.up_proj (the most dangerous),
2layers.1.linear_attn.in_proj_qkv,
3layers.51.mlp.up_proj (in the middle),
4layers.1.linear_attn.in_proj_b (almost the safest).
3 · From score to rank, line by line
These are the real lines from the solver (compute_hessian_danger_weight, lines 339–345 of 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py). The floor function repeats the same first three lines.
scored_names = [n for n in names if n in raw_scores] # tensors that HAVE a score
ranked = sorted(scored_names, key=lambda n: raw_scores[n]) # lowest raw score first
n = len(ranked) # how many tensors are scored
rank_of = {name: i / max(n - 1, 1) for i, name in enumerate(ranked)} # r = i / (n - 1)
return {name: 1.0 - rank_of[name] for name in scored_names} # w = 1 - r
A worked example you can check with a calculator
Take tensor 2layers.1.linear_attn.in_proj_qkv. In the sorted list of 496 scored tensors it sits at position i = 10 (10 tensors have a lower raw score). Here n = 496, so n − 1 = 495.
r = i ÷ (n − 1) = 10 ÷ 495 = 0.0202 w = 1 − r = 1 − 0.0202 = 0.9798
Tensor
Position i
Rank r = i ÷ 495
Danger weight w = 1 − r
1
layers.63.mlp.up_proj
0
0.0000
1.0000
2
layers.1.linear_attn.in_proj_qkv
10
0.0202
0.9798
3
layers.51.mlp.up_proj
212
0.4283
0.5717
4
layers.1.linear_attn.in_proj_b
427
0.8626
0.1374
Why rank instead of the raw score?
The raw scores span 1103-fold, from 0.000161 to 0.1780. Used directly, a single extreme value would dominate everything else. Rank squeezes every tensor into the same 0–1 range no matter how the raw values are spread. The cost is that rank only records order: the gap between two neighbouring tensors is lost.
Three properties of this calculation that are easy to miss
Ties are broken by file order, not by data
34 groups of tensors (75 tensors in all) share an identical raw score. Python’s sorted keeps tied tensors in the order they appear in the checkpoint file, so among equals the position, and therefore the rank, comes from list order.
Ranks move when coverage grows
n is the number of scored tensors, not 497. At 407 scored (when the plans were first run), tensor 2 had r = 0.0246 and w = 0.9754. Now, with 496 scored, the same tensor has r = 0.0202 and w = 0.9798. The tensor did not change; the list around it did. Percentages such as “the bottom 5%” therefore pick different tensors as the sweep progresses.
No score means weight 0, and 0 is the SAFEST end
A tensor missing from the score file gets danger weight w = 0 (the code comment: silence means no signal). On the danger-weight scale 0 is the safest end, and 1 is the most dangerous. So a missing score does not make a tensor look risky. It makes it look perfectly safe, and primary under-protects it. Do not confuse this with the raw score, where a low value means dangerous. Section 5.5 shows what it does to the 1 tensors that are unscored now, and section 5.6 the fix. Right now 2 tensors carry exactly 0: the 1 unscored ones, plus the one scored tensor with the highest raw score (its rank is exactly 1.0, so its weight is 1 − 1 = 0).
4 · The floor: a rule, not a preference
--hessian-floor-tiers "0.05:8,0.15:6" reads as: rank under 0.05 gets at least 8-bit; rank under 0.15 gets at least 6-bit.
The solver checks the tiers in order and the first one the rank falls under wins. With n = 496 the two thresholds land at positions 0.05 × 495 = 24.8 and 0.15 × 495 = 74.2. Positions 0 to 24 are below the first threshold, so 25 tensors get the 8-bit floor; positions 25 to 74 get the 6-bit floor, so 50 tensors; the other 421 scored tensors get no Hessian floor, and neither do the 1 unscored ones (they have no rank), so 422 of the 497 tensors have no Hessian floor now. That number falls as more tensors are scored, but slowly: the tiers are percentages, so at full coverage only 75 tensors get a floor and the other 422 never do. “No Hessian floor” does not mean 0 bits: those tensors may still go as low as 4-bit and are still held by the other floors below.
A floor is a step function of rank. It is added to the solver as a hard constraint (add_min_bits_constraint) before any optimizing. Whatever the objective wants, the tensor cannot go below its floor. The floor never says “exactly”; a floored tensor can still be raised.
Ties and the floor cutoffs
Right now no group of tensors with exactly the same raw score straddles a floor cutoff, so the ordering of ties does not change who gets a floor. It did when the plans were first run, and it can again as coverage grows.
The floor is not the only floor
Each tensor’s effective minimum is the highest of every floor that applies to it. In the tested plans (C2) those are:
Floor
Applies to
Minimum
Hessian tiers (this section)
75 tensors, by rank
8 or 6
Boundary protection
15 tensors in layers 0 and 63
6
lm_head protection
1 tensor
6
Vendor per-block rule (the OptiQ structural policy)
128 tensors, chosen from each block’s own KL data
5
Everything else
the lowest allowed candidate
4
5 · Primary, phase 1: how it picks a bit-width
This is the mechanism behind --hessian-primary, in full. A MILP (mixed-integer linear program) is a problem where you choose whole-number decisions to make one number as small as possible while obeying linear limits. The solver, HiGHS, finds the exact best answer.
5.1 What the solver chooses, and the three rules
For every tensor t and every candidate bit-width b in {4, 5, 6, 8, 16} there is one yes/no variable xt,b: it is 1 if tensor t is given b bits, and 0 otherwise. That is 497 × 5 = 2,485 variables, and the solver sets each to 0 or 1. The symbol Σ just means “add up”.
Σbxt,b = 1 for every tensor tRule 1, in words: each tensor gets exactly one bit-width. The five yes/no variables of a tensor must add up to 1.
xt,b = 0 whenever b < floortRule 2, in words: a tensor cannot get fewer bits than its floor (section 4).
Σt Σbxt,b · b · Pt ÷ Ptotal ≤ ceiling
Rule 3, in words (the budget): multiply each tensor’s bits by its size, add everything up and divide by the model’s total size. That average number of bits per weight (BPW) must stay under the ceiling. Pt is the number of parameters in tensor t; Ptotal = 25,621,954,560. The ceiling is the target 5.0281 plus a 0.0005 slack, 5.0286.
5.2 What phase 1 tries to minimize
Each option (tensor t, bit-width b) gets a deficit: how much danger is left unprotected by choosing b instead of the maximum, 16 bits. A fragile tensor (large w) left at few bits has a large deficit.
deficit(t, b) = wt × (16 − b)
The real code line: hessian_danger_weight.get(name, 0.0) * (max_bits - bits)
Phase 1 minimizes Σt Σbxt,b · deficit(t, b)
In words: choose the bit-widths so that the total unprotected danger across all tensors is as small as possible. KL plays no part in this step.
For tensor 2 (w = 0.9798) the five options are:
Bit-width
16 − bits
× 0.9798 = deficit
4-bit
12
11.76
5-bit
11
10.78
6-bit
10
9.80
8-bit
8
7.84
16-bit
0
0.00
Deficit falls with every extra bit and never flattens. So looking at one tensor alone, 16-bit is always the cheapest option, for every tensor. That is the important point: the tensor’s own options never decide. Only the shared budget in rule 3 says “not everyone can have 16”.
5.3 Where the budget goes: value for money
Suppose a tensor is raised to 16-bit. Two things happen, and they are measured in different currencies:
What it buys: the tensor’s deficit disappears. That is its danger weight w times the bits added. A fragile tensor buys more.
What it costs: budget, in proportion to the tensor’s size: bits added × Pt ÷ Ptotal. A large tensor costs more.
The four marked tensors, each raised to 16-bit from its floor:
Each bar is scaled to the largest value in its own column. Tensor 4 buys very little (1.65) but costs almost nothing (0.0001 bits per weight). Tensor 1 buys a lot (8.00) but costs 242 times as much. Value = what it buys ÷ what it costs, and tensor 4 wins by a factor of 49.8.
value = (wt × bits added) ÷ (bits added × Pt ÷ Ptotal) = wt × Ptotal ÷ PtIn words: the number of bits added cancels, and Ptotal is the same for every tensor. So tensors are ranked by wt ÷ Pt: danger divided by size. Written per million parameters: danger per million parameters = wt ÷ (Pt ÷ 1,000,000). This is the number in the last column of the figure and in the scatter below.
The deficit has no size in it, but the budget does. So the solver does not favour the most dangerous tensors; it favours the ones where protection is cheapest per unit of danger. What actually happened to these four:
Tensor
Parameters
Danger weight w
Danger per million parameters
Floor
C2 (floor + primary)
A2 (floor + KL)
Production
1
layers.63.mlp.up_proj
89,128,960
1.0000
0.0112
8
8
16
16
2
layers.1.linear_attn.in_proj_qkv
52,428,800
0.9798
0.0187
8
8
8
8
3
layers.51.mlp.up_proj
89,128,960
0.5717
0.0064
4
4
4
4
4
layers.1.linear_attn.in_proj_b
245,760
0.1374
0.5590
4
16
16
6
Tensor 4 is almost the safest in the model, but it is tiny, so under primary it gets 16-bit while tensor 1, the most dangerous, stays at 8-bit.
5.4 The budget as a shopping trip
Start every tensor at its floor. That plan uses 4.6308 bits per weight. The ceiling is 5.0286, so 0.3978 bits per weight are left to spend. Phase 1 spends them on upgrades to 16-bit, best value first, until the money runs out.
Reading the bar: the gray part is what the whole model costs if every tensor stays at its floor. The purple part is the spare allowance left under the ceiling, and it is all that can be spent on upgrades.
146 tensors fit; the small remainder buys 4 partial upgrades. So 493 of the 497 tensors are either at 16-bit or exactly on their floor. This is why the result looks all-or-nothing: a shopping list ranked by value for money, bought from the top until the money is gone. The scatter below shows every scored tensor and where the cut falls.
Every scored tensor in plan C2. Horizontal: parameters in the tensor (log scale). Vertical: danger weight. The tensors line up in vertical stripes because the model has only a few distinct tensor sizes (for example, all the 31,457,280-parameter tensors share one stripe). The dashed line marks danger per million parameters = 0.0344, the point where the spare budget ran out. Every tensor above it received 16-bit (filled purple), except 2 tiny in-between tensor(s) that took a partial upgrade. Many of the most dangerous tensors are large (right side), so they sit below the line and stay at their floors (rings). Hover a dot for its name.
Checked, not assumed
On this build (496 scored): the lowest danger per million parameters among the 146 tensors that got 16-bit is 0.0277; the highest among the tensors left at their floor is 0.0411. The two groups overlap, so value for money alone no longer explains the split. The median tensor that got 16-bit has only 245,760 parameters, small next to the 89,128,960 of the largest tensors. As a second test, an independent 40-line re-solve of phase 1 (same rules 1–3, same floors, same deficit) differs from the real solver plan on 0 of 497 tensors, and the minimum deficit is 2250.72.
5.5 What silence does: the unscored tensors
A tensor with no score has w = 0, so upgrading it removes no danger but still costs budget. Primary therefore never spends anything on it and it stays at its floor. 1 tensors are unscored now (5.0% of all parameters): 0 in_proj_a, 0 in_proj_b, and 1 others. Floor-plus-KL treats them very differently, because KL data exists for them.
The same 1 tensors in three plans. Plan A2 lets KL decide and keeps 0 of them at 16-bit. Plan C2 has no opinion about them, so they fall to their floors: 0 at 4-bit, 0 at 5-bit, 1 at 6-bit. Plan C3 gives them their KL ranking (section 5.6) and keeps 0 at 16-bit. In the same C2 plan, 88 of the 96 scoredin_proj_a/in_proj_b tensors get 16-bit. Whether a tiny gate tensor is kept at 16-bit therefore depends on whether it happened to be scored. Since the plans were first run, 89 more tensors have been scored, so 1 are unscored now (was 90); plain primary now pushes 0% of them to 4-bit (was 71%).
5.6 The fix: unscored tensors borrow their KL ranking
A missing score should not mean “safest”. The fix, added to the solver on 2026-09-19 as an opt-in flag, gives each unscored tensor a danger weight from the KL measurement it does have. Take the tensor’s own measured KL when squeezed to the lowest bit-width (4-bit, the hardest squeeze), sort all 497 tensors by that number from lowest to highest, and use the tensor’s position as a percentile:
wt = it ÷ (N − 1) for every tensor with no Hessian score
In words:N = 497 tensors, it is the tensor’s position when all of them are sorted by their 4-bit KL, lowest first. It is the same formula as section 3, only the source of the number changes: the most-damaged tensor gets 1.0, the least-damaged 0.0, so the direction matches the Hessian danger weight. Scored tensors and every floor are untouched.
Measured now (496 scored)
C2: plain primary
C3: primary + KL fallback
Unscored tensors kept at 16-bit
0
0
Unscored tensors at 6-bit
1
1
Scored tensors at 16-bit
146
146
Scored tensors whose bit-width changed
—
0
Tensors at 16-bit, all 497
146
146
Raw summed KL (production: 37.45)
40.27
40.27
The fix is cheap for the tensors that were already scored: only 0 of them move and 0 fewer reach 16-bit, because the unscored tensors are 5.0% of all parameters and most are tiny. It removes only 0.0 of the 2.8 points by which plain primary sits above production, so most of that penalty is not the unscored squeeze.
Honest limits of the fix
A Hessian rank and a KL rank are different signals placed on the same 0–1 scale. They disagree tensor by tensor (see the tensor explorer), so this is a sensible stand-in, not an equivalent.
It fades out by itself: every tensor the sweep scores leaves the fallback set, and at full coverage the flag changes nothing.
It has only been run as a solver plan, not built and benchmarked. An independent re-solve using the solver’s own fallback function differs from the real solver plan for C3 on 0 of 497 tensors; the default (zero) differs for plain primary (C2) on 0.
6 · Why a second phase exists
The plain idea: decide the first question, lock the answer, then settle the second question only among the answers that tie on the first. You will also see this called “lexicographic” or “dictionary” order; it is the same thing as sorting a phone book by surname and using the first name only when surnames match.
Schematic of the two phases (the bands show the idea, not measured counts).
The math
Two quantities are involved. f1 is the danger deficit (section 5). f2 is the weighted KL damage, using the measured damage of each option scaled by st, a structural weight (100 for the 16 boundary and output tensors, 1 otherwise when the soft Hessian weight is 0).
f1(x) = Σ xt,b · wt · (16 − b) f2(x) = Σ xt,b · KLt,b · st
Phase 1: f1* = the smallest possible f1(x), subject to rules 1–3
In words: find the best danger deficit any allowed allocation can reach, and call it f1*.
Phase 2: minimize f2(x), subject to rules 1–3 and f1(x) ≤ f1* + τ
In words: among only those allocations whose danger deficit is no worse than f1* (plus a tiny tolerance τ), pick the one with the lowest KL damage. The tolerance τ is the largest of 10−6, |f1*| × 10−9, and |f1*| times the solver’s achieved gap. On plan C2: f1* = 2250.72 and τ = 2.25e-06, so phase 2 may not accept anything worse than phase 1’s best.
Why not just add the two together?
The obvious alternative is one objective, f1 + λ · f2. The two numbers have different units (danger × bits versus summed KL), so λ has no natural value, and any λ lets one criterion buy back the other. That is exactly what failed with the soft weight: even at strength 1000 the KL term could not move the 20 most dangerous tensors. Lock-then-tie-break needs no λ.
What phase 2 actually did on this data
On plan C2 nothing. A fresh phase-1-only solve picks the same bit-width for all 497 tensors as the finished two-phase plan, so once danger was optimized, KL had no freedom left. The second phase is insurance for cases with real ties. The earlier note in RUNNING_GUIDE (“KL decides 3 of 497”) was measured at a different coverage.
7 · Is it better?
What has changed since the plans were first run
The same measurements, taken when the plans were first run and now, on the same solver recipes:
Measurement
When the plans were run (407 scored)
Now (496 scored)
Tensors with a fragility score
407
496
Tensors without one
90
1
Hessian floors (8-bit / 6-bit)
21 / 40
25 / 50
Tensors with no Hessian floor (scored + unscored)
436
422
Plain primary (C2): tensors at 16-bit
68
146
Plain primary (C2): unscored tensors at 4-bit
64 of 90
0 of 1
Plain primary (C2): raw summed KL
50.09
40.27
Primary + KL fallback (C3): tensors at 16-bit
149
146
Primary + KL fallback (C3): raw summed KL
39.41
40.27
Floors + KL (A2): tensors at 16-bit
142
139
Floors + KL (A2): raw summed KL
38.89
39.62
Lowest danger per million parameters among C2’s 16-bit tensors
0.0242
0.0277
The comparison
Nobody has measured that yet: no plan in this comparison has been built and benchmarked, and this page will not pretend otherwise. Here is what the real plans do show.
Raw summed KL is the very quantity a KL-first plan minimizes, so a Hessian-first plan scoring worse on it is expected; it is not a measure of real damage. Plan A2 (floors + KL) is about 2.2 above production. Plain primary (C2) is about 2.8 higher; B2 (frozen at the 407-tensor run) is 12.9 higher. The KL fallback (C3, section 5.6) narrows the gap only to 2.8.
What the evidence supports
Primary is a clear, testable statement of one hypothesis: curvature alone should decide who gets bits. It allocates by danger per parameter, as section 5 shows, with the exceptions noted there.
What it does not support yet
That this allocation protects quality better. No plan from this table has been benchmarked. Raw KL cannot settle it, because the two approaches disagree about which signal is right.
What has to happen first
Coverage, or the KL fallback while coverage is incomplete: with 1 of 497 tensors unscored, plain primary starves them (section 5.5). Then a real benchmark, naive first and the YAQA-corrected build only if promising.
Honest limits
Plans A2, C2 and C3 and every figure are computed at the last build (496 scored, 2026-09-30 01:14); a later build, or a fresh run of the commands in section 8 after the sweep moves on, gives slightly different numbers. Plan B2 is frozen at the 407-tensor run because its exact command is not known.
The allocation rule is verified only for plan C2 (floors + primary, no regional floors). Plan B2 adds regional floors on top and was not analysed the same way.
A2 and C2 differ in two ways, not one: primary on or off, and a 5-bit floor on 76 late-attention tensors (A2 has it, C2 does not). The comparison therefore does not isolate the effect of primary.
Raw KL here is summed isolated KL from the checkpoint, the same data the solver uses, not a measurement of the finished model.
8 · Which setting when
Both commands below already include the two fixes from 2026-09-18: --pareto none and --hessian-weight-strength 0. Run them from the project folder. Output paths are new names so no existing plan is overwritten. The score file keeps growing, so results will differ from this page.
Safest plan today
Floors + KL (plan A2)
Keeps KL as the judge for the 1 unscored tensors, and stays closest to production (39.62 raw KL versus 37.45). Plan A2 also carried a 5-bit floor on the 76 late-attention tensors.
Test the hypothesis
Floor + primary + KL fallback (plan C3)
The cleanest test of “does curvature predict damage?” while coverage is incomplete, because unscored tensors keep a KL-based weight instead of being starved. Plain primary (plan C2) is only meaningful at full coverage. Either one needs a benchmark first.
Do not use
Soft weight above 0
It changes nothing once a floor or primary is active, and it inflates the reported weighted KL.
--max-low-bit-run 999 switches off the separate run-guard, which is required alongside primary (see the MILP-aware guide). The first two commands differ in two ways, exactly as plans A2 and C2 did: the --hessian-primary flag and the late-attention floor (--late-attn-min-bits, 5 versus 0). The plan files do not record the run-guard setting used for A2, so it is kept disabled in both commands.
9 · Glossary
BPW (bits per weight)
The average number of bits spent per parameter across the whole model. The budget the solver may not exceed.
Hessian
A matrix describing how sharply a tensor’s output error curves as its weights change. The raw score summarizes how concentrated that curvature is.
KL (Kullback–Leibler divergence)
A measure of how far the quantized model’s output distribution moves from the original’s. Measured per tensor and per bit-width.
MILP
Mixed-integer linear program: pick whole-number (here yes/no) variables to minimize a linear total, subject to linear limits. Solved exactly by HiGHS.
Deficit
w × (16 − bits): how much danger a tensor leaves unprotected by not being at 16-bit.
Floor
A minimum bit-width forced onto a tensor before optimizing starts.
Danger per million parameters
w ÷ (parameters ÷ 1,000,000). The value-for-money ranking phase 1 effectively follows.
Tie
Two tensors with exactly the same raw score. The sort keeps them in checkpoint order, so file order settles their rank.
Pin tolerance τ
How much worse than the best phase-1 value phase 2 is allowed to accept. Here about 0.0000023.
Lexicographic
Decide on the first criterion; use the second only among ties. Same idea as dictionary order.
Unscored tensor
A tensor absent from the score file. It gets danger weight 0.
KL fallback
Opt-in flag --hessian-unscored-fallback kl: an unscored tensor gets the percentile rank of its own 4-bit KL as danger weight instead of 0.
Coverage
How many of the 497 tensors have a score. 407 when the plans were first run, 496 at the last build.
10 · How every number is produced and checked, and related pages
Every number and diagram on this page is produced by one script, hessian_primary_page.py, from the live score file and real solver runs. It calls the solver’s own compute_hessian_danger_weight and compute_hessian_floor_map, runs the real solver for plans A2, C2 and C3, cross-checks C2 and C3 with an independent re-solve of phase 1, and rebuilds this page whenever the score file, the solver or the sweep status changes. If a measured outcome changes, the sentence that describes it changes with it.
A working area, not a conclusion. Right now the allocation is driven by the Hessian score alone, and only 496 of 497 tensors have one; the godmode sweep that will add a second, richer signal is still running. This section asks one question on today’s data: what does adding a layered floor (8-bit, then 6-bit, then 5-bit) do to both signals at once, on each plan?
What is measured
Hessian danger deficit (over scored tensors): the sum of danger weight × (16 − bits). Lower means the tensors the Hessian calls fragile are better protected.
Raw summed KL: the summed measured damage of each chosen bit-width (the number a KL-first plan minimizes). Lower means less measured damage.
Read this before the chart: each plan is built to minimize one of these two numbers (primary minimizes the deficit, floors + KL minimizes weighted KL), so each wins its own metric by construction. Compare floor layouts within a plan, and read the gap between plans as the price of switching signal, not as one plan being better.
Mean bits of the critical sets: the average bit-width given to the 10% most Hessian-critical scored tensors and to the equally many most KL-critical scored tensors (the same population and the same size, so the two lists are comparable). This shows whether a layout protects one signal’s worry list at the expense of the other’s.
Every point is a real solver run at today’s coverage: six floor layouts on each of three plans, 17 runs. The weights behind the deficit do not depend on the floors, so the numbers are comparable across layouts.
Each dot is one run, numbered by floor layout (0 none, 1 top 5% only, 2 second tier at 5-bit, 3 current, 4 add a 5-bit layer to 30%, 5 add a 5-bit layer to 50%). Lower left is better on both signals. The dashed line joins the 2 of 17 runs that no other run beats on both.
Floor layout
Floored
Floor-only BPW
Hessian deficit
Raw KL
At 16-bit
Mean bits, Hessian-critical
Mean bits, KL-critical
Against the current layout
floors + KL (A2)
0 · no Hessian floors
0
4.317
2689
42.94
106
5.8
8.4
both signals worse (+292 deficit, +3.32 KL)
1 · top 5% only
25
4.484
2425
37.78
145
7.6
10.7
KL better, Hessian worse (+28 deficit, -1.84 KL)
2 · second tier at 5-bit
75
4.541
2414
38.36
143
7.8
10.7
KL better, Hessian worse (+17 deficit, -1.27 KL)
3 · current: 8 then 6
75
4.631
2397
39.62
139
8.2
9.2
baseline
4 · add a 5-bit layer to 30%
149
4.807
2462
41.01
119
7.8
8.9
both signals worse (+65 deficit, +1.39 KL)
5 · add a 5-bit layer to 50%
no solution: solver run for A2 failed: V3.3 weighted-KL solve did not meet acceptance criteria.
message=The problem is infeasible. (H
floor + primary (C2)
0 · no Hessian floors
0
4.317
2182
37.33
171
8.6
8.2
both signals better (-69 deficit, -2.94 KL)
1 · top 5% only
25
4.484
2217
38.39
160
9.9
8.4
both signals better (-34 deficit, -1.88 KL)
2 · second tier at 5-bit
75
4.541
2230
38.90
155
10.0
8.3
both signals better (-20 deficit, -1.37 KL)
3 · current: 8 then 6
75
4.631
2251
40.27
146
10.2
7.9
baseline
4 · add a 5-bit layer to 30%
149
4.807
2331
40.67
138
9.2
7.9
both signals worse (+80 deficit, +0.40 KL)
5 · add a 5-bit layer to 50%
248
5.021
2555
49.92
64
7.0
6.3
both signals worse (+304 deficit, +9.65 KL)
primary + KL fallback (C3)
0 · no Hessian floors
0
4.317
2182
37.33
171
8.6
8.2
both signals better (-69 deficit, -2.94 KL)
1 · top 5% only
25
4.484
2217
38.39
160
9.9
8.4
both signals better (-34 deficit, -1.88 KL)
2 · second tier at 5-bit
75
4.541
2230
38.90
155
10.0
8.3
both signals better (-20 deficit, -1.37 KL)
3 · current: 8 then 6
75
4.631
2251
40.27
146
10.2
7.9
baseline
4 · add a 5-bit layer to 30%
149
4.807
2331
40.67
138
9.2
7.9
both signals worse (+80 deficit, +0.40 KL)
5 · add a 5-bit layer to 50%
248
5.021
2555
49.92
64
7.0
6.3
both signals worse (+304 deficit, +9.65 KL)
Which layout does each signal prefer?
floors + KL (A2): the Hessian deficit is lowest with current: 8 then 6 (2397) and the raw KL is lowest with top 5% only (37.78). The two signals disagree on the best layout.
floor + primary (C2): the Hessian deficit is lowest with no Hessian floors (2182) and the raw KL is lowest with no Hessian floors (37.33). The two signals agree on the best layout.
primary + KL fallback (C3): the Hessian deficit is lowest with no Hessian floors (2182) and the raw KL is lowest with no Hessian floors (37.33). The two signals agree on the best layout.
What the Hessian floors themselves do
For the two primary plans the floors are extra constraints on the very number being minimized, so a lower deficit without floors is expected and proves nothing. What the floors buy is a guarantee that the Hessian-critical tensors keep at least 8 or 6 bits, which the deficit does not price (it counts every deficit point the same whatever the tensor’s size). For the floors + KL plan the story is the opposite: the floors help both numbers.
floors + KL (A2): the current floors move the Hessian deficit by -292 and the raw KL by -3.32 compared with no Hessian floors at all, and change the mean bits of the Hessian-critical tensors from 5.8 to 8.2 and of the KL-critical tensors from 8.4 to 9.2.
floor + primary (C2): the current floors move the Hessian deficit by +69 and the raw KL by +2.94 compared with no Hessian floors at all, and change the mean bits of the Hessian-critical tensors from 8.6 to 10.2 and of the KL-critical tensors from 8.2 to 7.9.
primary + KL fallback (C3): the current floors move the Hessian deficit by +69 and the raw KL by +2.94 compared with no Hessian floors at all, and change the mean bits of the Hessian-critical tensors from 8.6 to 10.2 and of the KL-critical tensors from 8.2 to 7.9.
The key experiment: adding a 5-bit layer
floors + KL (A2): adding the 5-bit layer moves the Hessian deficit by +65 (worse), raw KL by +1.39 (worse) and the number of 16-bit tensors by -20.
floor + primary (C2): adding the 5-bit layer moves the Hessian deficit by +80 (worse), raw KL by +0.40 (worse) and the number of 16-bit tensors by -8.
primary + KL fallback (C3): adding the 5-bit layer moves the Hessian deficit by +80 (worse), raw KL by +0.40 (worse) and the number of 16-bit tensors by -8.
Do the two signals agree?
Do the two signals agree on who is critical? Rank correlation between danger weight and 4-bit KL over the 496 scored tensors: -0.16. Of the 50 most Hessian-critical scored tensors, only 8 are also among the 50 most KL-critical scored tensors (about 5 would overlap by pure chance). The rest are the two disagreement corners the tensor explorer shows: fragile by curvature but harmless by KL, and the reverse. On the current floors the two objectives already pull apart: floors + KL (A2) reaches raw KL 39.62 with Hessian deficit 2397, while plain primary (C2) reaches deficit 2251 with raw KL 40.27.
What this cannot tell us
Proxies, not quality
Both measurements are proxies computed from isolated per-tensor data. Neither says which model is better. Only benchmarks of models built from these plans can. The existing suite (MMLU, GSM8K, IFEval, BFCL and HumanEval at about 100 questions each, ledger §09) may be too coarse for this: six real builds sit within 0.46 points of each other (91.68% to 92.14% on the 5-test mean), so plans that differ by a small amount will not separate at that sample size without larger samples or a finer measure such as divergence from the BF16 model on held-out text.
Unknown
What would settle it
Needs
Does Hessian-first allocation give a better model than KL-first at the same size?
Benchmark A2 against C3 (same budget, same checkpoint lineage)
Build both; larger samples or a finer measure
Do layered floors (the 5-bit layer) help or hurt real quality?
Benchmark C3 against C3 with the 5-bit layer to 30% (layout 4)
Build one more model; same measurement
Does the KL fallback for unscored tensors matter?
Benchmark C2 against C3
Only worth it while coverage is incomplete
When the two signals disagree, which one predicts real damage?
An ablation that protects only Hessian-critical/KL-benign tensors, and one that protects only KL-critical/Hessian-benign tensors
Two extra builds; the disagreement lists come from this section
Do today’s answers survive full coverage and the godmode sweep?
Rebuild this page after the sweep finishes and compare; then benchmark the shortlist
The sweep to finish
A sensible order, cheapest first: naive-affine builds of A2, C3 and C3 with the 5-bit layer; benchmark; build the YAQA-corrected version only for the plan that looks promising. Building and benchmarking is your call to trigger; the build steps are in RUNNING_GUIDE. When the godmode sweep finishes, this section is rebuilt automatically, so both solvers can be compared on full Hessian and full godmode data.
Built 2026-09-30 01:14 by hessian_primary_page.py for Hakim Ghelab, VegaLaboratories LTD. Numbers describe the score file at 496 of 497 tensors scored.