← Back to index
Mechanism reference · built 2026-09-30 01:14 · every number recomputed from live data

From a Hessian score to a bit-width: how the solver decides

The Hessian sweep hands the solver one number per tensor. This page follows that number step by step until it becomes a choice between 4, 5, 6, 8 or 16 bits, and shows what that choice does on the real plans.

How to read this page. Every term is explained where it first appears, and every calculation is shown so you can redo it by hand. Every number and diagram below is recomputed from the live Hessian score file (496 of 497 tensors scored at the last build, 2026-09-30 01:14) and from real solver runs, so it moves as the sweep progresses. The table “What has changed since the plans were first run” in section 7 shows the difference against the moment the plans were first run (407 scored), and the box below lists what changed since the previous build.
Where this sits

This is a mechanism reference, not part of the story. The story lives in the ledger: start at §08. The fixed commands are in the MILP-aware guide and RUNNING_GUIDE §11.8. This page explains why those flags behave the way they do.

Start here: what this page is about

Written so that an engineer can check every step and a non-specialist can follow the logic. On a first read you can skip the equations: the sentence under each one says the same thing in words.

A language model is mostly numbers. This one has about 25.6 billion of them (its weights), stored in 497 large tables called tensors. Storing every number with 16 bits is accurate but heavy. Storing them with 4 to 8 bits (quantization) makes the model several times smaller (and usually cheaper to run), at some cost in accuracy.

That cost is not spread evenly. Some tensors hardly notice being squeezed. Others are fragile: squeeze them and the model’s answers get worse. The model also has a fixed size limit (this project aims for an average of about 5.03 bits per number). The solver’s job is to decide, for each of the 497 tensors, whether it gets 4, 5, 6, 8 or 16 bits, so that the limit is respected and the damage is as small as possible.

Think of packing 497 items into luggage with a fixed weight allowance. Fragile items get heavy padding, sturdy ones almost none, and the allowance forces you to choose.

Fragility score from the Hessian sweep 496 of 497 tensors scored Damage measurements KL: output change when one tensor alone is squeezed Solver picks 4, 5, 6, 8 or 16 bits for each of the 497 tensors, inside a size limit Plan one bit-width per tensor (a JSON file) Build and test compress the model, measure real quality not done yet THIS PAGE
The pipeline in one picture. This page is about the solver box: how it turns the fragility score into a choice of bit-width. The last stage, building the compressed model and measuring its real quality, has not been done for any of the plans discussed here.

Fragility score

Also called the Hessian score

A measure, computed from calibration data, of how sharply the model’s error rises when a tensor’s numbers are disturbed. A low score means fragile.

KL damage measurement

Kullback–Leibler divergence

A measured number: how much the model’s output changes when only this tensor is squeezed to a given bit-width. Larger means more damage.

Solver

A mixed-integer linear program (MILP)

An exact optimizer. It is given rules and a target and finds the best whole-number choices. Here it uses the open-source HiGHS engine.

Where things stand
  • Verified The mechanism on this page. An independent re-solve reproduces a real plan exactly, tensor for tensor.
  • Not measured Whether any of these plans gives a better model. No plan compared here has been built and benchmarked yet.
  • Incomplete Fragility scores exist for 496 of 497 tensors (as of 2026-09-30; 407 when the plans were first run). Some behaviour below depends on that gap; an opt-in fix for the unscored tensors exists (section 5.6).
  • What would settle it Finish the score sweep, build the floor-plus-KL plan and the floor-plus-primary plan, and benchmark both against the production model.
Live status · built 2026-09-30 01:14
  • Coverage: 496 of 497 tensors have a fragility score (99.8%); 1 do not: 0 are the tiny in_proj_a/in_proj_b gates and 1 are others, together 5.0% of all parameters. Scored count over time: 479 → 481 → 483 → 485 → 488 → 496.
  • Sweep process: running (PID 98682, up 22:24).

Checks on this build

  • ✗ Every tensor left at its floor has lower danger per million parameters (0.0411 at most) than every tensor that reached 16-bit (0.0277 at least): 146 at 16-bit, 347 at their floor, 4 in between.
  • ✗ Plain primary squeezes most unscored tensors to 4-bit (0 of 1).
  • ✗ The KL fallback accounts for 0% of the raw-KL gap between plain primary and production (0.0 of 2.8 points).
  • ✓ An independent re-solve of phase 1 differs from the real solver on 0 tensors (C2) and 0 tensors (C3) out of 497.

Every number and diagram on this page is recomputed from the live score file and real solver runs by hessian_primary_page.py whenever the score file, the solver or the sweep changes. Plan B2 is the one exception: its exact command is not known, so it stays frozen at the 407-tensor run and is labelled so.

What changed since the previous build

Nothing measurable changed since the previous build.

The exact command, first

Most people learn a tool from one complete example. This is the recommended command, exactly as you would type it, then what each line means, then what it prints for real. Everything after this section explains why it behaves the way it does.

Copy and paste this. It only reads files and writes one new plan file; it does not build the model.
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 \
  --candidates 4,5,6,8,16 \
  --pareto none \
  --group-size 64 \
  --late-attn-min-bits 0 \
  --early-qkv-min-bits 0 \
  --lm-head-min-bits 6 \
  --boundary-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 0 \
  --hessian-floor-tiers "0.05:8,0.15:6" \
  --hessian-primary \
  --hessian-unscored-fallback kl \
  --max-low-bit-run 999 \
  --output 04_LANGUAGE_PLANS/HESSIAN_TEST/plan_floor_primary_klfallback_2026-09-19.json

What each line means

LineIn plain words
02_optimize_plan_….pyThe solver: the program that chooses a bit-width for every tensor.
…/cascaded_checkpoint_round1.jsonIts input: the measured KL damage of every tensor at every candidate bit-width (497 tensors × 5 bit-widths).
--target-bpw 5.0280…The size budget: the average number of bits per weight to aim for. This is the value of the production build, so plans compare like for like.
--candidates 4,5,6,8,16The bit-widths a tensor may be given.
--pareto noneKeep every candidate bit-width. Required whenever a Hessian score file is used; the default silently removes bit-widths a floor needs and can make the problem unsolvable.
--group-size 64Recorded in the plan for the later build step: how many numbers share one scale when a tensor is stored. It does not change which bit-width the solver picks.
--late-attn-min-bits 0
--early-qkv-min-bits 0
Two regional floors (a minimum number of bits for attention tensors in the last quarter of the model, and for the query/key/value tensors in the first quarter). 0 switches each one off, as in plan C2.
--lm-head-min-bits 6The output layer, lm_head, is never given fewer than 6 bits.
--boundary-min-bits 6The tensors in the first and last layer (0 and 63) are never given fewer than 6 bits.
--hessian-score-file ….jsonThe fragility score of every tensor the Hessian sweep has scored so far.
--hessian-weight-strength 0Switches off the older soft weight. It does nothing once a floor or primary is active (section 1).
--hessian-floor-tiers "0.05:8,0.15:6"The floor: the most fragile 5% of scored tensors get at least 8 bits, the next 10% at least 6 (section 4).
--hessian-primaryDecide the spare bits by fragility first and KL second (sections 5 and 6).
--hessian-unscored-fallback klTensors with no fragility score borrow their own KL ranking instead of counting as “safest” (section 5.6).
--max-low-bit-run 999Switches off a separate safety rule about long runs of low-bit blocks. Required alongside --hessian-primary.
--output ….jsonWhere the plan is written. A new file name, so no existing plan is overwritten.

What it prints, for real

This is the actual output of that command, from 2026-09-30 01:15 at 496 scored tensors (long lines shortened with …, the output path replaced by a placeholder). The two solves took about a second and a half.

[lm-head-floor] 1 lm_head tensor(s) floored at >= 6-bit (real fix, 2026-09-12 -- see IMPROVEMENT_LEDGER/04_LMHEAD_PROTECTION_INCIDENT.html)
[boundary-floor] 15 boundary tensor(s) (layer 0 + layer 63) floored at >= 6-bit (real fix, 2026-09-14 -- the flat 100x weight alone was not enough, see the flag's own help text)
[hessian-floor] real MILP-aware Hessian floor (2026-09-15): 0.05:8,0.15:6 -- 75 real tensor(s) floored, by tier bits: {6: 50, 8: 25}
[hessian-unscored-fallback] kl: 1 of 1 unscored tensor(s) given the percentile rank of their own isolated KL at 4-bit as danger weight (instead of 0.0 = safest); scored tensors and floors ...
[hessian-primary] real lexicographic solve ENABLED -- Phase 1 optimizes danger-weighted bit allocation for 496 real Hessian-scored tensor(s) plus 1 KL-fallback tensor(s) with zero KL infl ...
[P0_hessian_primary] optimal_or_gap_target_reached | 0.78s | gap=0.0 | nodes=893
[Q_global_weighted_KL] optimal_or_gap_target_reached | 7.68s | gap=1.512286573675527e-05 | nodes=684

Tensors:                    497
Requested BPW:              5.028075123
Hard quality ceiling BPW:   5.028575123
Selected / achieved BPW:    5.028573895
Raw Σ isolated KL:          40.272079778

Q4 :  235 tensors |  69.17% param mass
Q5 :   59 tensors |   8.55% param mass
Q6 :   35 tensors |  13.13% param mass
Q8 :   22 tensors |   5.22% param mass
Q16:  146 tensors |   3.93% param mass

Wrote: <the --output path you gave>
How to read the output
  • [hessian-floor] … 75 tensors floored: 25 tensors got an 8-bit floor and 50 a 6-bit floor.
  • [hessian-unscored-fallback] kl: 1 of 1: every tensor still missing a score borrowed its KL ranking.
  • [P0_hessian_primary] … gap=0.0: phase 1, the fragility solve, finished with the exact optimum. [Q_global_weighted_KL] is phase 2, the KL tie-break.
  • Selected / achieved BPW 5.028573895 must be at or under the ceiling 5.028575123: the size budget was respected.
  • Raw Σ isolated KL 40.27 is the damage estimate compared in section 7 (production: 37.45). Q16: 146 tensors is how many tensors were kept at full 16-bit precision.

The complete commands for the other two variants, floors plus KL and plain primary without the fallback, are in section 8. Nothing here builds or benchmarks a model; that is a separate step.

1 · The short answer

The solver receives one fragility (Hessian) score per tensor and can use it in three separate ways, selected by three command-line flags. They differ in when they act and in what the solver may still trade away.

--hessian-floor-tiers

A floor: a rule

Acts first, before any optimizing. It says “this tensor may never go below N bits.” The solver cannot trade it away for anything.

It only sets a minimum. It does not say how far above the minimum a tensor should go. In the current plans it covers 75 of 496 scored tensors (25 at 8-bit, 50 at 6-bit).

--hessian-weight-strength

A weight: a preference

Makes a tensor’s KL cost look bigger, so the solver is more reluctant to damage it. It is a preference, so it can still be traded away.

Measured: changes nothing once a floor or primary is active, and it inflates the reported weighted KL. Keep it at 0.

--hessian-primary

A different question

Changes what the solver optimizes first. Instead of “which allocation has the smallest KL?”, it asks “which allocation leaves the least Hessian danger unprotected?” KL only breaks ties.

It acts on the spare bits above the floors: who gets 16, who stays low.

How does primary choose between 4, 5, 6, 8 and 16?

Not tensor by tensor. Every tensor, judged alone, would take 16-bit, so its own candidates never make the choice. The choice comes from a shared budget:

  • Every tensor starts at its floor. That uses 4.6308 bits per weight; the ceiling is 5.0286, leaving 0.3978 to spend.
  • Upgrading a tensor to 16-bit removes danger in proportion to its danger weight, and costs budget in proportion to its size. The solver spends the spare budget on the tensors with the most danger per million parameters.
  • In plan C2 the split between 16-bit tensors and floor tensors no longer follows danger per million parameters cleanly; section 5.4 shows the check.
Two surprises worth knowing before you read on
  • The single most dangerous tensor in the model (layers.63.mlp.up_proj) does not get 16-bit under primary. It stays at 8-bit. A nearly-safe tiny tensor (layers.1.linear_attn.in_proj_b, 245,760 parameters) does get 16-bit.
  • Tensors with no score get danger weight 0, and on that scale 0 means safest, so primary pushes them down to their floors. 0 of the 1 unscored tensors end at 4-bit, where floor-plus-KL keeps 0 of them at 16-bit. This is fixable with an opt-in flag (section 5.6).

2 · Three numbers that all get called “danger”

“Danger” here just means fragility: a dangerous tensor is one that is easily damaged by squeezing. The code and the older docs use that one word for three different numbers, and they do not point the same way. That is the main source of confusion, so they get separate names here.

Raw score

Low means dangerous

The number from the Hessian sweep (hess_score): how spread out a tensor’s curvature is. A low score means the curvature is concentrated in a few directions, which is fragile.

Real range: 0.000161 to 0.1780, a 1103-fold span. How it is computed: Hessian Score — Origin.

Rank r

Low means dangerous

Sort all scored tensors from lowest raw score to highest. A tensor’s position i is its place in that list, counting from 0. Then r = i ÷ (n − 1), a number from 0 to 1.

r = 0 is the most dangerous tensor, r = 1 the safest. The floor uses r.

Danger weight w

High means dangerous

Simply w = 1 − r. The direction flips: w = 1 is the most dangerous tensor, w = 0 the safest. Primary uses w.

The solver calls this hessian_danger_weight.

1 2 3 4 Raw score low = dangerous log scale Rank r 0 = most dangerous r = i ÷ 495 Danger weight w 1 = most dangerous w = 1 − r 0.1780 0.000161 1.0 0.0 1.0 0.0 position i in the sorted list (0 = lowest raw score → 495 = highest)
The same 496 tensors, sorted from lowest to highest raw score, drawn three ways. The top curve is the raw score on a log scale. It climbs slowly and then steeply, because the values span 312-fold. The middle line is the rank, and the bottom line is the danger weight: both are straight lines, which is why the solver uses them instead of the raw score. Four real tensors are marked: 1layers.63.mlp.up_proj (the most dangerous), 2layers.1.linear_attn.in_proj_qkv, 3layers.51.mlp.up_proj (in the middle), 4layers.1.linear_attn.in_proj_b (almost the safest).

3 · From score to rank, line by line

These are the real lines from the solver (compute_hessian_danger_weight, lines 339–345 of 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py). The floor function repeats the same first three lines.

scored_names = [n for n in names if n in raw_scores]        # tensors that HAVE a score
ranked = sorted(scored_names, key=lambda n: raw_scores[n])  # lowest raw score first
n = len(ranked)                                             # how many tensors are scored
rank_of = {name: i / max(n - 1, 1) for i, name in enumerate(ranked)}   # r = i / (n - 1)
return {name: 1.0 - rank_of[name] for name in scored_names}            # w = 1 - r

A worked example you can check with a calculator

Take tensor 2 layers.1.linear_attn.in_proj_qkv. In the sorted list of 496 scored tensors it sits at position i = 10 (10 tensors have a lower raw score). Here n = 496, so n − 1 = 495.

r = i ÷ (n − 1) = 10 ÷ 495 = 0.0202
w = 1 − r = 1 − 0.0202 = 0.9798
TensorPosition iRank r = i ÷ 495Danger weight w = 1 − r
1layers.63.mlp.up_proj00.00001.0000
2layers.1.linear_attn.in_proj_qkv100.02020.9798
3layers.51.mlp.up_proj2120.42830.5717
4layers.1.linear_attn.in_proj_b4270.86260.1374

Why rank instead of the raw score?

The raw scores span 1103-fold, from 0.000161 to 0.1780. Used directly, a single extreme value would dominate everything else. Rank squeezes every tensor into the same 0–1 range no matter how the raw values are spread. The cost is that rank only records order: the gap between two neighbouring tensors is lost.

Three properties of this calculation that are easy to miss

Ties are broken by file order, not by data

34 groups of tensors (75 tensors in all) share an identical raw score. Python’s sorted keeps tied tensors in the order they appear in the checkpoint file, so among equals the position, and therefore the rank, comes from list order.

Ranks move when coverage grows

n is the number of scored tensors, not 497. At 407 scored (when the plans were first run), tensor 2 had r = 0.0246 and w = 0.9754. Now, with 496 scored, the same tensor has r = 0.0202 and w = 0.9798. The tensor did not change; the list around it did. Percentages such as “the bottom 5%” therefore pick different tensors as the sweep progresses.

No score means weight 0, and 0 is the SAFEST end

A tensor missing from the score file gets danger weight w = 0 (the code comment: silence means no signal). On the danger-weight scale 0 is the safest end, and 1 is the most dangerous. So a missing score does not make a tensor look risky. It makes it look perfectly safe, and primary under-protects it. Do not confuse this with the raw score, where a low value means dangerous. Section 5.5 shows what it does to the 1 tensors that are unscored now, and section 5.6 the fix. Right now 2 tensors carry exactly 0: the 1 unscored ones, plus the one scored tensor with the highest raw score (its rank is exactly 1.0, so its weight is 1 − 1 = 0).

4 · The floor: a rule, not a preference

--hessian-floor-tiers "0.05:8,0.15:6" reads as: rank under 0.05 gets at least 8-bit; rank under 0.15 gets at least 6-bit.

The solver checks the tiers in order and the first one the rank falls under wins. With n = 496 the two thresholds land at positions 0.05 × 495 = 24.8 and 0.15 × 495 = 74.2. Positions 0 to 24 are below the first threshold, so 25 tensors get the 8-bit floor; positions 25 to 74 get the 6-bit floor, so 50 tensors; the other 421 scored tensors get no Hessian floor, and neither do the 1 unscored ones (they have no rank), so 422 of the 497 tensors have no Hessian floor now. That number falls as more tensors are scored, but slowly: the tiers are percentages, so at full coverage only 75 tensors get a floor and the other 422 never do. “No Hessian floor” does not mean 0 bits: those tensors may still go as low as 4-bit and are still held by the other floors below.

4-bit 6-bit 8-bit lowest bit-width the solver may pick rank < 0.05 → at least 8-bit · 25 tensors rank < 0.15 → at least 6-bit · 50 tensors no Hessian floor · 421 scored tensors may go as low as 4-bit; the other floors below still apply. The 1 unscored tensors have no rank, so no Hessian floor either. 0 0.05 0.15 0.5 1.0 position 24.8 position 74.2 rank r (0 = most dangerous)
A floor is a step function of rank. It is added to the solver as a hard constraint (add_min_bits_constraint) before any optimizing. Whatever the objective wants, the tensor cannot go below its floor. The floor never says “exactly”; a floored tensor can still be raised.

Ties and the floor cutoffs

Right now no group of tensors with exactly the same raw score straddles a floor cutoff, so the ordering of ties does not change who gets a floor. It did when the plans were first run, and it can again as coverage grows.

The floor is not the only floor

Each tensor’s effective minimum is the highest of every floor that applies to it. In the tested plans (C2) those are:

FloorApplies toMinimum
Hessian tiers (this section)75 tensors, by rank8 or 6
Boundary protection15 tensors in layers 0 and 636
lm_head protection1 tensor6
Vendor per-block rule (the OptiQ structural policy)128 tensors, chosen from each block’s own KL data5
Everything elsethe lowest allowed candidate4

5 · Primary, phase 1: how it picks a bit-width

This is the mechanism behind --hessian-primary, in full. A MILP (mixed-integer linear program) is a problem where you choose whole-number decisions to make one number as small as possible while obeying linear limits. The solver, HiGHS, finds the exact best answer.

5.1  What the solver chooses, and the three rules

For every tensor t and every candidate bit-width b in {4, 5, 6, 8, 16} there is one yes/no variable xt,b: it is 1 if tensor t is given b bits, and 0 otherwise. That is 497 × 5 = 2,485 variables, and the solver sets each to 0 or 1. The symbol Σ just means “add up”.

Σb xt,b = 1   for every tensor t Rule 1, in words: each tensor gets exactly one bit-width. The five yes/no variables of a tensor must add up to 1.
xt,b = 0   whenever b < floort Rule 2, in words: a tensor cannot get fewer bits than its floor (section 4).
Σt Σb xt,b · b · Pt ÷ Ptotal ≤ ceiling Rule 3, in words (the budget): multiply each tensor’s bits by its size, add everything up and divide by the model’s total size. That average number of bits per weight (BPW) must stay under the ceiling. Pt is the number of parameters in tensor t; Ptotal = 25,621,954,560. The ceiling is the target 5.0281 plus a 0.0005 slack, 5.0286.

5.2  What phase 1 tries to minimize

Each option (tensor t, bit-width b) gets a deficit: how much danger is left unprotected by choosing b instead of the maximum, 16 bits. A fragile tensor (large w) left at few bits has a large deficit.

deficit(t, b) = wt × (16 − b) The real code line: hessian_danger_weight.get(name, 0.0) * (max_bits - bits)
Phase 1 minimizes  Σt Σb xt,b · deficit(t, b) In words: choose the bit-widths so that the total unprotected danger across all tensors is as small as possible. KL plays no part in this step.

For tensor 2 (w = 0.9798) the five options are:

Bit-width16 − bits× 0.9798 = deficit
4-bit1211.76
5-bit1110.78
6-bit109.80
8-bit87.84
16-bit00.00

Deficit falls with every extra bit and never flattens. So looking at one tensor alone, 16-bit is always the cheapest option, for every tensor. That is the important point: the tensor’s own options never decide. Only the shared budget in rule 3 says “not everyone can have 16”.

5.3  Where the budget goes: value for money

Suppose a tensor is raised to 16-bit. Two things happen, and they are measured in different currencies:

The four marked tensors, each raised to 16-bit from its floor:

WHAT IT BUYSdanger removed (w × bits added) WHAT IT COSTSbudget used (bits per weight) VALUEdanger per millionparameters 1 layers.63.mlp.up_proj 89,128,960 parameters · w = 1.0000 raise from 8-bit to 16-bit 8.00 0.0278 0.0112 stays at its floor 2 layers.1.linear_attn.in_proj_qkv 52,428,800 parameters · w = 0.9798 raise from 8-bit to 16-bit 7.84 0.0164 0.0187 stays at its floor 3 layers.51.mlp.up_proj 89,128,960 parameters · w = 0.5717 raise from 4-bit to 16-bit 6.86 0.0417 0.0064 stays at its floor 4 layers.1.linear_attn.in_proj_b 245,760 parameters · w = 0.1374 raise from 4-bit to 16-bit 1.65 0.0001 0.5590 gets 16-bit
Each bar is scaled to the largest value in its own column. Tensor 4 buys very little (1.65) but costs almost nothing (0.0001 bits per weight). Tensor 1 buys a lot (8.00) but costs 242 times as much. Value = what it buys ÷ what it costs, and tensor 4 wins by a factor of 49.8.
value = (wt × bits added) ÷ (bits added × Pt ÷ Ptotal) = wt × Ptotal ÷ Pt In words: the number of bits added cancels, and Ptotal is the same for every tensor. So tensors are ranked by wt ÷ Pt: danger divided by size. Written per million parameters: danger per million parameters = wt ÷ (Pt ÷ 1,000,000). This is the number in the last column of the figure and in the scatter below.

The deficit has no size in it, but the budget does. So the solver does not favour the most dangerous tensors; it favours the ones where protection is cheapest per unit of danger. What actually happened to these four:

TensorParametersDanger weight wDanger per million parametersFloorC2 (floor + primary)A2 (floor + KL)Production
1layers.63.mlp.up_proj89,128,9601.00000.0112881616
2layers.1.linear_attn.in_proj_qkv52,428,8000.97980.01878888
3layers.51.mlp.up_proj89,128,9600.57170.00644444
4layers.1.linear_attn.in_proj_b245,7600.13740.5590416166

Tensor 4 is almost the safest in the model, but it is tiny, so under primary it gets 16-bit while tensor 1, the most dangerous, stays at 8-bit.

5.4  The budget as a shopping trip

Start every tensor at its floor. That plan uses 4.6308 bits per weight. The ceiling is 5.0286, so 0.3978 bits per weight are left to spend. Phase 1 spends them on upgrades to 16-bit, best value first, until the money runs out.

Every tensor at its floor uses 4.6308 bits per weight Spare: 0.3978 bits per weight to spend ceiling 5.0286 4.0 4.5 5.0 spent on 146 upgrades to 16-bit (best value first) and 4 partial upgrades average bits per weight across the whole model (axis starts at 4.0)
Reading the bar: the gray part is what the whole model costs if every tensor stays at its floor. The purple part is the spare allowance left under the ceiling, and it is all that can be spent on upgrades.

146 tensors fit; the small remainder buys 4 partial upgrades. So 493 of the 497 tensors are either at 16-bit or exactly on their floor. This is why the result looks all-or-nothing: a shopping list ranked by value for money, bought from the top until the money is gone. The scatter below shows every scored tensor and where the cut falls.

0 0.25 0.5 0.75 1 0.3M 1M 3M 10M 30M 100M 16-bit territory: danger per million parameters ≥ 0.0344 layers.0.linear_attn.in_proj_a · 16-bit layers.0.linear_attn.in_proj_b · 16-bit layers.0.linear_attn.in_proj_qkv · 6-bit layers.0.linear_attn.in_proj_z · 6-bit layers.0.linear_attn.out_proj · 16-bit layers.0.mlp.down_proj · 6-bit layers.0.mlp.gate_proj · 6-bit layers.0.mlp.up_proj · 6-bit layers.1.linear_attn.in_proj_a · 16-bit layers.1.linear_attn.in_proj_b · 16-bit layers.1.linear_attn.in_proj_qkv · 8-bit layers.1.linear_attn.in_proj_z · 16-bit layers.1.linear_attn.out_proj · 5-bit layers.1.mlp.down_proj · 4-bit layers.1.mlp.gate_proj · 4-bit layers.1.mlp.up_proj · 4-bit layers.2.linear_attn.in_proj_a · 16-bit layers.2.linear_attn.in_proj_b · 16-bit layers.2.linear_attn.in_proj_qkv · 8-bit layers.2.linear_attn.in_proj_z · 16-bit layers.2.linear_attn.out_proj · 5-bit layers.2.mlp.down_proj · 4-bit layers.2.mlp.gate_proj · 4-bit layers.2.mlp.up_proj · 4-bit layers.3.self_attn.k_proj · 16-bit layers.3.self_attn.o_proj · 4-bit layers.3.self_attn.q_proj · 8-bit layers.3.self_attn.v_proj · 16-bit layers.3.mlp.down_proj · 4-bit layers.3.mlp.gate_proj · 4-bit layers.3.mlp.up_proj · 4-bit layers.4.linear_attn.in_proj_a · 16-bit layers.4.linear_attn.in_proj_b · 16-bit layers.4.linear_attn.in_proj_qkv · 6-bit layers.4.linear_attn.in_proj_z · 4-bit layers.4.linear_attn.out_proj · 5-bit layers.4.mlp.down_proj · 4-bit layers.4.mlp.gate_proj · 4-bit layers.4.mlp.up_proj · 4-bit layers.5.linear_attn.in_proj_a · 16-bit layers.5.linear_attn.in_proj_b · 16-bit layers.5.linear_attn.in_proj_qkv · 6-bit layers.5.linear_attn.in_proj_z · 16-bit layers.5.linear_attn.out_proj · 4-bit layers.5.mlp.down_proj · 4-bit layers.5.mlp.gate_proj · 4-bit layers.5.mlp.up_proj · 4-bit layers.6.linear_attn.in_proj_a · 16-bit layers.6.linear_attn.in_proj_b · 16-bit layers.6.linear_attn.in_proj_qkv · 8-bit layers.6.linear_attn.in_proj_z · 16-bit layers.6.linear_attn.out_proj · 4-bit layers.6.mlp.down_proj · 4-bit layers.6.mlp.gate_proj · 4-bit layers.6.mlp.up_proj · 4-bit layers.7.self_attn.k_proj · 16-bit layers.7.self_attn.o_proj · 4-bit layers.7.self_attn.q_proj · 8-bit layers.7.self_attn.v_proj · 16-bit layers.7.mlp.down_proj · 4-bit layers.7.mlp.gate_proj · 4-bit layers.7.mlp.up_proj · 4-bit layers.8.linear_attn.in_proj_a · 16-bit layers.8.linear_attn.in_proj_b · 16-bit layers.8.linear_attn.in_proj_qkv · 6-bit layers.8.linear_attn.in_proj_z · 16-bit layers.8.linear_attn.out_proj · 5-bit layers.8.mlp.down_proj · 4-bit layers.8.mlp.gate_proj · 4-bit layers.8.mlp.up_proj · 4-bit layers.9.linear_attn.in_proj_a · 16-bit layers.9.linear_attn.in_proj_b · 16-bit layers.9.linear_attn.in_proj_qkv · 8-bit layers.9.linear_attn.in_proj_z · 16-bit layers.9.linear_attn.out_proj · 5-bit layers.9.mlp.down_proj · 4-bit layers.9.mlp.gate_proj · 4-bit layers.9.mlp.up_proj · 4-bit layers.10.linear_attn.in_proj_a · 16-bit layers.10.linear_attn.in_proj_b · 16-bit layers.10.linear_attn.in_proj_qkv · 8-bit layers.10.linear_attn.in_proj_z · 16-bit layers.10.linear_attn.out_proj · 5-bit layers.10.mlp.down_proj · 4-bit layers.10.mlp.gate_proj · 4-bit layers.10.mlp.up_proj · 4-bit layers.11.self_attn.k_proj · 16-bit layers.11.self_attn.o_proj · 5-bit layers.11.self_attn.q_proj · 8-bit layers.11.self_attn.v_proj · 16-bit layers.11.mlp.down_proj · 4-bit layers.11.mlp.gate_proj · 4-bit layers.11.mlp.up_proj · 4-bit layers.12.linear_attn.in_proj_a · 16-bit layers.12.linear_attn.in_proj_b · 16-bit layers.12.linear_attn.in_proj_qkv · 8-bit layers.12.linear_attn.in_proj_z · 16-bit layers.12.linear_attn.out_proj · 4-bit layers.12.mlp.down_proj · 4-bit layers.12.mlp.gate_proj · 4-bit layers.12.mlp.up_proj · 4-bit layers.13.linear_attn.in_proj_a · 16-bit layers.13.linear_attn.in_proj_b · 16-bit layers.13.linear_attn.in_proj_qkv · 8-bit layers.13.linear_attn.in_proj_z · 16-bit layers.13.linear_attn.out_proj · 5-bit layers.13.mlp.down_proj · 4-bit layers.13.mlp.gate_proj · 4-bit layers.13.mlp.up_proj · 4-bit layers.14.linear_attn.in_proj_a · 16-bit layers.14.linear_attn.in_proj_b · 16-bit layers.14.linear_attn.in_proj_qkv · 8-bit layers.14.linear_attn.in_proj_z · 16-bit layers.14.linear_attn.out_proj · 4-bit layers.14.mlp.down_proj · 4-bit layers.14.mlp.gate_proj · 4-bit layers.14.mlp.up_proj · 4-bit layers.15.self_attn.k_proj · 16-bit layers.15.self_attn.o_proj · 4-bit layers.15.self_attn.q_proj · 8-bit layers.15.self_attn.v_proj · 16-bit layers.15.mlp.down_proj · 4-bit layers.15.mlp.gate_proj · 4-bit layers.15.mlp.up_proj · 4-bit layers.16.linear_attn.in_proj_a · 16-bit layers.16.linear_attn.in_proj_b · 16-bit layers.16.linear_attn.in_proj_qkv · 8-bit layers.16.linear_attn.in_proj_z · 16-bit layers.16.linear_attn.out_proj · 5-bit layers.16.mlp.down_proj · 4-bit layers.16.mlp.gate_proj · 4-bit layers.16.mlp.up_proj · 4-bit layers.17.linear_attn.in_proj_a · 16-bit layers.17.linear_attn.in_proj_b · 16-bit layers.17.linear_attn.in_proj_qkv · 6-bit layers.17.linear_attn.in_proj_z · 16-bit layers.17.linear_attn.out_proj · 4-bit layers.17.mlp.down_proj · 4-bit layers.17.mlp.gate_proj · 4-bit layers.17.mlp.up_proj · 4-bit layers.18.linear_attn.in_proj_a · 16-bit layers.18.linear_attn.in_proj_b · 16-bit layers.18.linear_attn.in_proj_qkv · 6-bit layers.18.linear_attn.in_proj_z · 16-bit layers.18.linear_attn.out_proj · 6-bit layers.18.mlp.down_proj · 4-bit layers.18.mlp.gate_proj · 4-bit layers.18.mlp.up_proj · 4-bit layers.19.self_attn.k_proj · 16-bit layers.19.self_attn.o_proj · 5-bit layers.19.self_attn.q_proj · 8-bit layers.19.self_attn.v_proj · 16-bit layers.19.mlp.down_proj · 4-bit layers.19.mlp.gate_proj · 4-bit layers.19.mlp.up_proj · 4-bit layers.20.linear_attn.in_proj_a · 16-bit layers.20.linear_attn.in_proj_b · 16-bit layers.20.linear_attn.in_proj_qkv · 6-bit layers.20.linear_attn.in_proj_z · 16-bit layers.20.linear_attn.out_proj · 8-bit layers.20.mlp.down_proj · 6-bit layers.20.mlp.gate_proj · 4-bit layers.20.mlp.up_proj · 6-bit layers.21.linear_attn.in_proj_a · 16-bit layers.21.linear_attn.in_proj_b · 16-bit layers.21.linear_attn.in_proj_qkv · 6-bit layers.21.linear_attn.in_proj_z · 16-bit layers.21.linear_attn.out_proj · 4-bit layers.21.mlp.down_proj · 4-bit layers.21.mlp.gate_proj · 4-bit layers.21.mlp.up_proj · 6-bit layers.22.linear_attn.in_proj_a · 16-bit layers.22.linear_attn.in_proj_b · 16-bit layers.22.linear_attn.in_proj_qkv · 4-bit layers.22.linear_attn.in_proj_z · 4-bit layers.22.linear_attn.out_proj · 5-bit layers.22.mlp.down_proj · 6-bit layers.22.mlp.gate_proj · 6-bit layers.22.mlp.up_proj · 4-bit layers.23.self_attn.k_proj · 16-bit layers.23.self_attn.o_proj · 4-bit layers.23.self_attn.q_proj · 8-bit layers.23.self_attn.v_proj · 16-bit layers.23.mlp.down_proj · 4-bit layers.23.mlp.gate_proj · 4-bit layers.23.mlp.up_proj · 6-bit layers.24.linear_attn.in_proj_a · 16-bit layers.24.linear_attn.in_proj_b · 16-bit layers.24.linear_attn.in_proj_qkv · 6-bit layers.24.linear_attn.in_proj_z · 16-bit layers.24.linear_attn.out_proj · 5-bit layers.24.mlp.down_proj · 4-bit layers.24.mlp.gate_proj · 4-bit layers.24.mlp.up_proj · 4-bit layers.25.linear_attn.in_proj_a · 16-bit layers.25.linear_attn.in_proj_b · 16-bit layers.25.linear_attn.in_proj_qkv · 6-bit layers.25.linear_attn.in_proj_z · 16-bit layers.25.linear_attn.out_proj · 5-bit layers.25.mlp.down_proj · 4-bit layers.25.mlp.gate_proj · 4-bit layers.25.mlp.up_proj · 4-bit layers.26.linear_attn.in_proj_a · 16-bit layers.26.linear_attn.in_proj_b · 16-bit layers.26.linear_attn.in_proj_qkv · 6-bit layers.26.linear_attn.in_proj_z · 16-bit layers.26.linear_attn.out_proj · 4-bit layers.26.mlp.down_proj · 4-bit layers.26.mlp.gate_proj · 4-bit layers.26.mlp.up_proj · 4-bit layers.27.self_attn.k_proj · 16-bit layers.27.self_attn.o_proj · 4-bit layers.27.self_attn.q_proj · 8-bit layers.27.self_attn.v_proj · 16-bit layers.27.mlp.down_proj · 4-bit layers.27.mlp.gate_proj · 4-bit layers.27.mlp.up_proj · 4-bit layers.28.linear_attn.in_proj_a · 16-bit layers.28.linear_attn.in_proj_b · 16-bit layers.28.linear_attn.in_proj_qkv · 6-bit layers.28.linear_attn.in_proj_z · 4-bit layers.28.linear_attn.out_proj · 5-bit layers.28.mlp.down_proj · 4-bit layers.28.mlp.gate_proj · 4-bit layers.28.mlp.up_proj · 4-bit layers.29.linear_attn.in_proj_a · 16-bit layers.29.linear_attn.in_proj_b · 16-bit layers.29.linear_attn.in_proj_qkv · 6-bit layers.29.linear_attn.in_proj_z · 4-bit layers.29.linear_attn.out_proj · 5-bit layers.29.mlp.down_proj · 4-bit layers.29.mlp.gate_proj · 4-bit layers.29.mlp.up_proj · 4-bit layers.30.linear_attn.in_proj_a · 16-bit layers.30.linear_attn.in_proj_b · 16-bit layers.30.linear_attn.in_proj_qkv · 5-bit layers.30.linear_attn.in_proj_z · 4-bit layers.30.linear_attn.out_proj · 4-bit layers.30.mlp.down_proj · 4-bit layers.30.mlp.gate_proj · 4-bit layers.30.mlp.up_proj · 4-bit layers.31.self_attn.k_proj · 16-bit layers.31.self_attn.o_proj · 4-bit layers.31.self_attn.q_proj · 8-bit layers.31.self_attn.v_proj · 16-bit layers.31.mlp.down_proj · 4-bit layers.31.mlp.gate_proj · 4-bit layers.31.mlp.up_proj · 4-bit layers.32.linear_attn.in_proj_a · 16-bit layers.32.linear_attn.in_proj_b · 16-bit layers.32.linear_attn.in_proj_qkv · 6-bit layers.32.linear_attn.in_proj_z · 5-bit layers.32.linear_attn.out_proj · 4-bit layers.32.mlp.down_proj · 4-bit layers.32.mlp.gate_proj · 4-bit layers.32.mlp.up_proj · 4-bit layers.33.linear_attn.in_proj_a · 16-bit layers.33.linear_attn.in_proj_b · 16-bit layers.33.linear_attn.in_proj_qkv · 4-bit layers.33.linear_attn.in_proj_z · 5-bit layers.33.linear_attn.out_proj · 5-bit layers.33.mlp.down_proj · 4-bit layers.33.mlp.gate_proj · 4-bit layers.33.mlp.up_proj · 4-bit layers.34.linear_attn.in_proj_a · 16-bit layers.34.linear_attn.in_proj_b · 16-bit layers.34.linear_attn.in_proj_qkv · 5-bit layers.34.linear_attn.in_proj_z · 16-bit layers.34.linear_attn.out_proj · 5-bit layers.34.mlp.down_proj · 4-bit layers.34.mlp.gate_proj · 4-bit layers.34.mlp.up_proj · 4-bit layers.35.self_attn.k_proj · 16-bit layers.35.self_attn.o_proj · 4-bit layers.35.self_attn.q_proj · 6-bit layers.35.self_attn.v_proj · 16-bit layers.35.mlp.down_proj · 4-bit layers.35.mlp.gate_proj · 4-bit layers.35.mlp.up_proj · 4-bit layers.36.linear_attn.in_proj_a · 16-bit layers.36.linear_attn.in_proj_b · 4-bit layers.36.linear_attn.in_proj_qkv · 4-bit layers.36.linear_attn.in_proj_z · 5-bit layers.36.linear_attn.out_proj · 5-bit layers.36.mlp.down_proj · 4-bit layers.36.mlp.gate_proj · 4-bit layers.36.mlp.up_proj · 4-bit layers.37.linear_attn.in_proj_a · 16-bit layers.37.linear_attn.in_proj_b · 16-bit layers.37.linear_attn.in_proj_qkv · 5-bit layers.37.linear_attn.in_proj_z · 4-bit layers.37.linear_attn.out_proj · 5-bit layers.37.mlp.down_proj · 4-bit layers.37.mlp.gate_proj · 4-bit layers.37.mlp.up_proj · 4-bit layers.38.linear_attn.in_proj_a · 16-bit layers.38.linear_attn.in_proj_b · 4-bit layers.38.linear_attn.in_proj_qkv · 5-bit layers.38.linear_attn.in_proj_z · 4-bit layers.38.linear_attn.out_proj · 5-bit layers.38.mlp.down_proj · 4-bit layers.38.mlp.gate_proj · 4-bit layers.38.mlp.up_proj · 4-bit layers.39.self_attn.k_proj · 16-bit layers.39.self_attn.o_proj · 4-bit layers.39.self_attn.q_proj · 5-bit layers.39.self_attn.v_proj · 16-bit layers.39.mlp.down_proj · 4-bit layers.39.mlp.gate_proj · 4-bit layers.39.mlp.up_proj · 4-bit layers.40.linear_attn.in_proj_a · 16-bit layers.40.linear_attn.in_proj_b · 16-bit layers.40.linear_attn.in_proj_qkv · 5-bit layers.40.linear_attn.in_proj_z · 4-bit layers.40.linear_attn.out_proj · 5-bit layers.40.mlp.down_proj · 4-bit layers.40.mlp.gate_proj · 4-bit layers.40.mlp.up_proj · 4-bit layers.41.linear_attn.in_proj_a · 16-bit layers.41.linear_attn.in_proj_b · 4-bit layers.41.linear_attn.in_proj_qkv · 5-bit layers.41.linear_attn.in_proj_z · 4-bit layers.41.linear_attn.out_proj · 5-bit layers.41.mlp.down_proj · 4-bit layers.41.mlp.gate_proj · 4-bit layers.41.mlp.up_proj · 4-bit layers.42.linear_attn.in_proj_a · 16-bit layers.42.linear_attn.in_proj_b · 6-bit layers.42.linear_attn.in_proj_qkv · 5-bit layers.42.linear_attn.in_proj_z · 4-bit layers.42.linear_attn.out_proj · 5-bit layers.42.mlp.down_proj · 4-bit layers.42.mlp.gate_proj · 4-bit layers.42.mlp.up_proj · 4-bit layers.43.self_attn.k_proj · 16-bit layers.43.self_attn.o_proj · 4-bit layers.43.self_attn.q_proj · 5-bit layers.43.self_attn.v_proj · 16-bit layers.43.mlp.down_proj · 4-bit layers.43.mlp.gate_proj · 4-bit layers.43.mlp.up_proj · 4-bit layers.44.linear_attn.in_proj_a · 16-bit layers.44.linear_attn.in_proj_b · 4-bit layers.44.linear_attn.in_proj_qkv · 4-bit layers.44.linear_attn.in_proj_z · 5-bit layers.44.linear_attn.out_proj · 5-bit layers.44.mlp.down_proj · 4-bit layers.44.mlp.gate_proj · 4-bit layers.44.mlp.up_proj · 4-bit layers.45.linear_attn.in_proj_a · 16-bit layers.45.linear_attn.in_proj_b · 16-bit layers.45.linear_attn.in_proj_qkv · 4-bit layers.45.linear_attn.in_proj_z · 4-bit layers.45.linear_attn.out_proj · 5-bit layers.45.mlp.down_proj · 4-bit layers.45.mlp.gate_proj · 4-bit layers.45.mlp.up_proj · 4-bit layers.46.linear_attn.in_proj_a · 16-bit layers.46.linear_attn.in_proj_b · 16-bit layers.46.linear_attn.in_proj_qkv · 5-bit layers.46.linear_attn.in_proj_z · 16-bit layers.46.linear_attn.out_proj · 5-bit layers.46.mlp.down_proj · 4-bit layers.46.mlp.gate_proj · 4-bit layers.46.mlp.up_proj · 4-bit layers.47.self_attn.k_proj · 16-bit layers.47.self_attn.o_proj · 4-bit layers.47.self_attn.q_proj · 4-bit layers.47.self_attn.v_proj · 16-bit layers.47.mlp.down_proj · 4-bit layers.47.mlp.gate_proj · 4-bit layers.47.mlp.up_proj · 4-bit layers.48.linear_attn.in_proj_a · 16-bit layers.48.linear_attn.in_proj_b · 16-bit layers.48.linear_attn.in_proj_qkv · 6-bit layers.48.linear_attn.in_proj_z · 16-bit layers.48.linear_attn.out_proj · 5-bit layers.48.mlp.down_proj · 4-bit layers.48.mlp.gate_proj · 4-bit layers.48.mlp.up_proj · 4-bit layers.49.linear_attn.in_proj_a · 16-bit layers.49.linear_attn.in_proj_b · 16-bit layers.49.linear_attn.in_proj_qkv · 6-bit layers.49.linear_attn.in_proj_z · 5-bit layers.49.linear_attn.out_proj · 5-bit layers.49.mlp.down_proj · 4-bit layers.49.mlp.gate_proj · 4-bit layers.49.mlp.up_proj · 4-bit layers.50.linear_attn.in_proj_a · 16-bit layers.50.linear_attn.in_proj_b · 16-bit layers.50.linear_attn.in_proj_qkv · 4-bit layers.50.linear_attn.in_proj_z · 5-bit layers.50.linear_attn.out_proj · 5-bit layers.50.mlp.down_proj · 4-bit layers.50.mlp.gate_proj · 4-bit layers.50.mlp.up_proj · 4-bit layers.51.self_attn.k_proj · 16-bit layers.51.self_attn.o_proj · 4-bit layers.51.self_attn.q_proj · 5-bit layers.51.self_attn.v_proj · 16-bit layers.51.mlp.down_proj · 4-bit layers.51.mlp.gate_proj · 4-bit layers.51.mlp.up_proj · 4-bit layers.52.linear_attn.in_proj_a · 16-bit layers.52.linear_attn.in_proj_b · 4-bit layers.52.linear_attn.in_proj_qkv · 5-bit layers.52.linear_attn.in_proj_z · 4-bit layers.52.linear_attn.out_proj · 5-bit layers.52.mlp.down_proj · 4-bit layers.52.mlp.gate_proj · 4-bit layers.52.mlp.up_proj · 4-bit layers.53.linear_attn.in_proj_a · 16-bit layers.53.linear_attn.in_proj_b · 16-bit layers.53.linear_attn.in_proj_qkv · 6-bit layers.53.linear_attn.in_proj_z · 4-bit layers.53.linear_attn.out_proj · 5-bit layers.53.mlp.down_proj · 4-bit layers.53.mlp.gate_proj · 4-bit layers.53.mlp.up_proj · 4-bit layers.54.linear_attn.in_proj_a · 16-bit layers.54.linear_attn.in_proj_b · 16-bit layers.54.linear_attn.in_proj_qkv · 5-bit layers.54.linear_attn.in_proj_z · 5-bit layers.54.linear_attn.out_proj · 4-bit layers.54.mlp.down_proj · 4-bit layers.54.mlp.gate_proj · 4-bit layers.54.mlp.up_proj · 4-bit layers.55.self_attn.k_proj · 16-bit layers.55.self_attn.o_proj · 6-bit layers.55.self_attn.q_proj · 4-bit layers.55.self_attn.v_proj · 16-bit layers.55.mlp.down_proj · 4-bit layers.55.mlp.gate_proj · 4-bit layers.55.mlp.up_proj · 4-bit layers.56.linear_attn.in_proj_a · 16-bit layers.56.linear_attn.in_proj_b · 16-bit layers.56.linear_attn.in_proj_qkv · 4-bit layers.56.linear_attn.in_proj_z · 16-bit layers.56.linear_attn.out_proj · 4-bit layers.56.mlp.down_proj · 4-bit layers.56.mlp.gate_proj · 4-bit layers.56.mlp.up_proj · 4-bit layers.57.linear_attn.in_proj_a · 16-bit layers.57.linear_attn.in_proj_b · 16-bit layers.57.linear_attn.in_proj_qkv · 5-bit layers.57.linear_attn.in_proj_z · 4-bit layers.57.linear_attn.out_proj · 4-bit layers.57.mlp.down_proj · 4-bit layers.57.mlp.gate_proj · 4-bit layers.57.mlp.up_proj · 4-bit layers.58.linear_attn.in_proj_a · 16-bit layers.58.linear_attn.in_proj_b · 4-bit layers.58.linear_attn.in_proj_qkv · 4-bit layers.58.linear_attn.in_proj_z · 5-bit layers.58.linear_attn.out_proj · 5-bit layers.58.mlp.down_proj · 4-bit layers.58.mlp.gate_proj · 4-bit layers.58.mlp.up_proj · 4-bit layers.59.self_attn.k_proj · 16-bit layers.59.self_attn.o_proj · 16-bit layers.59.self_attn.q_proj · 6-bit layers.59.self_attn.v_proj · 16-bit layers.59.mlp.down_proj · 4-bit layers.59.mlp.gate_proj · 4-bit layers.59.mlp.up_proj · 4-bit layers.60.linear_attn.in_proj_a · 16-bit layers.60.linear_attn.in_proj_b · 5-bit layers.60.linear_attn.in_proj_qkv · 5-bit layers.60.linear_attn.in_proj_z · 5-bit layers.60.linear_attn.out_proj · 4-bit layers.60.mlp.down_proj · 4-bit layers.60.mlp.gate_proj · 4-bit layers.60.mlp.up_proj · 4-bit layers.61.linear_attn.in_proj_a · 16-bit layers.61.linear_attn.in_proj_b · 16-bit layers.61.linear_attn.in_proj_qkv · 5-bit layers.61.linear_attn.in_proj_z · 4-bit layers.61.linear_attn.out_proj · 5-bit layers.61.mlp.down_proj · 6-bit layers.61.mlp.gate_proj · 4-bit layers.61.mlp.up_proj · 4-bit layers.62.linear_attn.in_proj_a · 16-bit layers.62.linear_attn.in_proj_b · 16-bit layers.62.linear_attn.in_proj_qkv · 4-bit layers.62.linear_attn.in_proj_z · 16-bit layers.62.linear_attn.out_proj · 4-bit layers.62.mlp.down_proj · 4-bit layers.62.mlp.gate_proj · 4-bit layers.62.mlp.up_proj · 8-bit layers.63.self_attn.k_proj · 16-bit layers.63.self_attn.o_proj · 16-bit layers.63.self_attn.q_proj · 8-bit layers.63.self_attn.v_proj · 16-bit layers.63.mlp.down_proj · 6-bit layers.63.mlp.gate_proj · 8-bit layers.63.mlp.up_proj · 8-bit 1 2 3 4 danger weight w (1 = most dangerous) parameters in the tensor (log scale) → 16-bit (146) 6- or 8-bit, held by floors (56) 4- or 5-bit (294)
Every scored tensor in plan C2. Horizontal: parameters in the tensor (log scale). Vertical: danger weight. The tensors line up in vertical stripes because the model has only a few distinct tensor sizes (for example, all the 31,457,280-parameter tensors share one stripe). The dashed line marks danger per million parameters = 0.0344, the point where the spare budget ran out. Every tensor above it received 16-bit (filled purple), except 2 tiny in-between tensor(s) that took a partial upgrade. Many of the most dangerous tensors are large (right side), so they sit below the line and stay at their floors (rings). Hover a dot for its name.
Checked, not assumed

On this build (496 scored): the lowest danger per million parameters among the 146 tensors that got 16-bit is 0.0277; the highest among the tensors left at their floor is 0.0411. The two groups overlap, so value for money alone no longer explains the split. The median tensor that got 16-bit has only 245,760 parameters, small next to the 89,128,960 of the largest tensors. As a second test, an independent 40-line re-solve of phase 1 (same rules 1–3, same floors, same deficit) differs from the real solver plan on 0 of 497 tensors, and the minimum deficit is 2250.72.

5.5  What silence does: the unscored tensors

A tensor with no score has w = 0, so upgrading it removes no danger but still costs budget. Primary therefore never spends anything on it and it stays at its floor. 1 tensors are unscored now (5.0% of all parameters): 0 in_proj_a, 0 in_proj_b, and 1 others. Floor-plus-KL treats them very differently, because KL data exists for them.

Plan A2 floors + KL 1 at 6-bit Plan C2 floor + primary 1 at 6-bit Plan C3 primary + KL fallback 1 at 6-bit Same 1 tensors in every bar: the ones with no Hessian score (0 in_proj_a, 0 in_proj_b, 1 others).
The same 1 tensors in three plans. Plan A2 lets KL decide and keeps 0 of them at 16-bit. Plan C2 has no opinion about them, so they fall to their floors: 0 at 4-bit, 0 at 5-bit, 1 at 6-bit. Plan C3 gives them their KL ranking (section 5.6) and keeps 0 at 16-bit. In the same C2 plan, 88 of the 96 scored in_proj_a/in_proj_b tensors get 16-bit. Whether a tiny gate tensor is kept at 16-bit therefore depends on whether it happened to be scored. Since the plans were first run, 89 more tensors have been scored, so 1 are unscored now (was 90); plain primary now pushes 0% of them to 4-bit (was 71%).

5.6  The fix: unscored tensors borrow their KL ranking

A missing score should not mean “safest”. The fix, added to the solver on 2026-09-19 as an opt-in flag, gives each unscored tensor a danger weight from the KL measurement it does have. Take the tensor’s own measured KL when squeezed to the lowest bit-width (4-bit, the hardest squeeze), sort all 497 tensors by that number from lowest to highest, and use the tensor’s position as a percentile:

wt = it ÷ (N − 1)   for every tensor with no Hessian score In words: N = 497 tensors, it is the tensor’s position when all of them are sorted by their 4-bit KL, lowest first. It is the same formula as section 3, only the source of the number changes: the most-damaged tensor gets 1.0, the least-damaged 0.0, so the direction matches the Hessian danger weight. Scored tensors and every floor are untouched.
Measured now (496 scored)C2: plain primaryC3: primary + KL fallback
Unscored tensors kept at 16-bit00
Unscored tensors at 6-bit11
Scored tensors at 16-bit146146
Scored tensors whose bit-width changed—0
Tensors at 16-bit, all 497146146
Raw summed KL (production: 37.45)40.2740.27

The fix is cheap for the tensors that were already scored: only 0 of them move and 0 fewer reach 16-bit, because the unscored tensors are 5.0% of all parameters and most are tiny. It removes only 0.0 of the 2.8 points by which plain primary sits above production, so most of that penalty is not the unscored squeeze.

Honest limits of the fix
  • A Hessian rank and a KL rank are different signals placed on the same 0–1 scale. They disagree tensor by tensor (see the tensor explorer), so this is a sensible stand-in, not an equivalent.
  • It fades out by itself: every tensor the sweep scores leaves the fallback set, and at full coverage the flag changes nothing.
  • It has only been run as a solver plan, not built and benchmarked. An independent re-solve using the solver’s own fallback function differs from the real solver plan for C3 on 0 of 497 tensors; the default (zero) differs for plain primary (C2) on 0.

6 · Why a second phase exists

The plain idea: decide the first question, lock the answer, then settle the second question only among the answers that tie on the first. You will also see this called “lexicographic” or “dictionary” order; it is the same thing as sorting a phone book by surname and using the first name only when surnames match.

Every way to give each tensor one bit-width 497 tensors × 5 candidates: 4, 5, 6, 8 or 16 … that respect every floor and stay under the bits-per-weight ceiling the feasible set … that also reach the lowest possible danger deficit Phase 1 finds f₁*, then locks f₁ ≤ f₁* + τ … the one with the lowest KL Phase 2
Schematic of the two phases (the bands show the idea, not measured counts).

The math

Two quantities are involved. f1 is the danger deficit (section 5). f2 is the weighted KL damage, using the measured damage of each option scaled by st, a structural weight (100 for the 16 boundary and output tensors, 1 otherwise when the soft Hessian weight is 0).

f1(x) = Σ xt,b · wt · (16 − b)    f2(x) = Σ xt,b · KLt,b · st
Phase 1:  f1* = the smallest possible f1(x), subject to rules 1–3 In words: find the best danger deficit any allowed allocation can reach, and call it f1*.
Phase 2:  minimize f2(x), subject to rules 1–3 and f1(x) ≤ f1* + τ In words: among only those allocations whose danger deficit is no worse than f1* (plus a tiny tolerance τ), pick the one with the lowest KL damage. The tolerance τ is the largest of 10−6, |f1*| × 10−9, and |f1*| times the solver’s achieved gap. On plan C2: f1* = 2250.72 and τ = 2.25e-06, so phase 2 may not accept anything worse than phase 1’s best.

Why not just add the two together?

The obvious alternative is one objective, f1 + λ · f2. The two numbers have different units (danger × bits versus summed KL), so λ has no natural value, and any λ lets one criterion buy back the other. That is exactly what failed with the soft weight: even at strength 1000 the KL term could not move the 20 most dangerous tensors. Lock-then-tie-break needs no λ.

What phase 2 actually did on this data

On plan C2 nothing. A fresh phase-1-only solve picks the same bit-width for all 497 tensors as the finished two-phase plan, so once danger was optimized, KL had no freedom left. The second phase is insurance for cases with real ties. The earlier note in RUNNING_GUIDE (“KL decides 3 of 497”) was measured at a different coverage.

7 · Is it better?

What has changed since the plans were first run

The same measurements, taken when the plans were first run and now, on the same solver recipes:

MeasurementWhen the plans were run (407 scored)Now (496 scored)
Tensors with a fragility score407496
Tensors without one901
Hessian floors (8-bit / 6-bit)21 / 4025 / 50
Tensors with no Hessian floor (scored + unscored)436422
Plain primary (C2): tensors at 16-bit68146
Plain primary (C2): unscored tensors at 4-bit64 of 900 of 1
Plain primary (C2): raw summed KL50.0940.27
Primary + KL fallback (C3): tensors at 16-bit149146
Primary + KL fallback (C3): raw summed KL39.4140.27
Floors + KL (A2): tensors at 16-bit142139
Floors + KL (A2): raw summed KL38.8939.62
Lowest danger per million parameters among C2’s 16-bit tensors0.02420.0277

The comparison

Nobody has measured that yet: no plan in this comparison has been built and benchmarked, and this page will not pretend otherwise. Here is what the real plans do show.

Raw summed KL (lower = smaller measured damage) Tensors kept at 16-bit production plan 37.45 153 A2 · floors + KL 39.62 139 B2 · primary + regional floors frozen: plan run at 407 scored 50.36 64 C2 · floor + primary 40.27 146 C3 · C2 + KL fallback 40.27 146
Raw summed KL is the very quantity a KL-first plan minimizes, so a Hessian-first plan scoring worse on it is expected; it is not a measure of real damage. Plan A2 (floors + KL) is about 2.2 above production. Plain primary (C2) is about 2.8 higher; B2 (frozen at the 407-tensor run) is 12.9 higher. The KL fallback (C3, section 5.6) narrows the gap only to 2.8.

What the evidence supports

Primary is a clear, testable statement of one hypothesis: curvature alone should decide who gets bits. It allocates by danger per parameter, as section 5 shows, with the exceptions noted there.

What it does not support yet

That this allocation protects quality better. No plan from this table has been benchmarked. Raw KL cannot settle it, because the two approaches disagree about which signal is right.

What has to happen first

Coverage, or the KL fallback while coverage is incomplete: with 1 of 497 tensors unscored, plain primary starves them (section 5.5). Then a real benchmark, naive first and the YAQA-corrected build only if promising.

Honest limits
  • Plans A2, C2 and C3 and every figure are computed at the last build (496 scored, 2026-09-30 01:14); a later build, or a fresh run of the commands in section 8 after the sweep moves on, gives slightly different numbers. Plan B2 is frozen at the 407-tensor run because its exact command is not known.
  • The allocation rule is verified only for plan C2 (floors + primary, no regional floors). Plan B2 adds regional floors on top and was not analysed the same way.
  • A2 and C2 differ in two ways, not one: primary on or off, and a 5-bit floor on 76 late-attention tensors (A2 has it, C2 does not). The comparison therefore does not isolate the effect of primary.
  • Raw KL here is summed isolated KL from the checkpoint, the same data the solver uses, not a measurement of the finished model.

8 · Which setting when

Both commands below already include the two fixes from 2026-09-18: --pareto none and --hessian-weight-strength 0. Run them from the project folder. Output paths are new names so no existing plan is overwritten. The score file keeps growing, so results will differ from this page.

Safest plan today

Floors + KL (plan A2)

Keeps KL as the judge for the 1 unscored tensors, and stays closest to production (39.62 raw KL versus 37.45). Plan A2 also carried a 5-bit floor on the 76 late-attention tensors.

Test the hypothesis

Floor + primary + KL fallback (plan C3)

The cleanest test of “does curvature predict damage?” while coverage is incomplete, because unscored tensors keep a KL-based weight instead of being starved. Plain primary (plan C2) is only meaningful at full coverage. Either one needs a benchmark first.

Do not use

Soft weight above 0

It changes nothing once a floor or primary is active, and it inflates the reported weighted KL.

Floors + KL: the recipe of plan A2
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 \
  --candidates 4,5,6,8,16 \
  --pareto none \
  --group-size 64 \
  --late-attn-min-bits 5 \
  --early-qkv-min-bits 0 \
  --lm-head-min-bits 6 \
  --boundary-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 0 \
  --hessian-floor-tiers "0.05:8,0.15:6" \
  --max-low-bit-run 999 \
  --output 04_LANGUAGE_PLANS/HESSIAN_TEST/plan_floor_kl_2026-09-19.json
Floor + primary: the recipe of plan C2
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 \
  --candidates 4,5,6,8,16 \
  --pareto none \
  --group-size 64 \
  --late-attn-min-bits 0 \
  --early-qkv-min-bits 0 \
  --lm-head-min-bits 6 \
  --boundary-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 0 \
  --hessian-floor-tiers "0.05:8,0.15:6" \
  --hessian-primary \
  --max-low-bit-run 999 \
  --output 04_LANGUAGE_PLANS/HESSIAN_TEST/plan_floor_primary_2026-09-19.json
Floor + primary + KL fallback: the recipe of plan C3 (recommended while coverage is incomplete)
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 \
  --candidates 4,5,6,8,16 \
  --pareto none \
  --group-size 64 \
  --late-attn-min-bits 0 \
  --early-qkv-min-bits 0 \
  --lm-head-min-bits 6 \
  --boundary-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 0 \
  --hessian-floor-tiers "0.05:8,0.15:6" \
  --hessian-primary \
  --hessian-unscored-fallback kl \
  --max-low-bit-run 999 \
  --output 04_LANGUAGE_PLANS/HESSIAN_TEST/plan_floor_primary_klfallback_2026-09-19.json

--max-low-bit-run 999 switches off the separate run-guard, which is required alongside primary (see the MILP-aware guide). The first two commands differ in two ways, exactly as plans A2 and C2 did: the --hessian-primary flag and the late-attention floor (--late-attn-min-bits, 5 versus 0). The plan files do not record the run-guard setting used for A2, so it is kept disabled in both commands.

9 · Glossary

BPW (bits per weight)

The average number of bits spent per parameter across the whole model. The budget the solver may not exceed.

Hessian

A matrix describing how sharply a tensor’s output error curves as its weights change. The raw score summarizes how concentrated that curvature is.

KL (Kullback–Leibler divergence)

A measure of how far the quantized model’s output distribution moves from the original’s. Measured per tensor and per bit-width.

MILP

Mixed-integer linear program: pick whole-number (here yes/no) variables to minimize a linear total, subject to linear limits. Solved exactly by HiGHS.

Deficit

w × (16 − bits): how much danger a tensor leaves unprotected by not being at 16-bit.

Floor

A minimum bit-width forced onto a tensor before optimizing starts.

Danger per million parameters

w ÷ (parameters ÷ 1,000,000). The value-for-money ranking phase 1 effectively follows.

Tie

Two tensors with exactly the same raw score. The sort keeps them in checkpoint order, so file order settles their rank.

Pin tolerance τ

How much worse than the best phase-1 value phase 2 is allowed to accept. Here about 0.0000023.

Lexicographic

Decide on the first criterion; use the second only among ties. Same idea as dictionary order.

Unscored tensor

A tensor absent from the score file. It gets danger weight 0.

KL fallback

Opt-in flag --hessian-unscored-fallback kl: an unscored tensor gets the percentile rank of its own 4-bit KL as danger weight instead of 0.

Coverage

How many of the 497 tensors have a score. 407 when the plans were first run, 496 at the last build.

10 · How every number is produced and checked, and related pages

Every number and diagram on this page is produced by one script, hessian_primary_page.py, from the live score file and real solver runs. It calls the solver’s own compute_hessian_danger_weight and compute_hessian_floor_map, runs the real solver for plans A2, C2 and C3, cross-checks C2 and C3 with an independent re-solve of phase 1, and rebuilds this page whenever the score file, the solver or the sweep status changes. If a measured outcome changes, the sentence that describes it changes with it.

Rebuild the page now from live data
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_primary_page.py build --force
Prove the page is honest (independent re-solve, tag balance, key numbers present); add --deep to also re-run the real solver on the frozen snapshot
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_primary_page.py check
Keep it updated in the background, checking every 5 minutes (survives closing the terminal; replaces an older watcher already running)
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_primary_page.py restart 300
Check on it, read its log, stop it
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_primary_page.py status
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_primary_page.py logs
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_primary_page.py stop

The page does not refresh itself: reload the tab (Cmd+R) to see the latest build. The time in the “Live status” box tells you when it was built.

11 · Research: layered floors and the two signals

A working area, not a conclusion. Right now the allocation is driven by the Hessian score alone, and only 496 of 497 tensors have one; the godmode sweep that will add a second, richer signal is still running. This section asks one question on today’s data: what does adding a layered floor (8-bit, then 6-bit, then 5-bit) do to both signals at once, on each plan?

What is measured

Every point is a real solver run at today’s coverage: six floor layouts on each of three plans, 17 runs. The weights behind the deficit do not depend on the floors, so the numbers are comparable across layouts.

2141 36.1 2288 39.8 2435 43.6 2583 47.4 2730 51.2 no Hessian floors · A2: deficit 2689, KL 42.94, 106 at 16-bit 0 no Hessian floors · C2: deficit 2182, KL 37.33, 171 at 16-bit 0 no Hessian floors · C3: deficit 2182, KL 37.33, 171 at 16-bit 0 top 5% only · A2: deficit 2425, KL 37.78, 145 at 16-bit 1 top 5% only · C2: deficit 2217, KL 38.39, 160 at 16-bit 1 top 5% only · C3: deficit 2217, KL 38.39, 160 at 16-bit 1 second tier at 5-bit · A2: deficit 2414, KL 38.36, 143 at 16-bit 2 second tier at 5-bit · C2: deficit 2230, KL 38.90, 155 at 16-bit 2 second tier at 5-bit · C3: deficit 2230, KL 38.90, 155 at 16-bit 2 current: 8 then 6 · A2: deficit 2397, KL 39.62, 139 at 16-bit 3 current: 8 then 6 · C2: deficit 2251, KL 40.27, 146 at 16-bit 3 current: 8 then 6 · C3: deficit 2251, KL 40.27, 146 at 16-bit 3 add a 5-bit layer to 30% · A2: deficit 2462, KL 41.01, 119 at 16-bit 4 add a 5-bit layer to 30% · C2: deficit 2331, KL 40.67, 138 at 16-bit 4 add a 5-bit layer to 30% · C3: deficit 2331, KL 40.67, 138 at 16-bit 4 add a 5-bit layer to 50% · C2: deficit 2555, KL 49.92, 64 at 16-bit 5 add a 5-bit layer to 50% · C3: deficit 2555, KL 49.92, 64 at 16-bit 5 Hessian danger deficit over scored tensors (lower = fragile tensors better protected) → worse raw summed KL (lower = less measured damage) → worse floors + KL (A2) floor + primary (C2) primary + KL fallback (C3) ringed = current floors
Each dot is one run, numbered by floor layout (0 none, 1 top 5% only, 2 second tier at 5-bit, 3 current, 4 add a 5-bit layer to 30%, 5 add a 5-bit layer to 50%). Lower left is better on both signals. The dashed line joins the 2 of 17 runs that no other run beats on both.
Floor layoutFlooredFloor-only BPWHessian deficitRaw KLAt 16-bitMean bits, Hessian-criticalMean bits, KL-criticalAgainst the current layout
floors + KL (A2)
0 · no Hessian floors04.317268942.941065.88.4both signals worse (+292 deficit, +3.32 KL)
1 · top 5% only254.484242537.781457.610.7KL better, Hessian worse (+28 deficit, -1.84 KL)
2 · second tier at 5-bit754.541241438.361437.810.7KL better, Hessian worse (+17 deficit, -1.27 KL)
3 · current: 8 then 6754.631239739.621398.29.2baseline
4 · add a 5-bit layer to 30%1494.807246241.011197.88.9both signals worse (+65 deficit, +1.39 KL)
5 · add a 5-bit layer to 50%no solution: solver run for A2 failed: V3.3 weighted-KL solve did not meet acceptance criteria. message=The problem is infeasible. (H
floor + primary (C2)
0 · no Hessian floors04.317218237.331718.68.2both signals better (-69 deficit, -2.94 KL)
1 · top 5% only254.484221738.391609.98.4both signals better (-34 deficit, -1.88 KL)
2 · second tier at 5-bit754.541223038.9015510.08.3both signals better (-20 deficit, -1.37 KL)
3 · current: 8 then 6754.631225140.2714610.27.9baseline
4 · add a 5-bit layer to 30%1494.807233140.671389.27.9both signals worse (+80 deficit, +0.40 KL)
5 · add a 5-bit layer to 50%2485.021255549.92647.06.3both signals worse (+304 deficit, +9.65 KL)
primary + KL fallback (C3)
0 · no Hessian floors04.317218237.331718.68.2both signals better (-69 deficit, -2.94 KL)
1 · top 5% only254.484221738.391609.98.4both signals better (-34 deficit, -1.88 KL)
2 · second tier at 5-bit754.541223038.9015510.08.3both signals better (-20 deficit, -1.37 KL)
3 · current: 8 then 6754.631225140.2714610.27.9baseline
4 · add a 5-bit layer to 30%1494.807233140.671389.27.9both signals worse (+80 deficit, +0.40 KL)
5 · add a 5-bit layer to 50%2485.021255549.92647.06.3both signals worse (+304 deficit, +9.65 KL)

Which layout does each signal prefer?

What the Hessian floors themselves do

For the two primary plans the floors are extra constraints on the very number being minimized, so a lower deficit without floors is expected and proves nothing. What the floors buy is a guarantee that the Hessian-critical tensors keep at least 8 or 6 bits, which the deficit does not price (it counts every deficit point the same whatever the tensor’s size). For the floors + KL plan the story is the opposite: the floors help both numbers.

The key experiment: adding a 5-bit layer

Do the two signals agree?

Do the two signals agree on who is critical? Rank correlation between danger weight and 4-bit KL over the 496 scored tensors: -0.16. Of the 50 most Hessian-critical scored tensors, only 8 are also among the 50 most KL-critical scored tensors (about 5 would overlap by pure chance). The rest are the two disagreement corners the tensor explorer shows: fragile by curvature but harmless by KL, and the reverse. On the current floors the two objectives already pull apart: floors + KL (A2) reaches raw KL 39.62 with Hessian deficit 2397, while plain primary (C2) reaches deficit 2251 with raw KL 40.27.

What this cannot tell us

Proxies, not quality

Both measurements are proxies computed from isolated per-tensor data. Neither says which model is better. Only benchmarks of models built from these plans can. The existing suite (MMLU, GSM8K, IFEval, BFCL and HumanEval at about 100 questions each, ledger §09) may be too coarse for this: six real builds sit within 0.46 points of each other (91.68% to 92.14% on the 5-test mean), so plans that differ by a small amount will not separate at that sample size without larger samples or a finer measure such as divergence from the BF16 model on held-out text.

UnknownWhat would settle itNeeds
Does Hessian-first allocation give a better model than KL-first at the same size?Benchmark A2 against C3 (same budget, same checkpoint lineage)Build both; larger samples or a finer measure
Do layered floors (the 5-bit layer) help or hurt real quality?Benchmark C3 against C3 with the 5-bit layer to 30% (layout 4)Build one more model; same measurement
Does the KL fallback for unscored tensors matter?Benchmark C2 against C3Only worth it while coverage is incomplete
When the two signals disagree, which one predicts real damage?An ablation that protects only Hessian-critical/KL-benign tensors, and one that protects only KL-critical/Hessian-benign tensorsTwo extra builds; the disagreement lists come from this section
Do today’s answers survive full coverage and the godmode sweep?Rebuild this page after the sweep finishes and compare; then benchmark the shortlistThe sweep to finish

A sensible order, cheapest first: naive-affine builds of A2, C3 and C3 with the 5-bit layer; benchmark; build the YAQA-corrected version only for the plan that looks promising. Building and benchmarking is your call to trigger; the build steps are in RUNNING_GUIDE. When the godmode sweep finishes, this section is rebuilt automatically, so both solvers can be compared on full Hessian and full godmode data.

Built 2026-09-30 01:14 by hessian_primary_page.py for Hakim Ghelab, VegaLaboratories LTD. Numbers describe the score file at 496 of 497 tensors scored.