← Back to index ← Back to research index
Precision incident analysis · real data only, every number sourced below · 2026-09-12, updated 2026-09-13

Why lm_head landed at Q4, exactly what it broke, and exactly what's still valid

A step-by-step, code-verified trace of the mechanism, the exact real blast radius (497 tensors checked, not estimated), and a precise verdict on what can be salvaged for free versus what genuinely requires re-measurement.

Where this was actually found (added 2026-09-17)

This incident surfaced while overlaying the real, live stratified-calibration cascade documented on The Free Signal Hiding in Your Hessians with the Hessian score — that page is the real substrate this incident happened inside, not a separate investigation.

The exact code mechanism, verified line by line

From 06_measure_sensitivity_cascaded_V4.py, the real per-tensor cascade loop. Read at lines 261-304.

for layer_idx, (layer_name, module) in enumerate(quantizable): # causal order, lm_head LAST for bits in candidate_bits: module.weight = _simulate_quantize(orig_w, bits, group_size) # ... measure this tensor's own KL at each candidate bit ... planned = baseline_bits[layer_name] # THE CONTEXT plan's choice for THIS tensor module.weight = quantize(orig_w, planned) # locked in AFTER measurement, before the NEXT tensor

Two real, load-bearing facts follow directly from this order:

✓ lm_head's own context value never leaks

Because lm_head is last in causal order, its lock-in step runs after every other tensor has already been measured. No other tensor's forward pass ever sees lm_head at anything but BF16. Whatever P0 or P1 said about lm_head specifically had zero effect on any other tensor's measured sensitivity, in any round.

✗ every OTHER tensor's context value does leak forward

Any tensor that changes bit-width between the original and corrected P1 gets locked in at its new value before every causally-later tensor is measured in the next round. If a floored P1 changes tensors other than lm_head, round 2's checkpoint — measured against the original P1 — is stale for everything downstream of the earliest changed tensor.

Milestone 1 — does P0→P1 vary "enormously"? Real answer, not an adjective.

Re-ran the real V3.3 MILP against the real, unmodified cascaded_checkpoint_round1.json with --force-regex "lm_head=6" added. Diffed all 497 tensors against the original plan_v4_round1.json, one by one.

Verdict
No — not by tensor count. Yes — by exactly which tensors moved.

12 of 497 tensors changed (2.4%) — not "enormous" in scale, but not zero either, and the identity of what changed matters more than the count.

TensorOriginal P1Floored P1Params
language_model.lm_headQ4Q61271.4M
layers.3.self_attn.o_projQ5Q431.5M
layers.5.linear_attn.in_proj_zQ6Q531.5M
layers.52.linear_attn.out_projQ16Q531.5M
layers.53.linear_attn.in_proj_zQ16Q531.5M
layers.53.linear_attn.out_projQ16Q531.5M
layers.54.linear_attn.in_proj_zQ16Q531.5M
layers.54.linear_attn.out_projQ16Q531.5M
layers.55.self_attn.o_projQ16Q531.5M
layers.56.linear_attn.out_projQ16Q531.5M
layers.3.self_attn.k_projQ16Q65.2M
layers.1.linear_attn.in_proj_bQ16Q60.2M
Read this table exactly, don't round it off

Most of the 11 non-lm_head changes are late layers (52-56) that were sitting at BF16 above the late-attention floor and got pulled down to exactly the floor's own minimum (Q5) to help fund lm_head — the hard floor held, nothing broke a constraint, but real headroom that used to exist there is now gone. The two that matter most for what comes next are layers 1 and 3 — early tensors, which is exactly the problem in Milestone 2.

Milestone 2 — is round 2's existing checkpoint still valid once P1 is corrected?

Verdict
No. This is the real cheat you caught.

Layers 1 and 3 are early in causal order. Every tensor measured after them in round 2's sweep — which is nearly the entire model — was measured with layer 1 and layer 3 locked at the original P1's bit-widths (Q16 and Q5 respectively), not the corrected ones (Q6 and Q4). cascaded_checkpoint_round2.json reflects a context that no longer exists once P1 is fixed. Re-optimizing against it and calling the result "round 2, lm_head protected" is not a valid reconstruction — it's exactly the shortcut you flagged.

RoundMeasured against contextStill valid after the lm_head fix?Why
Round 1 checkpointP0 (untouched, fixed baseline)VALIDP0 never changes. Re-optimizing this real checkpoint with the floor added is a legitimate, free re-derivation.
Round 2 checkpointOriginal P1 (now superseded)STALELayers 1 & 3's context values changed under the floor. Nearly every round-2 measurement ran with the wrong upstream context.
Round 3 (currently running)Original P2 (built from stale round 2)STALEInherits round 2's problem one layer further down the chain.

What the native vendor OptiQ algorithm does on the exact same real checkpoint

Not simulated — actually ran optiq.core.optimizer.optimize_mixed_precision, the vendor's own unmodified greedy allocator, against the real cascaded_checkpoint_round1.json, same 5.028 BPW target.

Solverlm_headQ4Q5Q6Q8Q16
Our V3.3 MILP (original)Q423685113162
Our V3.3 MILP (lm_head floored)Q623692133153
Vendor native greedy (unmodified)Q6211109480129

The vendor's own algorithm — same protected-tensor list, same 100x weight, same real cascaded data, zero modification by us — never put lm_head below Q6. Its overall shape also differs more broadly from ours (fewer Q4, far more Q6, no Q8 at all) — a real, separate finding about how differently a myopic greedy allocator and a joint MILP solver carve up the same budget, worth its own investigation later but not the subject of this document.

What this actually costs, precisely — no rounding up or down

Round 1
Fully salvageable, zero new GPU cost.

Real checkpoint already on disk, measured against the untouched P0. Re-optimizing it with the floor took under 0.1 seconds. Already done — /tmp/plan_v4_round1_LMHEAD_FLOORED.json is a real, valid, corrected P1.

Round 2
Not salvageable. Requires a real re-measurement pass.

Must reload the model fresh, lock in the CORRECTED P1's bits (not the original), and re-run the full 497-tensor cascaded sweep for real. This is the expensive step — this project's own pages elsewhere cite the full 497-tensor cascaded measurement at 40+ real hours.

Round 3
Currently live, running on a stale foundation.

It's measuring against P2, which was itself built from the stale round-2 checkpoint. Letting it finish still produces real, informative data — but a P3 built on top of it inherits the same staleness one layer further down.

What I am and am not claiming

Established with real, verified evidence
  • The exact code mechanism (causal-order lock-in), read at the real line numbers.
  • The exact 12-tensor diff, computed from two real re-solves, not estimated.
  • Round 1's correction is legitimate; round 2 onward is not, without new measurement.
  • The vendor's own unmodified algorithm agrees with our floored version, not our original.
Still open, not decided here
  • Whether to let the live round 3 finish anyway, or stop it now.
  • Whether to redo round 2 fresh from the corrected P1 before touching anything else.
  • The real GPU-hour cost of a fresh round 2 on this exact machine — not measured yet, only cited from an earlier full-run figure elsewhere in this project.
Author: Hakim Ghelab, VegaLaboratories LTD · Every number on this page came from a real command run against real files on this machine during this analysis — none estimated, none carried over from a different regime without being re-verified here.