Why lm_head landed at Q4, exactly what it broke, and exactly what's still valid
A step-by-step, code-verified trace of the mechanism, the exact real blast radius (497 tensors checked, not estimated), and a precise verdict on what can be salvaged for free versus what genuinely requires re-measurement.
This incident surfaced while overlaying the real, live stratified-calibration cascade documented on The Free Signal Hiding in Your Hessians with the Hessian score — that page is the real substrate this incident happened inside, not a separate investigation.
The exact code mechanism, verified line by line
From 06_measure_sensitivity_cascaded_V4.py, the real per-tensor cascade loop. Read at lines 261-304.
Two real, load-bearing facts follow directly from this order:
Because lm_head is last in causal order, its lock-in step runs after every other tensor has already been measured. No other tensor's forward pass ever sees lm_head at anything but BF16. Whatever P0 or P1 said about lm_head specifically had zero effect on any other tensor's measured sensitivity, in any round.
Any tensor that changes bit-width between the original and corrected P1 gets locked in at its new value before every causally-later tensor is measured in the next round. If a floored P1 changes tensors other than lm_head, round 2's checkpoint — measured against the original P1 — is stale for everything downstream of the earliest changed tensor.
Milestone 1 — does P0→P1 vary "enormously"? Real answer, not an adjective.
Re-ran the real V3.3 MILP against the real, unmodified cascaded_checkpoint_round1.json with --force-regex "lm_head=6" added. Diffed all 497 tensors against the original plan_v4_round1.json, one by one.
12 of 497 tensors changed (2.4%) — not "enormous" in scale, but not zero either, and the identity of what changed matters more than the count.
| Tensor | Original P1 | Floored P1 | Params |
|---|---|---|---|
| language_model.lm_head | Q4 | Q6 | 1271.4M |
| layers.3.self_attn.o_proj | Q5 | Q4 | 31.5M |
| layers.5.linear_attn.in_proj_z | Q6 | Q5 | 31.5M |
| layers.52.linear_attn.out_proj | Q16 | Q5 | 31.5M |
| layers.53.linear_attn.in_proj_z | Q16 | Q5 | 31.5M |
| layers.53.linear_attn.out_proj | Q16 | Q5 | 31.5M |
| layers.54.linear_attn.in_proj_z | Q16 | Q5 | 31.5M |
| layers.54.linear_attn.out_proj | Q16 | Q5 | 31.5M |
| layers.55.self_attn.o_proj | Q16 | Q5 | 31.5M |
| layers.56.linear_attn.out_proj | Q16 | Q5 | 31.5M |
| layers.3.self_attn.k_proj | Q16 | Q6 | 5.2M |
| layers.1.linear_attn.in_proj_b | Q16 | Q6 | 0.2M |
Most of the 11 non-lm_head changes are late layers (52-56) that were sitting at BF16 above the late-attention floor and got pulled down to exactly the floor's own minimum (Q5) to help fund lm_head — the hard floor held, nothing broke a constraint, but real headroom that used to exist there is now gone. The two that matter most for what comes next are layers 1 and 3 — early tensors, which is exactly the problem in Milestone 2.
Milestone 2 — is round 2's existing checkpoint still valid once P1 is corrected?
Layers 1 and 3 are early in causal order. Every tensor measured after them in round 2's sweep — which is nearly the entire model — was measured with layer 1 and layer 3 locked at the original P1's bit-widths (Q16 and Q5 respectively), not the corrected ones (Q6 and Q4). cascaded_checkpoint_round2.json reflects a context that no longer exists once P1 is fixed. Re-optimizing against it and calling the result "round 2, lm_head protected" is not a valid reconstruction — it's exactly the shortcut you flagged.
| Round | Measured against context | Still valid after the lm_head fix? | Why |
|---|---|---|---|
| Round 1 checkpoint | P0 (untouched, fixed baseline) | VALID | P0 never changes. Re-optimizing this real checkpoint with the floor added is a legitimate, free re-derivation. |
| Round 2 checkpoint | Original P1 (now superseded) | STALE | Layers 1 & 3's context values changed under the floor. Nearly every round-2 measurement ran with the wrong upstream context. |
| Round 3 (currently running) | Original P2 (built from stale round 2) | STALE | Inherits round 2's problem one layer further down the chain. |
What the native vendor OptiQ algorithm does on the exact same real checkpoint
Not simulated — actually ran optiq.core.optimizer.optimize_mixed_precision, the vendor's own unmodified greedy allocator, against the real cascaded_checkpoint_round1.json, same 5.028 BPW target.
| Solver | lm_head | Q4 | Q5 | Q6 | Q8 | Q16 |
|---|---|---|---|---|---|---|
| Our V3.3 MILP (original) | Q4 | 236 | 85 | 11 | 3 | 162 |
| Our V3.3 MILP (lm_head floored) | Q6 | 236 | 92 | 13 | 3 | 153 |
| Vendor native greedy (unmodified) | Q6 | 211 | 109 | 48 | 0 | 129 |
The vendor's own algorithm — same protected-tensor list, same 100x weight, same real cascaded data, zero modification by us — never put lm_head below Q6. Its overall shape also differs more broadly from ours (fewer Q4, far more Q6, no Q8 at all) — a real, separate finding about how differently a myopic greedy allocator and a joint MILP solver carve up the same budget, worth its own investigation later but not the subject of this document.
What this actually costs, precisely — no rounding up or down
Real checkpoint already on disk, measured against the untouched P0. Re-optimizing it with the floor took under 0.1 seconds. Already done — /tmp/plan_v4_round1_LMHEAD_FLOORED.json is a real, valid, corrected P1.
Must reload the model fresh, lock in the CORRECTED P1's bits (not the original), and re-run the full 497-tensor cascaded sweep for real. This is the expensive step — this project's own pages elsewhere cite the full 497-tensor cascaded measurement at 40+ real hours.
It's measuring against P2, which was itself built from the stale round-2 checkpoint. Letting it finish still produces real, informative data — but a P3 built on top of it inherits the same staleness one layer further down.
What I am and am not claiming
- The exact code mechanism (causal-order lock-in), read at the real line numbers.
- The exact 12-tensor diff, computed from two real re-solves, not estimated.
- Round 1's correction is legitimate; round 2 onward is not, without new measurement.
- The vendor's own unmodified algorithm agrees with our floored version, not our original.
- Whether to let the live round 3 finish anyway, or stop it now.
- Whether to redo round 2 fresh from the corrected P1 before touching anything else.
- The real GPU-hour cost of a fresh round 2 on this exact machine — not measured yet, only cited from an earlier full-run figure elsewhere in this project.