The lm_head fix: exact commands, in exact order, for everything that comes next
The optimizer script itself is fixed (real, default-on hard floor, verified working). This page is the exact sequence to actually use it — a new folder is already seeded and ready.
A full cascade re-run with the boundary floor on was judged not worth the real compute cost. The real fix kept the existing cascaded checkpoint from the real, live P0→P3 cascade and rebuilt only the plan from it with the boundary floor added — a clean cascaded, lm_head-protected plan, with the Hessian not yet involved. That exact plan is what §09 later benchmarked.
Step 1 — resume the cascade in the new, clean folder
Already done for you: 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/ exists and already contains the real, valid cascaded_checkpoint_round1.json (copied unchanged from the live run — proven valid in the incident doc, since it was measured against P0, which never changes).
Run this exact command when you're ready (identical to the live process's own command, only --output-dir changed):
Round 1: sees all 497 tensors already have a real result in the seeded checkpoint, skips straight to the optimizer — which is now fixed by default, so it writes a corrected plan_v4_round1.json (lm_head=Q6) in under a second. Zero new GPU time.
Round 2: reloads the model fresh, locks in the corrected P1's bits, and does a real, full 497-tensor cascaded measurement — this is genuine new GPU time, no way around it, because layers 1 and 3 also changed under the floor and everything downstream of them needs re-measuring against the real new context.
Round 3: same pattern again, real GPU time, using round 2's real output as context.
You can run this in parallel with the live (uncorrected) round-3 process if you want both data points — they write to different folders and don't conflict.
Step 2 — your test sequence, as you laid it out
| # | Action | Purpose |
|---|---|---|
| 1 | Let the live (uncorrected) round 3 finish | Real data point on the original, un-floored cascade |
| 2 | Run the corrected cascade above (this page) | Real P1/P2/P3 with lm_head genuinely protected throughout |
| 3 | Pick a plan to benchmark — you're leaning P2, based on the flat-N8 precedent (P0→P1 changed a lot, P1→P2 much less, P2→P3 mostly oscillation) | One real regression benchmark settles the lm_head question empirically, not just on paper |
| 4 | Compare against your existing reference models on the same benchmark suite | Confirms (or corrects) everything this whole investigation predicted |
| 5 | Only then: bake the final YAQA-corrected model | Stage 2 sits on top of a validated Stage 1 plan, not a provisional one |
Your question: GODMODE multi-bit calibration first, or build with what we have?
You don't need GODMODE for this specific correction. Here's why: GODMODE precomputes YAQA correction at every candidate bit-width (4,5,6,8,16) for every tensor up front, so any future plan change costs zero additional compute, ever. That's real value if you expect to keep changing individual tensors' bits repeatedly. But for what you're actually doing — one plan (P2, or whichever you pick) built once, corrected once — the existing resume-cache mechanism already gives you the same practical benefit for free (see Step 3 below). GODMODE would be solving a problem you don't currently have. Save it for if/when you're doing repeated exploratory re-bakes of many candidate plans.
Your question: does baking YAQA for a corrected plan mean starting from scratch?
05_full_model_quantize.py's resume-skip logic (read at the real line numbers, lines 733-758) checks bits, group_size, and curvature_version per tensor by name against whatever plan you point it at. Point a new build at the same --resume-dir as an existing YAQA build made from a different plan, and:
- Any tensor whose bit-width is identical between the old and new plan is automatically skipped — its already-computed corrected weights are reused as-is.
- Any tensor whose bit-width differs is automatically recomputed — nothing stale gets silently reused.
This matches exactly what you remembered: YAQA correction is per-tensor independent, and the Hessian itself is bit-width-independent and reusable. Given the real diff we measured — 12 of 497 tensors changed between the original and lm_head-floored P1 (2.4%) — if you already have (or build) a resume-dir cache from any closely-related prior plan, pointing a new build at that same directory should reuse roughly 95%+ of the work automatically. The one thing to check before relying on this: which plan your existing YAQA resume directories were actually built from, so you know how much real overlap to expect — I haven't verified that against your current final plan choice yet.