← Back to index ← Back to research index
Operational playbook — copy-pasteable, nothing to remember · 2026-09-12, updated 2026-09-13

The lm_head fix: exact commands, in exact order, for everything that comes next

The optimizer script itself is fixed (real, default-on hard floor, verified working). This page is the exact sequence to actually use it — a new folder is already seeded and ready.

Real context (added 2026-09-17)

A full cascade re-run with the boundary floor on was judged not worth the real compute cost. The real fix kept the existing cascaded checkpoint from the real, live P0→P3 cascade and rebuilt only the plan from it with the boundary floor added — a clean cascaded, lm_head-protected plan, with the Hessian not yet involved. That exact plan is what §09 later benchmarked.

Step 1 — resume the cascade in the new, clean folder

Already done for you: 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/ exists and already contains the real, valid cascaded_checkpoint_round1.json (copied unchanged from the live run — proven valid in the incident doc, since it was measured against P0, which never changes).

Run this exact command when you're ready (identical to the live process's own command, only --output-dir changed):

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 scripts/06_measure_sensitivity_cascaded_V4.py \ /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara \ --baseline-plan 04_LANGUAGE_PLANS/V3_3_APPLES_TO_APPLES_0005/plan_Qwen38_V3_3_5.028075122774708_BPW.json \ --target-bpw 5.028075122774708 \ --optimizer-script scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \ --optimizer-extra-args "--early-qkv-min-bits 6" \ --max-rounds 3 \ --n-calibration 4 \ --calibration-mode stratified \ --candidates 4,5,6,8,16 \ --group-size 64 \ --output-dir 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected
What happens automatically, no intervention needed

Round 1: sees all 497 tensors already have a real result in the seeded checkpoint, skips straight to the optimizer — which is now fixed by default, so it writes a corrected plan_v4_round1.json (lm_head=Q6) in under a second. Zero new GPU time.

Round 2: reloads the model fresh, locks in the corrected P1's bits, and does a real, full 497-tensor cascaded measurement — this is genuine new GPU time, no way around it, because layers 1 and 3 also changed under the floor and everything downstream of them needs re-measuring against the real new context.

Round 3: same pattern again, real GPU time, using round 2's real output as context.

You can run this in parallel with the live (uncorrected) round-3 process if you want both data points — they write to different folders and don't conflict.

Step 2 — your test sequence, as you laid it out

#ActionPurpose
1Let the live (uncorrected) round 3 finishReal data point on the original, un-floored cascade
2Run the corrected cascade above (this page)Real P1/P2/P3 with lm_head genuinely protected throughout
3Pick a plan to benchmark — you're leaning P2, based on the flat-N8 precedent (P0→P1 changed a lot, P1→P2 much less, P2→P3 mostly oscillation)One real regression benchmark settles the lm_head question empirically, not just on paper
4Compare against your existing reference models on the same benchmark suiteConfirms (or corrects) everything this whole investigation predicted
5Only then: bake the final YAQA-corrected modelStage 2 sits on top of a validated Stage 1 plan, not a provisional one

Your question: GODMODE multi-bit calibration first, or build with what we have?

Direct answer

You don't need GODMODE for this specific correction. Here's why: GODMODE precomputes YAQA correction at every candidate bit-width (4,5,6,8,16) for every tensor up front, so any future plan change costs zero additional compute, ever. That's real value if you expect to keep changing individual tensors' bits repeatedly. But for what you're actually doing — one plan (P2, or whichever you pick) built once, corrected once — the existing resume-cache mechanism already gives you the same practical benefit for free (see Step 3 below). GODMODE would be solving a problem you don't currently have. Save it for if/when you're doing repeated exploratory re-bakes of many candidate plans.

Your question: does baking YAQA for a corrected plan mean starting from scratch?

No — verified directly in the real code, not assumed

05_full_model_quantize.py's resume-skip logic (read at the real line numbers, lines 733-758) checks bits, group_size, and curvature_version per tensor by name against whatever plan you point it at. Point a new build at the same --resume-dir as an existing YAQA build made from a different plan, and:

  • Any tensor whose bit-width is identical between the old and new plan is automatically skipped — its already-computed corrected weights are reused as-is.
  • Any tensor whose bit-width differs is automatically recomputed — nothing stale gets silently reused.

This matches exactly what you remembered: YAQA correction is per-tensor independent, and the Hessian itself is bit-width-independent and reusable. Given the real diff we measured — 12 of 497 tensors changed between the original and lm_head-floored P1 (2.4%) — if you already have (or build) a resume-dir cache from any closely-related prior plan, pointing a new build at that same directory should reuse roughly 95%+ of the work automatically. The one thing to check before relying on this: which plan your existing YAQA resume directories were actually built from, so you know how much real overlap to expect — I haven't verified that against your current final plan choice yet.

Author: Hakim Ghelab, VegaLaboratories LTD · Companion to IMPROVEMENT_LEDGER/04_LMHEAD_PROTECTION_INCIDENT.html · Every command on this page is the real command, copied from the real running process or verified against real code, not reconstructed from memory.