← Back to index
Real, standalone guide · 2026-09-14

Closing the Hessian coverage gap

A complete, plain-language guide — every command, every piece of every command, and exactly where each piece came from. No step assumes you already know the last one.

THE GOAL

What this whole thing is actually for

At the moment this build was launched: 360 of 497 real tensors already had real risk data, 136 didn't. This guide runs the build that closes that gap for all but one of them. Live count as this page was last corrected (2026-09-17): 391 scored, 105 still missing — the number keeps moving in the scored direction because the real build below is already running and its own live merge process (Step 3) is filling in new tensors as they finish, faster than this page can describe a fixed snapshot. Check the real, current number with the command in Step 3's closing note.

The real problem, in plain terms

Of the 497 real tensors in this model, most already have a real "Hessian danger score" — a number that says how risky it is to compress that specific tensor. At launch time, 136 tensors didn't have one — simply because, in the one real build that ever measured this, those 136 tensors were kept at full precision (16-bit), and the measurement only happens as a side effect of actually correcting a tensor, so anything left untouched never got measured. One tensor, lm_head, is never going to get this data, by a real, deliberate decision (already tried, never worked better than the simple fallback) — nothing below ever touches it.

The 3 steps below run one small, real, extra build that specifically re-measures those originally-136 tensors — nothing else — so you end up with real data for almost the whole model instead of just most of it.

Two different real files, easy to conflate — read this before Step 1

This guide talks about two separate, real data stores, not one. The flat score file (hess_scores_qwen38_v2_fp32_mtpcorrected.json) holds one real number per tensor — a single danger score, independent of bit-width. That's the file the "goal" above and the 136/105 count refer to. The godmode sensitivity checkpoint (godmode_sensitivity_checkpoint.jsonl, inside the resume folder Step 2 creates) is a completely different, richer file — real, measured rounding error at every candidate bit-width, per tensor, not one flat number. The same real build below produces both, but they are not the same thing, and they don't have the same real count. The flat file is what closes the coverage gap this guide is named after; the godmode file is the newer, more valuable output — see the closing section for where it actually goes.

BEFORE THE STEPS

Glossary — every term used below

If a word in the steps below is unfamiliar, it's defined here.

Tensor
One named block of numbers inside the model — the model is made of 497 of these. Each one gets its own real decision about how many bits to store it in.
Bits / bit-width
How much real storage space one tensor's numbers get. Lower = smaller and faster, but more real risk of losing accuracy. 16 means "don't compress it at all."
Plan
A real file (JSON) listing, for all 497 tensors, how many bits each one gets. Every real build reads one of these to know what to do.
Hessian / Hessian score
A real, measured number describing how "narrow and risky" vs. "broad and safe" a tensor's own error-sensitivity is. Lower = more dangerous to compress.
Correction
The real math step that actually reduces a tensor's compression error, done automatically during a real build — the thing that produces the Hessian score as a side effect.
Resume-dir / .yaqa_resume
A real working folder a build automatically creates next to its output, holding its logs and per-tensor progress. Step 3 reads this folder.
Godmode
A real, existing feature that measures a tensor at several bit-widths in one pass, instead of just one — used here to get richer real data per tensor.
Godmode sensitivity checkpoint
The real file godmode's multi-bit measurement writes to (godmode_sensitivity_checkpoint.jsonl, inside the resume folder). One real entry per tensor, holding a real measured error number at every candidate bit-width — not the same file, and not the same shape, as the single-number flat score file this guide's coverage gap refers to. See the callout above Step 1.
GPU
The real hardware doing the heavy computation. Only one real heavy job should run on it at a time — that's why Step 2 has to wait for your benchmark.
CHECK FIRST

Before you touch anything — check the GPU

Your benchmark (IFEval) is a real, heavy job using the same GPU. Step 1 below is 100% safe to run any time — it only edits a small JSON file, no model, no GPU. Step 2 is the one that needs the GPU free. Check with this exact command:

ps aux | grep "optiq eval" | grep -v grep

If it prints a line back, the benchmark is still running — wait. If it prints nothing, the GPU is free and Step 2 is safe to run.

STEP 1

Build the probe plan already run, confirmed working

This step does not touch the model at all. It reads a real, existing plan (a list of "this tensor gets this many bits"), and changes only the 136 missing tensors to 8-bit — a safe, moderate value whose only real job is to not be 16, so the real build below will actually process them. Every other tensor, including lm_head, is left exactly as it already was.

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara python3 scripts/yaqa_port/10_build_hessian_probe_plan.py \ scripts/yaqa_port/research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/plans_P1_both_off/plan_baseline_noHessian.json \ --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \ --probe-bits 8 \ --output 04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.json
PieceWhat it is / where it came from
cd .../qwen38-27b-heretic-araMoves into the real project's root folder — every path after this is written relative to here.
python3 10_build_hessian_probe_plan.pyA new, small script I wrote and tested today, specifically for this one job. It's real Python code, not a shortcut — I ran it and confirmed the real output.
(the plan path)The real, existing plan to start from — the one already used to build the model you're benchmarking right now (P1, no Hessian).
--hessian-score-fileThe real, existing file that already has 360 tensors' worth of real data — used to figure out which 136 are missing.
--probe-bits 8The bit-width to force the 136 missing tensors to. 8 was chosen as the safest real choice — low real risk, still enough to trigger real measurement.
--output ...Where the new, edited plan gets written. This exact file is what Step 2 reads.
Real, confirmed result

Real, confirmed result of running this exact command: 136 real tensor(s) overridden to 8-bit, file written and verified on disk.

STEP 2

Run the real build wait for the GPU check above

This is the one real, multi-hour job. It loads the real model, and for every tensor the plan above marks as non-16-bit (the 136 missing ones), it runs the real correction process — which, as a side effect, prints the real curvature data we're after. Every tensor the plan leaves at 16-bit (everything already scored, plus lm_head) is skipped entirely, exactly like every other real build you've already done.

Corrected 2026-09-15 — now a real, saved script, not a command to paste

The original single-line command on this page broke repeatedly in real terminal use — long multi-line pastes kept losing their \ line-continuations somewhere in transit, splitting one command into several and producing confusing, unrelated-looking errors. The exact same command now lives in a real file on disk, so there's nothing left to paste or retype:

/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/run_hessian_probe.sh

That's the whole command to type — no flags, no paths, no cd needed first (every path inside the script is absolute). Run it from anywhere.

What that script actually contains, in full — real, verbatim, re-read directly from the file just now, not recalled from memory. A script name alone doesn't teach you the command; this is the command, so this page stays complete even if that file ever moves, gets renamed, or is lost:

PROJECT_ROOT=/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara "$PROJECT_ROOT/scripts/yaqa_port/run_full_yaqa.sh" \ /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE \ --skip-lm-head-gptq \ --plan "$PROJECT_ROOT/04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.json" \ --n-calibration 4 \ --incoherence none \ --correct-mtp \ --godmode-multi-bit-checkpoint \ --godmode-candidate-bits 3,4,5,6,8
PieceWhat it is / where it came from
run_hessian_probe.shA real, new, saved script (scripts/yaqa_port/run_hessian_probe.sh) — every piece below is inside it, fixed and already-verified, not something you type each time.
run_full_yaqa.shInside the script: the real, existing orchestrator this project always uses for a real build — never call the underlying Python script directly (it crashes on a real memory limit if you do).
.../HESSIAN-PROBEWhere the real output of this build gets saved. A brand-new name, chosen so it can never overwrite anything real you already have. This is a scratch/data-gathering build, not something to ship.
--skip-lm-head-gptqSkips a slow, already-known-futile attempt to correct lm_head. Copied exactly from the real command that built the model you're benchmarking right now — same reason applies here.
--plan <absolute path>The exact file Step 1 created — an absolute path this time, not a relative one, so it can never break depending on which folder you happen to run from. This is exactly the piece that broke twice before this fix.
--n-calibration 4How much real calibration data to use per tensor. 4 is this project's own established real number — used in every real build so far, confirmed by checking the actual project files, not assumed.
--incoherence noneA real setting about a rotation trick before quantizing. none is this project's real, standard, production choice — copied from the same proven real command.
--correct-mtpAlso corrects the real MTP (multi-token-prediction) sidecar tensors, not just the main trunk — this model has both, and a genuine build needs both done.
--godmode-multi-bit-checkpointThe new part: turns on real measurement at every candidate bit-width for each of the 136 tensors, not just the single one the plan lists.
--godmode-candidate-bits 3,4,5,6,8Which real bit-widths to measure at. Matches exactly what you asked for earlier — every real candidate except 2 (judged too aggressive).
Do not run this until the GPU check at the top prints nothing

Running this while the benchmark is still active risks slowing or corrupting both real jobs.

STEP 3

Pull the new data out and merge it after Step 2 finishes

Step 2 writes its real findings into log files. This one step now reads those logs, pulls out the real curvature numbers, and merges them straight into your permanent file — real fix, 2026-09-16: this used to be two separate steps (extract to a temp file, then a hand-written Python merge script). Both are now one real subcommand of the consolidated toolkit, so there's no temp file and no separate merge step to forget. Full toolkit details: RUNNING_GUIDE.md §11.9.

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \ watch \ /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume \ /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \ --no-loop
PieceWhat it is / where it came from
08_extract_hessian_scores.py watchThe real, consolidated toolkit's live-merge subcommand — extracts and merges in one real pass, only ever adds new tensors, never overwrites or loses an existing real entry.
...HESSIAN-PROBE.yaqa_resumeThe exact real name from Step 2, with .yaqa_resume added — that's the real folder Step 2 writes its working logs into, automatically.
hess_scores_qwen38_v2_fp32_mtpcorrected.jsonThe real, permanent file this whole project uses — read AND written directly, no temp file in between.
--no-loopRun once and exit. Drop this flag (and add --interval 30) to keep watching and merging continuously while Step 2 is still running, instead of waiting for it to finish first.

Once this build fully completes, re-running the same command should report 496 real tensors scored — every one except lm_head, exactly as intended.

What this real data unlocks next

Closing this coverage gap is what actually makes the next real mechanism usable on (almost) every tensor, not just the 360 that already had a score: Making the Solver Listen to Curvature, Not Just KL — the real per-tensor hard floor and lexicographic Hessian-primary solve this guide's flat-score data feeds directly into.

Where the richer godmode data goes (2026-09-17)

The flat score above is only one of the two real files this build produces — see the callout near the top of this page. The richer godmode_sensitivity_checkpoint.jsonl output feeds a newer, separate real mechanism: the Hessian Hybrid Checkpoint design merges it straight into what the solver reads, tensor by tensor, falling back to cascaded KL only where this build hasn't reached yet. Where that mechanism is ultimately headed — a real, open hypothesis, not yet confirmed — is written up in Improvement Ledger §10, the pure sensitivity hypothesis.

Author: Hakim Ghelab, VegaLaboratories LTD