A complete, plain-language guide — every command, every piece of every command, and exactly where each piece came from. No step assumes you already know the last one.
At the moment this build was launched: 360 of 497 real tensors already had real risk data, 136 didn't. This guide runs the build that closes that gap for all but one of them. Live count as this page was last corrected (2026-09-17): 391 scored, 105 still missing — the number keeps moving in the scored direction because the real build below is already running and its own live merge process (Step 3) is filling in new tensors as they finish, faster than this page can describe a fixed snapshot. Check the real, current number with the command in Step 3's closing note.
Of the 497 real tensors in this model, most already have a real "Hessian danger score" —
a number that says how risky it is to compress that specific tensor. At launch time, 136
tensors didn't have one — simply because, in the one real build that ever measured this,
those 136 tensors were kept at full precision (16-bit), and the measurement only happens
as a side effect of actually correcting a tensor, so anything left untouched never got
measured. One tensor, lm_head, is never going to get this data, by a
real, deliberate decision (already tried, never worked better than the simple fallback) —
nothing below ever touches it.
The 3 steps below run one small, real, extra build that specifically re-measures those originally-136 tensors — nothing else — so you end up with real data for almost the whole model instead of just most of it.
This guide talks about two separate, real data stores, not one. The flat score file
(hess_scores_qwen38_v2_fp32_mtpcorrected.json) holds one real number per tensor —
a single danger score, independent of bit-width. That's the file the "goal" above and the
136/105 count refer to. The godmode sensitivity checkpoint
(godmode_sensitivity_checkpoint.jsonl, inside the resume folder Step 2 creates)
is a completely different, richer file — real, measured rounding error at every
candidate bit-width, per tensor, not one flat number. The same real build below produces
both, but they are not the same thing, and they don't have the same real count. The flat
file is what closes the coverage gap this guide is named after; the godmode file is the
newer, more valuable output — see the closing section for where it actually goes.
If a word in the steps below is unfamiliar, it's defined here.
.yaqa_resumegodmode_sensitivity_checkpoint.jsonl, inside the resume folder). One real entry per tensor, holding a real measured error number at every candidate bit-width — not the same file, and not the same shape, as the single-number flat score file this guide's coverage gap refers to. See the callout above Step 1.Your benchmark (IFEval) is a real, heavy job using the same GPU. Step 1 below is 100% safe to run any time — it only edits a small JSON file, no model, no GPU. Step 2 is the one that needs the GPU free. Check with this exact command:
If it prints a line back, the benchmark is still running — wait. If it prints nothing, the GPU is free and Step 2 is safe to run.
This step does not touch the model at all. It reads a real, existing plan (a list of "this
tensor gets this many bits"), and changes only the 136 missing tensors to 8-bit — a
safe, moderate value whose only real job is to not be 16, so the real build below will
actually process them. Every other tensor, including lm_head, is left exactly
as it already was.
| Piece | What it is / where it came from |
|---|---|
| cd .../qwen38-27b-heretic-ara | Moves into the real project's root folder — every path after this is written relative to here. |
| python3 10_build_hessian_probe_plan.py | A new, small script I wrote and tested today, specifically for this one job. It's real Python code, not a shortcut — I ran it and confirmed the real output. |
| (the plan path) | The real, existing plan to start from — the one already used to build the model you're benchmarking right now (P1, no Hessian). |
| --hessian-score-file | The real, existing file that already has 360 tensors' worth of real data — used to figure out which 136 are missing. |
| --probe-bits 8 | The bit-width to force the 136 missing tensors to. 8 was chosen as the safest real choice — low real risk, still enough to trigger real measurement. |
| --output ... | Where the new, edited plan gets written. This exact file is what Step 2 reads. |
Real, confirmed result of running this exact command: 136 real tensor(s) overridden to 8-bit, file written and verified on disk.
This is the one real, multi-hour job. It loads the real model, and for every tensor the
plan above marks as non-16-bit (the 136 missing ones), it runs the real correction process —
which, as a side effect, prints the real curvature data we're after. Every tensor the plan
leaves at 16-bit (everything already scored, plus lm_head) is skipped entirely,
exactly like every other real build you've already done.
The original single-line command on this page broke repeatedly in real terminal use — long
multi-line pastes kept losing their \ line-continuations somewhere in transit,
splitting one command into several and producing confusing, unrelated-looking errors. The
exact same command now lives in a real file on disk, so there's nothing left to paste or
retype:
That's the whole command to type — no flags, no paths, no cd needed first (every
path inside the script is absolute). Run it from anywhere.
What that script actually contains, in full — real, verbatim, re-read directly from the file just now, not recalled from memory. A script name alone doesn't teach you the command; this is the command, so this page stays complete even if that file ever moves, gets renamed, or is lost:
| Piece | What it is / where it came from |
|---|---|
| run_hessian_probe.sh | A real, new, saved script (scripts/yaqa_port/run_hessian_probe.sh) — every piece below is inside it, fixed and already-verified, not something you type each time. |
| run_full_yaqa.sh | Inside the script: the real, existing orchestrator this project always uses for a real build — never call the underlying Python script directly (it crashes on a real memory limit if you do). |
| .../HESSIAN-PROBE | Where the real output of this build gets saved. A brand-new name, chosen so it can never overwrite anything real you already have. This is a scratch/data-gathering build, not something to ship. |
| --skip-lm-head-gptq | Skips a slow, already-known-futile attempt to correct lm_head. Copied exactly from the real command that built the model you're benchmarking right now — same reason applies here. |
| --plan <absolute path> | The exact file Step 1 created — an absolute path this time, not a relative one, so it can never break depending on which folder you happen to run from. This is exactly the piece that broke twice before this fix. |
| --n-calibration 4 | How much real calibration data to use per tensor. 4 is this project's own established real number — used in every real build so far, confirmed by checking the actual project files, not assumed. |
| --incoherence none | A real setting about a rotation trick before quantizing. none is this project's real, standard, production choice — copied from the same proven real command. |
| --correct-mtp | Also corrects the real MTP (multi-token-prediction) sidecar tensors, not just the main trunk — this model has both, and a genuine build needs both done. |
| --godmode-multi-bit-checkpoint | The new part: turns on real measurement at every candidate bit-width for each of the 136 tensors, not just the single one the plan lists. |
| --godmode-candidate-bits 3,4,5,6,8 | Which real bit-widths to measure at. Matches exactly what you asked for earlier — every real candidate except 2 (judged too aggressive). |
Running this while the benchmark is still active risks slowing or corrupting both real jobs.
Step 2 writes its real findings into log files. This one step now reads those logs, pulls
out the real curvature numbers, and merges them straight into your permanent file — real fix, 2026-09-16:
this used to be two separate steps (extract to a temp file, then a hand-written Python merge script). Both are
now one real subcommand of the consolidated toolkit, so there's no temp file and no separate merge step to
forget. Full toolkit details: RUNNING_GUIDE.md §11.9.
| Piece | What it is / where it came from |
|---|---|
| 08_extract_hessian_scores.py watch | The real, consolidated toolkit's live-merge subcommand — extracts and merges in one real pass, only ever adds new tensors, never overwrites or loses an existing real entry. |
| ...HESSIAN-PROBE.yaqa_resume | The exact real name from Step 2, with .yaqa_resume added — that's the real folder Step 2 writes its working logs into, automatically. |
| hess_scores_qwen38_v2_fp32_mtpcorrected.json | The real, permanent file this whole project uses — read AND written directly, no temp file in between. |
| --no-loop | Run once and exit. Drop this flag (and add --interval 30) to keep watching and merging continuously while Step 2 is still running, instead of waiting for it to finish first. |
Once this build fully completes, re-running the same command should report 496 real tensors scored — every one except lm_head, exactly as intended.
Closing this coverage gap is what actually makes the next real mechanism usable on (almost) every tensor, not just the 360 that already had a score: Making the Solver Listen to Curvature, Not Just KL — the real per-tensor hard floor and lexicographic Hessian-primary solve this guide's flat-score data feeds directly into.
The flat score above is only one of the two real files this build produces —
see the callout near the top of this page. The richer godmode_sensitivity_checkpoint.jsonl output feeds a newer, separate real
mechanism: the Hessian Hybrid Checkpoint design merges it straight into what the
solver reads, tensor by tensor, falling back to cascaded KL only where this build hasn't reached yet. Where that mechanism is ultimately headed
— a real, open hypothesis, not yet confirmed — is written up in
Improvement Ledger §10, the pure sensitivity hypothesis.