← Back to index

YAQA MLX Port — Complete Running Guide

Hakim Ghelab, VegaLaboratories LTD

2026-09-05

Everything needed to run this codebase, start to finish, nothing skipped. Real path: /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/

1. Environment

PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port

Real, verified versions (requirements.txt): mlx==0.32.2, mlx-lm==0.31.3, numpy==2.4.6. Apple Silicon Mac required — MLX is Metal-based, no CPU-only or CUDA path exists.

To set up fresh elsewhere:

uv venv --python 3.11
uv pip install -r requirements.txt

2. Fast sanity check (always run this first, ~1-2 min, no real model needed)

"$PYBIN" test_regression.py

Must print ALL REGRESSION TESTS PASSED. Eleven real, independently-cross-checked guards — every bug found this session gets a permanent test so it can’t silently come back.

3. Steps 1-4 (synthetic + one real layer, not the full pipeline)

"$PYBIN" 01_gradient_collection_test.py
"$PYBIN" 02_sketch_b_synthetic_test.py
"$PYBIN" 03_rounding_update_test.py
"$PYBIN" 04_real_layer_test.py                # one real layer, end to end
"$PYBIN" 04_real_layer_test.py --from-cache   # skip model load/backward on repeat runs

4. Step 5, comparison-only (no save — just proves the correction math on real tensors)

"$PYBIN" 05_full_model_quantize.py --subset default
"$PYBIN" 05_full_model_quantize.py --subset "language_model.model.layers.63.mlp.up_proj"

5. The real full run — production, the one that actually produces a usable model

Never call 05_full_model_quantize.py directly with --output for a full run. It will hit real memory/Metal crashes documented in PORT_LEDGER.md. Always use the orchestrator:

./run_full_yaqa.sh <output_dir> [flags...]

# The real command actually used for this project's model:
./run_full_yaqa.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw \
  --n-calibration 4 --incoherence none

What run_full_yaqa.sh actually does, in order

  1. Queries the real remaining batch count (accounts for anything already done via resume).
  2. Runs one batch at a time, each as its own separate process — the real fix for a hard Metal resource-handle limit that crashes if too many batches share one process. Each batch writes its results immediately to <output_dir>.yaqa_resume/manifest.json.
  3. Re-queries the real remaining count after every batch, until it hits zero. (Real bug fixed 2026-09-05: this used to be a fixed loop count, which broke once resume-skip made the remaining count shrink mid-run — see PORT_LEDGER.md.)
  4. Runs 06_lm_head_gptq.py automatically — real, scoped one-sided GPTQ correction for lm_head, the one tensor YAQA’s two-sided correction structurally cannot reach.
  5. Runs the final assembly (--assemble-from-resume) — reads every real result from the resume cache and writes the actual, complete, loadable model to <output_dir>.

Real flags that matter

Flag What it does Real effect
--source <path> BF16 source model to quantize (was hardcoded, now a real override) Must be consistent across every step of one run
--plan <path> Real per-tensor bit-allocation plan JSON (was hardcoded, now a real override) A different --source almost certainly needs its own real plan too
--n-calibration N Real stratified sequences per domain (6 domains, so total = 6×N) More wall-clock time, NOT more peak memory (see below)
--incoherence {none,hadamard} none = production, real mlx_lm-loadable today. hadamard = research branch, needs an unbuilt loader none is the real default; don’t use hadamard unless you’re building the loader too
--hessian-batch-budget-gb Real memory budget per batch (default 8GB) Lower = more batches, less peak memory per batch; raise only if you have real headroom
--lm-head-name <name> Real dotted name of the vocabulary-output tensor Auto-detected by default — you never need to pass this. Two real, cross-checked signals: the model’s real vocab size (found anywhere in its config) matched against which plan tensor’s real output shape equals it, cross-checked against the .lm_head naming convention. Only pass this explicitly if it raises an error (shows exactly what each signal found)
--correct-mtp Real YAQA correction for the MTP speculative-decode sidecar (added 2026-09-08) Off by default — the sidecar is otherwise naive-quantized. See section 6, Scenario C below
--reuse-cached-fallback Skip lm_head’s damping-search retry if a prior run already recorded naive_fallback (added 2026-09-08) Off by default — see section 9 below
--skip-lm-head-gptq Never call 06_lm_head_gptq.py at all for this run — lm_head goes straight to Part 5c’s plain native mx.quantize fallback (added 2026-09-13, run_full_yaqa.sh only) Off by default — see section 9 below for exactly when to use this vs. --reuse-cached-fallback vs. neither

Full list of every real flag on both scripts, cross-checked against --help: YAQA_COMMAND_REFERENCE.md’s “Every real flag” section — not duplicated here.

Calibration, precisely — is it built in, or passed?

Passed, not built in. --n-calibration sets k_per_domain for the real stratified calibration loader (6 domains: agent, code, instruct, prose, thought, tool). At the default of 4, that’s 24 total real sequences (N=24) — matching this project’s own established convention (same as 08_gptq_apply_plan.py’s own default).

Does it affect memory? Not the way you’d think. H_I/H_O are fixed-size matrices (in_features²/out_features²) no matter how much calibration data is used — more calibration only adds more terms into the same-sized running sum, it never grows the matrix. What actually changes is how many real forward+backward chunks get run (--batch-size, default 1 sequence at a time) — so --n-calibration affects real wall-clock time (linearly), not peak memory. The real memory pressure seen during this project’s own full run came from which tensors happened to be in a batch (the large 17,408-dimension MLP tensors), not from the calibration count.

6. Common scenarios, worked end to end

Real situations, each a complete, copy-pasteable example — the difference between them is entirely about whether <output_dir>.yaqa_resume already exists, what’s in it, and which real plan you’re building from (this list has grown as real, new situations came up — current count: five, A through E).

Scenario A — brand new model, nothing exists yet

Nothing to reuse: no output directory, no resume cache. This is the plain, default path.

NEW_OUTPUT="/Users/hghelab/.mtplx/models/My-New-Model-YAQA-5bpw"
./run_full_yaqa.sh "$NEW_OUTPUT" --n-calibration 4 --incoherence none

run_full_yaqa.sh creates ${NEW_OUTPUT}.yaqa_resume fresh, runs every real batch from zero (see section 5’s “what it actually does” above), then lm_head, then assembles. No special flags needed — this is the every-tensor-from-scratch case.

Scenario B — resuming or reusing an EXISTING cache (same output)

A run was interrupted (crash, Ctrl+C, machine slept), or you just want to change a flag like --hessian-batch-budget-gb and continue. Run the exact same command again, unchanged (same <output_dir>):

./run_full_yaqa.sh "$NEW_OUTPUT" --n-calibration 4 --incoherence none

Every tensor already completed (checked by name against the real manifest in ${NEW_OUTPUT}.yaqa_resume/manifest.json, matching bits/group_size) is skipped automatically and picked back up where it left off — this works even if --hessian-batch-budget-gb changes between runs, since the skip check is per-tensor, not per-batch-index. Nothing needs to be copied or renamed for this case — that’s only for Scenario C below, which deliberately reuses a cache under a different name.

Scenario C — adding YAQA-corrected MTP to an ALREADY-COMPLETED build

Different from Scenario B: here the trunk build already finished successfully (a real, working model exists), and you want to add the MTP sidecar correction (--correct-mtp, see the flags table in section 5) without recomputing the whole multi-hour trunk correction. This reuses the existing resume cache under a new name, so the original completed model is never touched or put at risk.

# Step 1 -- copy the existing resume cache under a NEW output name (non-destructive)
OLD="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32"
NEW="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected"
cp -a "${OLD}.yaqa_resume" "${NEW}.yaqa_resume"

# Step 2 -- launch: skips every already-cached trunk/lm_head tensor, goes straight to
# reassembly + the new MTP correction
./run_full_yaqa.sh "$NEW" --n-calibration 4 --incoherence none --correct-mtp --reuse-cached-fallback

# Step 3 -- watch it live, in a second terminal tab
./monitor.sh "${NEW}.yaqa_resume"

Real pitfall — only run Step 1 once. cp -a src dst copies src’s contents into a new dst only if dst doesn’t already exist; if dst already exists (e.g. Step 1 was already run once), it instead nests src inside dst, silently doubling real disk usage with no error (confirmed: 16GB → 32GB). Check ls -d "${NEW}.yaqa_resume" first — if it already exists, skip straight to Step 2, or rm -rf it before re-copying.

Full flag-by-flag rationale (why --reuse-cached-fallback exists, what --mtp-bits defaults to, etc.): YAQA_COMMAND_REFERENCE.md’s dedicated MTP section — not duplicated here, this is the same real 3-step procedure either way.

Scenario D — godmode: a multi-bit rich sensitivity checkpoint (new, 2026-09-11)

Different question from Scenarios A-C: those build a real model at the plan’s one assigned bit per tensor. Godmode instead asks, for every tensor, “what would YAQA’s real corrected error be at every candidate bit-width” — 2, 3, 4, 5, 6, and 8-bit by default, not just the one the plan picked. Reuses the same H_I/H_O Part 5a already computes per tensor (confirmed bit-width-independent) — no extra forward/backward calibration passes, only extra (cheap) correction/safety-gate calls. Does not touch or replace the existing single-bit build — it’s a pure addition, on by explicit flag only.

# Fresh checkpoint, from scratch, default candidate bits (2,3,4,5,6,8), no 16-bit
# (16-bit is the BF16 reference itself -- zero rounding error by construction, already
# known, not worth spending real time re-measuring)
GODMODE="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-GODMODE"
"$PYBIN" 05_full_model_quantize.py \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara \
  --plan 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE/plan_v4_round2.json \
  --resume-dir "${GODMODE}.yaqa_resume" \
  --godmode-multi-bit-checkpoint \
  --n-calibration 4 --incoherence none

Stop: Ctrl+C, or kill the process — same safety property as the main run. Each real tensor’s godmode sweep writes to ${GODMODE}.yaqa_resume/godmode_sensitivity_checkpoint.jsonl immediately after that tensor’s last candidate bit finishes (one JSON object per line, appended) — a kill loses at most the tensor currently in flight, never anything already written.

Resume: run the exact same command again, unchanged. On startup it reads the existing .jsonl, builds the set of tensor names already present, and skips them — matching this file’s own resume-cache convention, not a new mechanism.

Real cost, verified, not estimated: the existing single-bit trunk correction took 27.56 real hours for all 362 tensors (measured directly from real batch_*.log file timestamps, 2026-09-11) — that’s the real floor, since Part 5a’s expensive calibration pass is unchanged. Godmode’s own added overhead (looping 6 candidates’ worth of cheap correction calls instead of 1, per tensor) has not been measured yet — the first real run is the first real measurement of it, not a promise made in advance.

Override the candidate set with --godmode-candidate-bits "4,5,6" (comma-separated, any subset); override the output path with --godmode-checkpoint-path.

Scenario E — building from the lm_head-protected stratified cascade (new, 2026-09-13)

Real background, not a hypothetical: the live stratified-N24 cascade (V4_CLEAN_STRATIFIED_CASCADE) ran its full three rounds and finished (497/497 tensors measured in round 3, no crash, real plan written). Separately, a real bug was found and fixed during that same investigation: language_model.lm_head — 1.27 billion parameters, the single largest tensor in the model — was landing at Q4 in every real cascaded round (round 1 and round 2 both), despite carrying a 100x structural_weight KL-protection multiplier in 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py. Root cause, verified directly: cascaded measurement at lm_head (measured last in causal order, on top of an already-degraded upstream context) is nearly flat across candidate bits (0.2326 at 4-bit vs. 0.2289 at 8-bit — a 1.6% range), so a mathematically correct MILP finds it “cheap” to spend that tensor’s budget elsewhere even with the 100x weight applied. Fixed by adding a real, hard minimum-bits floor (--lm-head-min-bits, default 6, on by default) to the optimizer script itself — full trace, real numbers, and independent confirmation (the vendor optiq package’s own unmodified allocator never puts lm_head below Q6 on the same real data) in research_hadamard_blowup/IMPROVEMENT_LEDGER/04_LMHEAD_PROTECTION_INCIDENT.html.

Re-optimizing round 1’s real, already-measured, still-valid checkpoint (valid because it was measured against P0, which never changes) through the now-fixed optimizer produced a corrected P1 with lm_head=6 and only 11 other tensors shifted (12 total, 2.4% of 497) — real, exact diff in the incident doc above. Rounds 2 and 3 were not re-measured with the floor (that needs real new GPU time, a deliberate decision, not an oversight — see IMPROVEMENT_LEDGER/05_LMHEAD_FIX_NEXT_STEPS.html), so this scenario builds from the corrected P1 specifically.

Why Step 1 below matters — this is not optional. run_full_yaqa.sh always derives its resume folder as "${OUTPUT_DIR}.yaqa_resume", straight from the --output name — there is no flag to point it at a different, pre-existing resume-dir. The uncorrected P1 build already computed YAQA correction for 363 real tensors under Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume. Only 12 of 497 tensors differ between the uncorrected and corrected P1 plans (2.4%), and YAQA’s per-tensor correction is independent of which plan produced the target bit-width (05_full_model_quantize.py’s resume-skip check, real line numbers ~733-758, matches on bits/group_size/ curvature_version per tensor name) — so almost all of that prior work is reusable. But a brand-new --output name means a brand-new, empty .yaqa_resume folder with nothing in it to reuse, unless the old one is physically copied to the exact path the new run will look for first.

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port

# Step 1 -- seed the new resume-dir with the existing, closely-related cache (non-destructive:
# copies TO a new path, never touches the original). Check real disk headroom first (df -h) --
# this duplicates the full 16GB cache size before anything can be deduplicated away.
SRC=~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume
DST=~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6.yaqa_resume
cp -R "$SRC" "$DST"

# Step 2 -- launch. Real, verified result of Step 1 on this project's own data: 236 of 343
# real targets resume-skip (already match), leaving only 107 tensors (8 real batches) to
# actually recompute -- checked directly via --print-num-batches before committing to the run.
./run_full_yaqa.sh \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6 \
  --skip-lm-head-gptq \
  --plan /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/plan_v4_round1.json \
  --n-calibration 4 \
  --incoherence none \
  --correct-mtp

# Step 3 -- watch it live, in a second terminal tab
./monitor.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6.yaqa_resume

Real pitfall — only run Step 1 once, same as Scenario C: cp -R src dst copies src’s contents into a new dst only if dst doesn’t already exist; if dst already exists (Step 1 already ran), it nests src inside dst instead, silently doubling real disk usage with no error. Check ls -d "$DST" first — if it already exists, skip straight to Step 2, or rm -rf it before re-copying.

Why --skip-lm-head-gptq specifically, here — not --reuse-cached-fallback, not neither: this command’s --output name is new, so <output>.yaqa_resume is a brand-new folder (even seeded via Step 1, nothing in it is tagged for lm_head specifically — the old cache never ran GPTQ against this exact plan). There is nothing cached for lm_head to reuse. --reuse-cached-fallback only skips a repeat of the damping search when a matching prior attempt is sitting in the same resume-dir already — pointed at a folder with nothing cached for this tensor, it finds nothing and the real, full, ~15-20-minute damping search would run anyway, for a tensor already proven (real manifest evidence, both the old build and this project’s own real testing) to always end in naive_fallback regardless. --skip-lm-head-gptq is the only one of the three real options that guarantees zero wasted time here, because it never calls 06_lm_head_gptq.py at all — see section 9 below for the complete decision guide covering all three real situations, spelled out, not summarized.

What every unmentioned flag actually defaults to

None of the three examples above pass every real flag — every flag not mentioned silently uses its own real default, and those defaults are never “off”/“nothing”, they are specific, real values:

The complete, exhaustive list (every flag on both scripts, cross-checked directly against real --help output): YAQA_COMMAND_REFERENCE.md’s “Every real flag” section.

7. Watching a run live

Recommended (added 2026-09-16): just run watch.sh with no arguments — it lists every real *.yaqa_resume job it finds under /Users/hghelab/.mtplx/models/, numbered, and prompts for which one to watch. No path to remember, no hardcoded wrapper script per job (that pattern — one watch_<job>.sh per job — was tried and replaced; it doesn’t scale past a couple of jobs).

./watch.sh

Or skip the prompt if you already know which job:

./watch.sh HESSIAN-PROBE        # by name fragment
./watch.sh 2                    # by number from the list

Under the hood this still calls monitor.sh (see below) — watch.sh is just the real, generic entry point so you never have to construct that path by hand.

monitor.sh directly, if you already have the exact path

./monitor.sh <output_dir>.yaqa_resume/logs

Point it at the logs folder (not a specific file) — it auto-detects and follows whichever batch is currently live, switching automatically as the run progresses. Color-coded: green = running, red = stopped/crashed, yellow = memory getting low. Shows real ps stats, real system memory/swap, and the last 15 real log lines.

Live visual dashboard (real, auto-generated, added 2026-09-05)

Every time monitor.sh refreshes (every 4s), it also regenerates a real, single-file HTML dashboard at:

<output_dir>.yaqa_resume/dashboard.html

Open that file once in a browser — it auto-refreshes itself every 8s via a meta-refresh tag, so it stays live without you re-opening it. Real example (this project’s own model):

file:///Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw.yaqa_resume/dashboard.html

Everything on it is computed from real data each refresh — no simulated numbers:

Generated by generate_dashboard.py (called automatically from inside monitor.sh – you never run it directly). Reuses the exact same real interpretation logic as interpret_manifest.py (section 8 below), just rendered visually instead of as text. As of 2026-09-08, monitor.sh also serves two live, no-flicker versions over a local HTTP server (file:// pages can’t auto-refresh via fetch() in Chrome): the same internal dashboard, and a separate customer-facing view — both show the same real milestones, including a dedicated YAQA-correcting MTP sidecar step when --correct-mtp is in use.

Real fix, 2026-09-16 — the ports are no longer fixed. They used to always be 8765/8766, which meant two real jobs being watched at the same time fought over the same two ports and kept silently swapping which job’s dashboard you were actually looking at (caught live, mid-session, from the real self-heal log line). The internal dashboard’s port is now derived from a hash of the specific job’s own resume-dir path, so two jobs never collide — monitor.sh prints the real, resolved port for the job you’re watching on every refresh cycle; read it from there rather than assuming 8765. The customer-facing port stays fixed at 8766 (it always serves the same shared brand/ folder regardless of job, so there’s no real collision to avoid there), but the file it writes is now named per-job (live-status-<job-name>.html) instead of a single shared file every job used to overwrite.

8. Checking real progress at any time (while a run is in progress, or after)

"$PYBIN" interpret_manifest.py <output_dir>.yaqa_resume

Real example (this project’s own model):

"$PYBIN" interpret_manifest.py /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw.yaqa_resume

Explains Metric A (err_yaqa/err_naive — real Hessian-weighted reconstruction error, not KL divergence) vs. Metric B (plain Frobenius distance) in plain terms, and breaks real progress down by tensor kind and bit-width.

9. lm_head — what it gets, why it’s separate, and exactly which flag to use

What lm_head is and why it needs its own path

lm_head is the model’s final output-vocabulary projection (1.27B parameters). YAQA’s normal two-sided correction cannot reach it: its output-side Hessian alone (tied to the 248,320-word vocabulary) needs 246.7GB — a real memory limit, not a tunable one. 06_lm_head_gptq.py exists to give it something better than nothing: a real, scoped, one-sided GPTQ correction instead (GPTQ only ever needs the small input-side Hessian, which is affordable). This is a genuine, intended, standard part of the pipeline — run_full_yaqa.sh calls it automatically, every time, right after all the other batches finish — not an optional extra step you’re expected to skip by default.

If GPTQ’s real search finds a damping ratio that passes its safety gate, the manifest entry is tagged "method": "gptq". If every real damping ratio it tries fails that gate, it falls back and tags the entry "method": "naive_fallback" — plain native mx.quantize, the same thing 05_full_model_quantize.py’s own Part 5c would have done for lm_head anyway if GPTQ were skipped entirely. Either way lm_head still ends up quantized at the plan’s own bits/group_size (6-bit/G64 for this project’s corrected plans) — it is never left at BF16. Only the correction method differs, not the bit-width.

On this project’s own real tensor, GPTQ’s search has been run for real, multiple times, and has recorded naive_fallback every single time — every real damping ratio it tried failed the safety gate. That is not a hypothesis; it is the observed, repeated outcome on this specific tensor with this specific Hessian.

Three real situations, and which flag (if any) fits each one

Situation A — you’re building a genuinely new corrected plan for the first time, and lm_head’s known-futile outcome is not going to change. This is the common case: a new plan, a brand-new --output name, therefore a brand-new, empty <output>.yaqa_resume folder. Nothing is cached in that folder yet for anything, lm_head included. In this situation, letting 06_lm_head_gptq.py run means it re-derives the exact same naive_fallback result it always gets on this tensor — burning ~15-20 real minutes to arrive at a result you already know in advance. Use --skip-lm-head-gptq. This skips calling 06_lm_head_gptq.py at all; lm_head still gets set correctly, at the plan’s real bits, via 05_full_model_quantize.py’s own built-in Part 5c native-quantize fallback — numerically identical to what GPTQ’s naive_fallback path would have produced anyway, just without wasting the real time re-deriving it.

Situation B — you’re re-running against a resume-dir that genuinely already has a cached lm_head GPTQ attempt from a previous run against the exact same tensor and Hessian (for example: a prior run on this same output crashed partway through, after lm_head’s GPTQ step had already completed and been written to the manifest, and you’re resuming it). Here, and only here, --reuse-cached-fallback does real work: it checks the manifest for a matching prior lm_head entry and, if found, skips straight past re-running the damping-search loop instead of repeating it. Use --reuse-cached-fallback in this situation specifically. Passing it against an empty resume-dir (Situation A) is not wrong, exactly, but it is a no-op — there is nothing cached yet for it to find, so 06_lm_head_gptq.py still runs its full real search from scratch regardless of the flag.

Situation C — you have a specific, real reason to believe GPTQ might behave differently this time (a materially different plan’s lm_head bit-width, a changed calibration set, a code change to the damping search itself) and you want a genuine, un-skipped attempt. Pass neither flag. 06_lm_head_gptq.py runs its real search fresh, with no shortcut and no cached result to fall back on, and whatever it finds (gptq or naive_fallback) is recorded as a new, real manifest entry.

The three flags are not interchangeable and do not stack usefully: --skip-lm-head-gptq prevents 06_lm_head_gptq.py from running at all, so --reuse-cached-fallback has nothing left to apply to if both are passed together — pick the one that matches your actual situation, not both.

Reusing prior YAQA work when the resume-dir would otherwise be empty

run_full_yaqa.sh always derives its resume folder as "${OUTPUT_DIR}.yaqa_resume" — directly from whatever --output/first-positional-arg name you give it. There is no flag to point it at a different existing resume-dir. This matters because YAQA’s per-tensor correction is independent of which plan produced the target bit-width (checked by 05_full_model_quantize.py‘s resume-skip logic, real line numbers ~733-758: it matches on bits, group_size, and curvature_version per tensor name) — so a new plan that only changes a handful of tensors’ bits (this project’s real lm_head-floor fix changed 12 of 497, 2.4%) should be able to reuse almost everything already computed under a previous, closely-related plan. But the orchestrator has no built-in way to know that a prior resume-dir it should look at even exists — a brand-new output name always starts from a brand-new, empty .yaqa_resume folder, full stop.

The real fix, done the same way this project already handles reusing a prior cascade checkpoint: physically copy the old resume-dir to the exact path the new run will look for, before running it.

SRC=~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume
DST=~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6.yaqa_resume
cp -R "$SRC" "$DST"

After the copy, running 05_full_model_quantize.py --print-num-batches (or just letting run_full_yaqa.sh start normally) against the new plan will report real resume-skip hits for every tensor whose bits/group_size/curvature_version match — verified directly on this project’s own real P1-vs-P1-floored plan pair: 236 of 343 real targets matched and were skipped, leaving only 107 tensors (8 real batches) actually needing recomputation. Check real disk headroom first (df -h) — the copy duplicates the full resume-dir size (16GB in this project’s case) before any of it can be deduplicated away.

10. After the run finishes

The real output model is a complete, standard directory at <output_dir> — safetensors shards, config.json, vision/audio sidecars, MTP head, all reattached automatically via the same real mechanism 08_gptq_apply_plan.py already uses in production. Loadable with plain mlx_lm.load(). Verify it:

"$PYBIN" -c "
from mlx_lm.utils import load
import mlx.core as mx
model, tok, cfg = load('<output_dir>')
ids = tok.encode('The capital of France is')
logits = model(mx.array([ids]))
mx.eval(logits)
print(tok.decode([int(mx.argmax(logits[0, -1]))]))
"

Should print a coherent next token. This exact check has been run and passed on this project’s own real saved output already.

11. The Hessian-aware hybrid solver (new, 2026-09-14) — exact mechanics, exact commands

This is a real, separate addition to 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py — the MILP that turns real cascaded sensitivity data into a bits-per-tensor plan. It answers a real, specific question: the solver’s own protection for “important” tensors (lm_head, first/last-layer boundary tensors) was, until now, a single flat 100× KL-weight multiplier applied to a hand-picked category — a real, working patch, but blind to whether any given tensor’s own real curvature actually justified that multiplier. This section documents the real fix: a second, genuinely independent signal (real Hessian geometry, not another KL measurement) that can replace the flat multiplier with a continuous, per-tensor one.

11.1 The two real flags, exactly as implemented

--hessian-score-file HESSIAN_SCORE_FILE
    Optional path to a real JSON {tensor_name: hess_score} map (e.g. exported from
    hessian_story_lib.load_hessian()). When given, multiplies each named tensor's
    structural_weight by a real, bounded factor derived from its score's percentile
    rank among all scored tensors -- the most dangerous tensor (lowest score) gets
    the largest boost, the safest (highest score) gets none. Tensors with no real
    score in the file are left at factor 1.0 (unchanged). Default: no file, no effect.

--hessian-weight-strength HESSIAN_WEIGHT_STRENGTH
    Real multiplier range applied via --hessian-score-file: the single most
    dangerous tensor's structural_weight is scaled by (1 + this), the safest by
    1.0, linearly interpolated by real percentile rank in between. Only read when
    --hessian-score-file is given. (default: 2.0)

Both are off by default — omit --hessian-score-file entirely and the solver’s output is byte-identical to before this feature existed (verified directly: same real plan, zero diff, with the flag omitted).

11.2 Where the real score comes from, and the exact command to (re)generate it

hess_score(tensor) = mean(effective_rank(H_I)/dim_in, effective_rank(H_O)/dim_out), where effective_rank(H) = trace(H)² / Σ(H²) — a real spectral participation-ratio diagnostic computed from each tensor’s own real H_I/H_O (the same Hessians YAQA’s own two-sided correction already builds; nothing new is measured). Low score = the Hessian’s energy is concentrated in very few directions (narrow, dangerous). High score = spread broadly (forgiving). This is parsed directly from the real "... real effective rank -- H_I=.../..., H_O=.../..." log lines every real YAQA build batch already prints — hessian_story_lib.load_hessian() already does this parsing; nothing new to write.

Real automation (added 2026-09-14, direct user request — “we might need to streamline an orchestrator so the Hessian score is always there”): the raw per-tensor data was already being printed for free by every real correction pass — the only real gap was that nobody ever consolidated it automatically. That gap is now closed:

Exact real command (this is what run_full_yaqa.sh now runs automatically, and what you’d run by hand against any older, already-completed build):

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  extract \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume \
  --output /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json

Real fix, 2026-09-16: this script is now the consolidated toolkit (§11.9) — the extract subcommand above is required (the bare form without it now errors).

The real, full lifecycle of this exact command — same command, two different moments, two different answers:

Moment Is this command correct? Why
Day zero — no master score file exists yet Yes — this exact command This is precisely what creates the file for the first time, from one finished build’s real logs.
Any time after — the file already has real data (e.g. from a later watch run against a different build) No — do not re-run this extract reads only the one resume-dir you point it at and overwrites the output. Real tensors from any other build (only watch adds those) aren’t in what it just read, so they’re gone from what it just wrote. Use watch instead (§11.9) — it reads the existing file first and only ever adds.

This project’s own score file already had its real day zero before this section was written — so today, the correct command for it is always watch, never this one.

Verified byte-for-byte identical to the original manual extraction this session. Real, already-generated output, kept permanently in the repo (not a scratch file): research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json — 360 real tensors scored (every real YAQA-corrected trunk tensor from that build; lm_head and any bits=16 tensor are absent from this file by real construction, not by omission — see 11.4). Real scope note: this score is a property of the model’s own real weights and calibration data, not of which plan built it — one real extraction is reusable for any future plan against the same model, it does not need to be regenerated per plan.

11.3 The four real experiments this session, exact commands, exact results

Honesty note first, direct user question, direct answer: these four runs are NOT a systematic strength sweep (never ran strength = 1, 3, 4, 5, 6, 7 — only 2.0 and 8.0 were ever actually run). 2.0 is the flag’s own coded default (--hessian-weight-strength, default: 2.0) — a reasonable starting value, never empirically tuned or optimized against any real eval. 8.0 was one deliberate second data point, chosen to be “notably higher” specifically to see how the effect scales — not a search. Real summary table, all four:

# Real config (beyond target-bpw/group-size/candidates, held constant) Strength Compared against Real result
A --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 (no Hessian file at all) — real shipped production plan 69/497 differ. lm_head self-promotes 6→16, zero Hessian involved.
B same floors as A + --hessian-score-file ... --hessian-weight-strength 2.0 2.0 plan A Only 3/497 differ from A. lm_head stays at 16 (unchanged from A). Real, meaningful move: layers.55.self_attn.k_proj: 16→4 — matches the original role-based “safe to compress” recommendation cascaded-KL alone had refused.
C same floors as A + --hessian-score-file ... --hessian-weight-strength 8.0 8.0 plan A 40/497 differ from A. Broad promotion of layers 34–62 to full precision. lm_head drops back down from 16 to exactly 6 — cranking the signal too far dilutes its own relative specialness.
D --late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 --no-protect-first-last + --hessian-score-file ... --hessian-weight-strength 2.0 2.0 real shipped production plan 34/497 differ. lm_head self-promotes again, 6→16, this time with real regional floors still active.

Experiment A — pure cascaded KL, no floors at all, no Hessian signal. The real zero-Hessian baseline:

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \
  --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \
  --output 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/plan_v4_round1_NOFLOORS.json

Experiment B — add the real Hessian signal at the default strength (2.0), floors still off:

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \
  --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 2.0 \
  --output /tmp/plan_v4_round1_NOFLOORS_HYBRID.json

Experiment C — same, but strength cranked to 8.0 (the real cautionary result):

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \
  --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 8.0 \
  --output /tmp/plan_v4_round1_NOFLOORS_HYBRID_STRONG.json

Experiment D — the actual “hybrid” configuration: production floors kept ON, only the flat 100× boundary category replaced by the real Hessian score:

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \
  --late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 \
  --no-protect-first-last \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 2.0 \
  --output 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/plan_v4_round1_HESSIAN_HYBRID.json

--no-protect-first-last is the real, pre-existing flag that disables the flat 100× first/ last-layer boundary category (identify_boundary_layers()); --hessian-score-file supplies the real, continuous replacement. --late-attn-min-bits 5 --early-qkv-min-bits 6 are left exactly at the real production values — not disabled here, see 11.4 for why.

Real result: 34 of 497 tensors differ from the real shipped production plan. L0/L63 MLP tensors (previously locked at 16-bit purely by the flat boundary rule) drop to their real curvature-justified bits (mostly 4-bit); lm_head again self-promotes, this time all the way to 16-bit; several mid-depth self_attn.k_proj/v_proj tensors that the boundary rule never touched settle exactly at the real late-attention floor (5-bit) instead of being accidentally over-protected. Same real BPW budget throughout — a reallocation, not “spend more.”

11.4 The direct answer: should the regional floors stay ON or come OFF?

Keep --late-attn-min-bits and --early-qkv-min-bits at their real production defaults (5 and 6). Do not disable them. Real reasoning, not a guess:

Concrete, current recommendation: run production builds with the region floors ON (unchanged), --lm-head-min-bits 6 ON (unchanged), and --no-protect-first-last --hessian-score-file ... --hessian-weight-strength 2.0 as the one real, justified substitution — replacing only the boundary category’s flat 100×, nothing else. This has not yet been used to build a real shipped model; do that (and run the real IFEval/MMLU regression checks against it) before making it any script’s own default.

11.5 The exact granularity mechanism — real worked example, not the abstract version

The real formula, in three steps, exactly as implemented in apply_hessian_weighting():

  1. Rank all 360 real scored tensors by their real hess_score, most dangerous (lowest score) to safest (highest score).
  2. Convert each tensor’s position in that ranking into a percentile r from 0 (most dangerous) to 1 (safest): r = rank_index / (360 - 1).
  3. Compute its real weight: factor = 1.0 + strength × (1 − r).

That is the entire mechanism. Real numbers, pulled directly from the real 360-tensor score file, at strength=2.0:

Real tensor Real hess_score Percentile rank r Real weight factor
layers.63.mlp.up_proj (most dangerous of all 360) 0.000161 0.000 3.000×
layers.25.mlp.up_proj 0.000499 0.251 2.499×
layers.15.self_attn.o_proj (dead middle) 0.001130 0.501 1.997×
layers.36.mlp.up_proj 0.002000 0.752 1.496×
layers.55.self_attn.k_proj (safest of all 360) 0.012979 1.000 1.000×

Five genuinely different real weights, out of 360 total — every scored tensor gets its own precise number. The old flat system would have given all five of these 1.0 (none of them are named lm_head or a boundary layer), blind to the fact that layers.63.mlp.up_proj is objectively far more dangerous than layers.36.mlp.up_proj by real measured curvature.

11.6 Naive 100× vs. real Hessian score — the exact mechanical difference

Flat 100× category weight Real Hessian percentile weight
What decides the multiplier A binary category match (name pattern / depth position) This tensor’s own real measured curvature
Granularity Same multiplier for every tensor in the category Continuous, unique per tensor
Can it protect a tensor the category misses? No — if it’s not in the category, it gets 1.0× regardless of real risk Yes — any of the 360 real scored tensors, category or not
Can it over-protect a tensor that doesn’t need it? Yes, systematically — every category member gets full weight even if genuinely safe (real proof: L0/L63 MLP tensors sat at 16-bit for no curvature reason) No — the safest real tensor gets exactly 1.0×, no boost at all
Works for lm_head? Yes, but insufficient alone (real, measured: still lands at Q4 despite 100×) No real data exists — structurally exempt from this whole mechanism
Real, current status Still governs the region floors’ own logic (untouched) Opt-in replacement, boundary category only, not yet a shipped default

11.7 Closing the Hessian coverage gap (2026-09-14) — the exact real commands

The real gap, confirmed by direct audit: of the 497 real solver targets, only 360 have a real Hessian score. 137 do not — 136 of those (lm_head is the 137th, and stays permanently excluded, see below) simply because the build that produced hess_scores.json held them at bits=16 in its own plan, so YAQA’s real correction pass — the thing that prints the effective-rank data this whole mechanism is built from — never ran on them. This is a real, structural chicken-and-egg problem: the tensors an old protection rule kept at full precision are exactly the tensors with zero real data on whether that protection was ever justified.

Real breakdown of the 136 missing (excluding lm_head), by tensor family:

Family Count
in_proj_a 48
in_proj_b 48
v_proj 10
k_proj 9
out_proj 6
in_proj_qkv 4
in_proj_z 4
down_proj 2
up_proj 2
o_proj 2
gate_proj 1

Why lm_head stays excluded, permanently, by direct decision (not a gap to fix): a real one-sided GPTQ correction was already tried on it — it never rounded better than the naive fallback, needed heavy real padding, and lm_head is already known-sensitive from the literature. The real decision: keep it at its naive floor (6-bit) or BF16, never attempt YAQA/GPTQ correction on it. This script never touches lm_head, under any circumstance.

Why a real build is needed even though the Hessian math itself doesn’t depend on bit-width: H_I/H_O are built from real calibration activations, computed before any rounding decision — they don’t know or care what bit-width you’re about to pick. But 05_full_model_quantize.py’s own real tensor-selection gate, load_plan_targets(), skips any tensor a plan marks bits=16 — real correction (and the effective-rank print that is its free byproduct) never runs on it. That’s a real, deliberate pipeline shortcut (why spend real GPU time on a tensor you’re not touching), not a mathematical requirement — but it’s the gate that exists, so it has to be satisfied. Confirmed directly: even --godmode-multi-bit-checkpoint (the real, already-existing multi-bit sweep that reuses one real collected Hessian across every candidate in {2,3,4,5,6,8} without re-running calibration) inherits this same gate — it does not bypass it.

Step 1 — build a real “probe” plan. New, tested script: scripts/yaqa_port/10_build_hessian_probe_plan.py. Takes a real, already-built plan and forces only the currently-unscored, non-lm_head tensors to a single real candidate bit-width (default 8 — the safest real choice; its only job is to be != 16 so the tensor clears the pipeline’s selection gate). Every other tensor, including lm_head, is left completely untouched.

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/yaqa_port/10_build_hessian_probe_plan.py \
    scripts/yaqa_port/research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/plans_P1_both_off/plan_baseline_noHessian.json \
    --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
    --probe-bits 8 \
    --output 04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.json

Real, verified output (ran this exact command): 136 real tensor(s) overridden to 8-bit (everything else, including lm_head, left untouched).

Step 2 — run the real build with godmode multi-bit sweeping enabled, so every probed tensor’s real effective rank is captured at every real candidate in {3,4,5,6,8} in one real pass (not just at the single forced probe bit):

Corrected 2026-09-15 — this originally omitted the required cd and used a plan path only correct from the project root; ./run_full_yaqa.sh only runs from inside scripts/yaqa_port/, two real directories deeper than the project root. That relative-path version then broke a second way in real use: pasted as a multi-line command, the \ line-continuations kept getting lost in transit, silently splitting one command into several. Real, permanent fix: a saved script with every path absolute, nothing left to paste or retype — scripts/yaqa_port/run_hessian_probe.sh:

/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/run_hessian_probe.sh

That’s the whole command — no flags, no cd first, run from anywhere. Its real, fixed contents (everything below already inside the script):

/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/run_full_yaqa.sh \
    /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE \
    --skip-lm-head-gptq \
    --plan /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.json \
    --n-calibration 4 \
    --incoherence none \
    --correct-mtp \
    --godmode-multi-bit-checkpoint \
    --godmode-candidate-bits 3,4,5,6,8

This is a real, multi-hour GPU job — only 136 of the 497 real tensors need correction here (everything already scored keeps its cached real result if run against the same resume-dir lineage; confirm real resume-dir semantics before assuming full reuse), but real godmode sweeping across 5 candidates per tensor is real, additional compute per tensor, not free.

Real, verified fix required before this step works (2026-09-14, found and fixed the same day): godmode’s own per-candidate correction calls originally passed verbose=False (real code, 05_full_model_quantize.py), which silently suppressed the effective-rank print this whole mechanism depends on — its richer data only ever reached the separate godmode_sensitivity_checkpoint.jsonl file, which 08_extract_hessian_scores.py never reads. Fixed to verbose=True. That alone was not enough: the real label passed to each godmode call was f"{name}[godmode bits={cand_bits}]" — checked directly against hessian_story_lib.RANK_RE, and the space inside "[godmode bits=" breaks the regex’s name capture entirely (confirmed: zero match, not a garbled one). Real fix: the label is now the clean tensor name alone. This loses nothing real — effective_rank(H_I)/(H_O) does not depend on candidate bit-width (established earlier: the Hessian comes from real calibration activations, before any rounding decision), so every real godmode candidate for the same tensor prints the identical real number under the same real name; load_hessian()’s overwrite-on-duplicate-key behavior is a correct no-op here, not data loss. cand_bits itself is preserved separately, as godmode_candidates’s own real dict key. Both fixes are already in 05_full_model_quantize.py — verified via a direct regex test against the real corrected print format before this section was written.

Step 3 — extract and merge the new real scores. As of 2026-09-16 this is one real command, not a manual two-step (extract, then hand-merge) — the watch subcommand does both, safe to run live while this build is still in progress:

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  watch \
  <this build's real resume-dir> \
  /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --interval 30

Only ever adds real, newly-available tensors (or updates changed ones with --overwrite) — never removes or silently clobbers an existing real entry. Real, permanent output either way: scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json.

Real end state once this completes: 496 of 497 real tensors scored (only lm_head permanently excluded) — the region floors (--late-attn-min-bits, --early-qkv-min-bits) can then be turned off for real, with every tensor they used to blanket-protect instead getting real, individual, Hessian-informed protection — the actual, original goal this whole investigation was aimed at.

11.8 MILP-aware Hessian (2026-09-15) — hard floor + lexicographic priority, now part of the solver

Real problem this closes: §11.5/§11.6 above showed the soft percentile weight is a real, continuous, per-tensor signal — a genuine improvement over the flat 100× category weight. But a direct check found it is still not a guarantee: even the single most dangerous real tensor in the whole model (layers.63.mlp.up_proj) only reached 8-bit in the real production plan, and pushing the multiplier to strength=1000 (500× stronger than the tested default) changed zero outcomes in the top 20 most dangerous real tensors. The reason: these tensors’ real, raw KL cost is close enough to flat across candidate bit-widths that no multiplier of an already-small number wins against tensors with genuinely large KL cost, under the same fixed BPW budget. A soft weight cannot fix a signal that’s fundamentally too small to matter in the objective it’s added to.

The fix is two new, independent, opt-in flags — neither changes existing behavior when omitted (verified directly: identical output, 0/497 tensors differ, with both flags left out):

--hessian-floor-tiers "0.05:8,0.15:6" — a real, per-tensor MILP hard floor, computed from each tensor’s real Hessian percentile rank (same ranking as §11.5). Any real tensor whose rank falls under 0.05 gets a genuine minimum of 8-bit; under 0.15 gets a minimum of 6-bit — enforced the exact same way --lm-head-min-bits/--boundary-min-bits already work (add_min_bits_constraint, a real LinearConstraint, not a weight). The solver remains free to go above the floor — real, confirmed: of the 54 tensors floored in the tested run, 29 ended up at full 16-bit anyway. The real feasibility ceiling for this exact production budget (5.028075 BPW) was found by bisection: coverage beyond ~18.7% of scored tensors (67 of 360) makes the solve genuinely infeasible; the tested two-tier spec (54 tensors) sits safely under that.

--hessian-primary (+ --hessian-primary-gap, default 1e-7) — a real, lexicographic two-phase solve. Phase 1 solves a new MILP minimizing danger-weighted bit deficit, KL playing no role at all; its result is pinned (to a provably tight gap, not an accepted approximation — confirmed gap=0.0 on the tested run) before the existing weighted-KL phase runs, now constrained to only Phase-1-optimal solutions. This is the same two-phase mechanism this file already uses for --raw-tiebreak, just with Hessian danger promoted to the primary objective and KL demoted to tie-break. Real, measured consequence: KL ends up deciding 0 of the 360 real Hessian-scored tensors — Phase 1’s ranking is precise enough that ties among them are effectively never real — and only 3 of the 497 total (both unscored, no Hessian data at all).

Floor alone and lexicographic alone each fix half the real gap; combined, they dominate both individually on every real metric checked (kept-at-16-bit and cut-to-≤5-bit, across every danger tier from top-10 to top-108) — see research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/MILP_AWARE_HESSIAN_GUIDE.html for the full real comparison.

The exact real command, everything else matching the already-benchmarked production checkpoint (--max-low-bit-run 999 is required alongside --hessian-primary — without it, the separate run-guard mechanism piles extra constraints on top each iteration and the combined solve goes infeasible; unrelated to the Hessian mechanism itself):

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 \
  --candidates 4,5,6,8,16 \
  --pareto none \
  --group-size 64 \
  --late-attn-min-bits 0 \
  --early-qkv-min-bits 0 \
  --lm-head-min-bits 6 \
  --boundary-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 0 \
  --hessian-floor-tiers "0.05:8,0.15:6" \
  --hessian-primary \
  --max-low-bit-run 999 \
  --output <your_output_path>.json

How these flags actually pick bit-widths (added 2026-09-19). The rank, the danger weight, the floor tiers and the two-phase solve are explained step by step, with diagrams and checked numbers, in How the Solver Turns a Hessian Score into a Bit-Width. The one-line version: under --hessian-primary, spare bits above the floors go to the tensors with the most danger per million parameters, so a tiny nearly-safe tensor can reach 16-bit while the single most dangerous (large) tensor stays at its floor, and tensors with no score are pushed down to their floors.

--hessian-unscored-fallback {zero,kl} (added 2026-09-19, default zero = unchanged behaviour). Under --hessian-primary, a tensor with no score in the Hessian score file gets danger weight 0, and on that scale 0 is the safest end, so it is under-protected: on the 407-scored snapshot 64 of the 90 unscored tensors ended at 4-bit where a KL-driven plan keeps 89 of them at 16-bit. --hessian-unscored-fallback kl gives each unscored tensor the percentile rank of its own isolated KL at the lowest candidate bit (over all tensors, 1.0 = most dangerous) as its danger weight instead. Scored tensors and every floor are unchanged; it fades out as the sweep scores more tensors. Measured: 84 of 90 unscored tensors at 16-bit, 8 scored tensors move, 3 fewer scored tensors reach 16-bit, raw summed KL 50.09 to 39.41 (production 37.45). Full explanation and evidence: How the Solver Turns a Hessian Score into a Bit-Width, section 5.6. The complete, exact command with the fallback (recommended while coverage is incomplete; copy and paste as is, it writes one new plan file and builds nothing):

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 \
  --candidates 4,5,6,8,16 \
  --pareto none \
  --group-size 64 \
  --late-attn-min-bits 0 \
  --early-qkv-min-bits 0 \
  --lm-head-min-bits 6 \
  --boundary-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 0 \
  --hessian-floor-tiers "0.05:8,0.15:6" \
  --hessian-primary \
  --hessian-unscored-fallback kl \
  --max-low-bit-run 999 \
  --output 04_LANGUAGE_PLANS/HESSIAN_TEST/plan_floor_primary_klfallback_2026-09-19.json

Independent code review (2026-09-15) confirmed the core design correct and properly gated, and found two real, since-fixed issues: the floor’s own audit field originally reported requested tiers rather than actually-shipped bits (fixed — now reports real shipped bits plus an explicit violations list, [] on the tested run), and Phase 1’s gap tolerance was originally looser than the pin’s own framing implied (fixed via --hessian-primary-gap, default 1e-7, independent of the general --quality-mip-rel-gap).

Warning, real and confirmed (2026-09-18) — --pareto none is required whenever a Hessian file drives allocation

--pareto defaults to meaningful, which removes any candidate bit-width a tensor’s own isolated-KL data says isn’t a meaningful improvement over a cheaper one — a real, useful filter when KL is the only signal deciding allocation, but it has no idea a Hessian floor is about to demand a specific bit-width for a completely different, KL-independent reason. Real, checked at 405/497 coverage: 55 real tensors had their Hessian-floor-required bit filtered out by exactly this mechanism, forcing several to jump all the way to 16-bit instead of the tier they actually needed, and in some real runs this combined with --hessian-primary’s own tight re-pinning to produce genuine MILP infeasibility (HiGHS Status 8: Infeasible) — worse as real coverage grows, not better. --pareto none fixes it: confirmed real re-runs of the same production budget, previously-failing configurations solved cleanly once added.

Separately, --hessian-weight-strength was directly, isolatedly tested (2026-09-18): re-running the identical command with only that flag toggled between its default (2.0) and 0, --pareto none held constant, produced bit-for-bit identical results — the soft weight changes nothing once the floor and/or --hessian-primary are active. It was only ever inflating the reported OptiQ-weighted ΣKL statistic (by up to 3× per tensor), never the real allocation. Full real evidence: MILP_AWARE_HESSIAN_GUIDE.html.

11.9 The Hessian score toolkit, consolidated (2026-09-16) — and what the real Hessian-vs-KL investigation actually found

The toolkit used to be three separate scripts in two different directories. That was a real, reported problem (“scripts all over the place… nightmare to keep track”) — now consolidated into one file, one location, three subcommands. The two old scripts (08b_merge_live_hessian_scores.py, 08c_build_hessian_hybrid_checkpoint.py) no longer exist; everything they did now lives here:

# One-shot: extract scores from a finished build's own logs
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  extract \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume \
  --output /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json

# Keep growing that same file live, from any running job (Ctrl+C to stop; add --no-loop for one pass)
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  watch \
  <resume_dir> \
  /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --interval 30

# Merge a real KL checkpoint with a real godmode sweep into one file the
# UNMODIFIED MILP solver (section 3.3 below is unaffected) already reads natively
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  hybrid \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  <resume_dir>/godmode_sensitivity_checkpoint.jsonl \
  --output /tmp/hessian_hybrid_checkpoint.json

This exact filename and location (scripts/yaqa_port/, not the research_hadamard_blowup/ subfolder) is load-bearing — 05_full_model_quantize.py imports it by file path for its own automatic end-of-build hook. Do not rename or relocate it without also updating that hook.

The real question this whole toolkit exists to answer: does the free Hessian score agree closely enough with the expensive cascaded-KL measurement to replace some of it? Investigated directly tonight, real data throughout. Honest summary — full trail, diagrams, and every command: research_hadamard_blowup/exports/HESSIAN_MASTER_REFERENCE.html.

12. What is NOT yet done (honest, from PORT_LEDGER.md)

See PORT_LEDGER.md for the complete, detailed engineering log, and CHANGELOG.md for the short chronological list of every real change.


© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.