2026-09-05
Everything needed to run this codebase, start to finish, nothing
skipped. Real path:
/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_portReal, verified versions (requirements.txt):
mlx==0.32.2, mlx-lm==0.31.3,
numpy==2.4.6. Apple Silicon Mac required — MLX is
Metal-based, no CPU-only or CUDA path exists.
To set up fresh elsewhere:
uv venv --python 3.11
uv pip install -r requirements.txt"$PYBIN" test_regression.pyMust print ALL REGRESSION TESTS PASSED. Eleven real,
independently-cross-checked guards — every bug found this session gets a
permanent test so it can’t silently come back.
"$PYBIN" 01_gradient_collection_test.py
"$PYBIN" 02_sketch_b_synthetic_test.py
"$PYBIN" 03_rounding_update_test.py
"$PYBIN" 04_real_layer_test.py # one real layer, end to end
"$PYBIN" 04_real_layer_test.py --from-cache # skip model load/backward on repeat runs"$PYBIN" 05_full_model_quantize.py --subset default
"$PYBIN" 05_full_model_quantize.py --subset "language_model.model.layers.63.mlp.up_proj"Never call 05_full_model_quantize.py directly
with --output for a full run. It will hit real
memory/Metal crashes documented in PORT_LEDGER.md. Always
use the orchestrator:
./run_full_yaqa.sh <output_dir> [flags...]
# The real command actually used for this project's model:
./run_full_yaqa.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw \
--n-calibration 4 --incoherence nonerun_full_yaqa.sh actually does, in order<output_dir>.yaqa_resume/manifest.json.PORT_LEDGER.md.)06_lm_head_gptq.py automatically — real, scoped
one-sided GPTQ correction for lm_head, the one tensor
YAQA’s two-sided correction structurally cannot reach.--assemble-from-resume) —
reads every real result from the resume cache and writes the actual,
complete, loadable model to <output_dir>.| Flag | What it does | Real effect |
|---|---|---|
--source <path> |
BF16 source model to quantize (was hardcoded, now a real override) | Must be consistent across every step of one run |
--plan <path> |
Real per-tensor bit-allocation plan JSON (was hardcoded, now a real override) | A different --source almost certainly needs its own
real plan too |
--n-calibration N |
Real stratified sequences per domain (6 domains, so total = 6×N) | More wall-clock time, NOT more peak memory (see below) |
--incoherence {none,hadamard} |
none = production, real mlx_lm-loadable
today. hadamard = research branch, needs an unbuilt
loader |
none is the real default; don’t use
hadamard unless you’re building the loader too |
--hessian-batch-budget-gb |
Real memory budget per batch (default 8GB) | Lower = more batches, less peak memory per batch; raise only if you have real headroom |
--lm-head-name <name> |
Real dotted name of the vocabulary-output tensor | Auto-detected by default — you never need to pass
this. Two real, cross-checked signals: the model’s real vocab
size (found anywhere in its config) matched against which plan tensor’s
real output shape equals it, cross-checked against the
.lm_head naming convention. Only pass this explicitly if it
raises an error (shows exactly what each signal found) |
--correct-mtp |
Real YAQA correction for the MTP speculative-decode sidecar (added 2026-09-08) | Off by default — the sidecar is otherwise naive-quantized. See section 6, Scenario C below |
--reuse-cached-fallback |
Skip lm_head’s damping-search retry if a prior run already recorded
naive_fallback (added 2026-09-08) |
Off by default — see section 9 below |
--skip-lm-head-gptq |
Never call 06_lm_head_gptq.py at all for this run —
lm_head goes straight to Part 5c’s plain native
mx.quantize fallback (added 2026-09-13,
run_full_yaqa.sh only) |
Off by default — see section 9 below for exactly when to use this
vs. --reuse-cached-fallback vs. neither |
Full list of every real flag on both scripts, cross-checked against
--help: YAQA_COMMAND_REFERENCE.md’s “Every
real flag” section — not duplicated here.
Passed, not built in. --n-calibration
sets k_per_domain for the real stratified calibration
loader (6 domains: agent, code, instruct, prose, thought, tool). At the
default of 4, that’s 24 total real sequences (N=24) — matching this
project’s own established convention (same as
08_gptq_apply_plan.py’s own default).
Does it affect memory? Not the way you’d think.
H_I/H_O are fixed-size matrices
(in_features²/out_features²) no matter how
much calibration data is used — more calibration only adds more terms
into the same-sized running sum, it never grows the matrix. What
actually changes is how many real forward+backward chunks get
run (--batch-size, default 1 sequence at a time) —
so --n-calibration affects real wall-clock time (linearly),
not peak memory. The real memory pressure seen during this project’s own
full run came from which tensors happened to be in a batch (the
large 17,408-dimension MLP tensors), not from the calibration count.
Real situations, each a complete, copy-pasteable example — the
difference between them is entirely about whether
<output_dir>.yaqa_resume already exists, what’s in
it, and which real plan you’re building from (this list has grown as
real, new situations came up — current count: five, A through E).
Nothing to reuse: no output directory, no resume cache. This is the plain, default path.
NEW_OUTPUT="/Users/hghelab/.mtplx/models/My-New-Model-YAQA-5bpw"
./run_full_yaqa.sh "$NEW_OUTPUT" --n-calibration 4 --incoherence nonerun_full_yaqa.sh creates
${NEW_OUTPUT}.yaqa_resume fresh, runs every real batch from
zero (see section 5’s “what it actually does” above), then lm_head, then
assembles. No special flags needed — this is the
every-tensor-from-scratch case.
A run was interrupted (crash, Ctrl+C, machine slept), or
you just want to change a flag like
--hessian-batch-budget-gb and continue. Run the
exact same command again, unchanged (same
<output_dir>):
./run_full_yaqa.sh "$NEW_OUTPUT" --n-calibration 4 --incoherence noneEvery tensor already completed (checked by name against the real
manifest in ${NEW_OUTPUT}.yaqa_resume/manifest.json,
matching bits/group_size) is skipped automatically and picked back up
where it left off — this works even if
--hessian-batch-budget-gb changes between runs, since the
skip check is per-tensor, not per-batch-index. Nothing needs to
be copied or renamed for this case — that’s only for Scenario C
below, which deliberately reuses a cache under a different
name.
Different from Scenario B: here the trunk build already finished
successfully (a real, working model exists), and you want to add the MTP
sidecar correction (--correct-mtp, see the flags table in
section 5) without recomputing the whole multi-hour trunk correction.
This reuses the existing resume cache under a new name,
so the original completed model is never touched or put at risk.
# Step 1 -- copy the existing resume cache under a NEW output name (non-destructive)
OLD="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32"
NEW="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected"
cp -a "${OLD}.yaqa_resume" "${NEW}.yaqa_resume"
# Step 2 -- launch: skips every already-cached trunk/lm_head tensor, goes straight to
# reassembly + the new MTP correction
./run_full_yaqa.sh "$NEW" --n-calibration 4 --incoherence none --correct-mtp --reuse-cached-fallback
# Step 3 -- watch it live, in a second terminal tab
./monitor.sh "${NEW}.yaqa_resume"Real pitfall — only run Step 1 once.
cp -a src dst copies src’s contents into a new
dst only if dst doesn’t already exist; if
dst already exists (e.g. Step 1 was already run once), it
instead nests src inside dst,
silently doubling real disk usage with no error (confirmed: 16GB →
32GB). Check ls -d "${NEW}.yaqa_resume" first — if it
already exists, skip straight to Step 2, or rm -rf it
before re-copying.
Full flag-by-flag rationale (why --reuse-cached-fallback
exists, what --mtp-bits defaults to, etc.):
YAQA_COMMAND_REFERENCE.md’s dedicated MTP section — not
duplicated here, this is the same real 3-step procedure either way.
Different question from Scenarios A-C: those build a real model at
the plan’s one assigned bit per tensor. Godmode instead asks, for every
tensor, “what would YAQA’s real corrected error be at every
candidate bit-width” — 2, 3, 4, 5, 6, and 8-bit by default, not just the
one the plan picked. Reuses the same H_I/H_O
Part 5a already computes per tensor (confirmed bit-width-independent) —
no extra forward/backward calibration passes, only extra (cheap)
correction/safety-gate calls. Does not touch or replace
the existing single-bit build — it’s a pure addition, on by explicit
flag only.
# Fresh checkpoint, from scratch, default candidate bits (2,3,4,5,6,8), no 16-bit
# (16-bit is the BF16 reference itself -- zero rounding error by construction, already
# known, not worth spending real time re-measuring)
GODMODE="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-GODMODE"
"$PYBIN" 05_full_model_quantize.py \
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara \
--plan 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE/plan_v4_round2.json \
--resume-dir "${GODMODE}.yaqa_resume" \
--godmode-multi-bit-checkpoint \
--n-calibration 4 --incoherence noneStop: Ctrl+C, or kill the process —
same safety property as the main run. Each real tensor’s godmode sweep
writes to
${GODMODE}.yaqa_resume/godmode_sensitivity_checkpoint.jsonl
immediately after that tensor’s last candidate bit finishes (one JSON
object per line, appended) — a kill loses at most the tensor currently
in flight, never anything already written.
Resume: run the exact same command again, unchanged.
On startup it reads the existing .jsonl, builds the set of
tensor names already present, and skips them — matching this file’s own
resume-cache convention, not a new mechanism.
Real cost, verified, not estimated: the existing
single-bit trunk correction took 27.56 real hours for all 362 tensors
(measured directly from real batch_*.log file timestamps,
2026-09-11) — that’s the real floor, since Part 5a’s expensive
calibration pass is unchanged. Godmode’s own added overhead (looping 6
candidates’ worth of cheap correction calls instead of 1, per tensor)
has not been measured yet — the first real run is the first real
measurement of it, not a promise made in advance.
Override the candidate set with
--godmode-candidate-bits "4,5,6" (comma-separated, any
subset); override the output path with
--godmode-checkpoint-path.
Real background, not a hypothetical: the live stratified-N24 cascade
(V4_CLEAN_STRATIFIED_CASCADE) ran its full three rounds and
finished (497/497 tensors measured in round 3, no crash, real plan
written). Separately, a real bug was found and fixed during that same
investigation: language_model.lm_head — 1.27 billion
parameters, the single largest tensor in the model — was landing at Q4
in every real cascaded round (round 1 and round 2 both), despite
carrying a 100x structural_weight KL-protection multiplier
in 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py. Root
cause, verified directly: cascaded measurement at lm_head
(measured last in causal order, on top of an already-degraded upstream
context) is nearly flat across candidate bits (0.2326 at 4-bit
vs. 0.2289 at 8-bit — a 1.6% range), so a mathematically correct MILP
finds it “cheap” to spend that tensor’s budget elsewhere even with the
100x weight applied. Fixed by adding a real, hard minimum-bits floor
(--lm-head-min-bits, default 6, on by default) to the
optimizer script itself — full trace, real numbers, and independent
confirmation (the vendor optiq package’s own unmodified
allocator never puts lm_head below Q6 on the same real
data) in
research_hadamard_blowup/IMPROVEMENT_LEDGER/04_LMHEAD_PROTECTION_INCIDENT.html.
Re-optimizing round 1’s real, already-measured, still-valid
checkpoint (valid because it was measured against P0, which never
changes) through the now-fixed optimizer produced a corrected P1 with
lm_head=6 and only 11 other tensors shifted (12 total, 2.4%
of 497) — real, exact diff in the incident doc above. Rounds 2 and 3
were not re-measured with the floor (that needs real
new GPU time, a deliberate decision, not an oversight — see
IMPROVEMENT_LEDGER/05_LMHEAD_FIX_NEXT_STEPS.html), so this
scenario builds from the corrected P1 specifically.
Why Step 1 below matters — this is not optional.
run_full_yaqa.sh always derives its resume folder as
"${OUTPUT_DIR}.yaqa_resume", straight from the
--output name — there is no flag to point it at a
different, pre-existing resume-dir. The uncorrected P1 build already
computed YAQA correction for 363 real tensors under
Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume.
Only 12 of 497 tensors differ between the uncorrected and corrected P1
plans (2.4%), and YAQA’s per-tensor correction is independent of which
plan produced the target bit-width
(05_full_model_quantize.py’s resume-skip check, real line
numbers ~733-758, matches on bits/group_size/
curvature_version per tensor name) — so almost all of that
prior work is reusable. But a brand-new --output name means
a brand-new, empty .yaqa_resume folder with
nothing in it to reuse, unless the old one is physically copied to the
exact path the new run will look for first.
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
# Step 1 -- seed the new resume-dir with the existing, closely-related cache (non-destructive:
# copies TO a new path, never touches the original). Check real disk headroom first (df -h) --
# this duplicates the full 16GB cache size before anything can be deduplicated away.
SRC=~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume
DST=~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6.yaqa_resume
cp -R "$SRC" "$DST"
# Step 2 -- launch. Real, verified result of Step 1 on this project's own data: 236 of 343
# real targets resume-skip (already match), leaving only 107 tensors (8 real batches) to
# actually recompute -- checked directly via --print-num-batches before committing to the run.
./run_full_yaqa.sh \
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6 \
--skip-lm-head-gptq \
--plan /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/plan_v4_round1.json \
--n-calibration 4 \
--incoherence none \
--correct-mtp
# Step 3 -- watch it live, in a second terminal tab
./monitor.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6.yaqa_resumeReal pitfall — only run Step 1 once, same as
Scenario C: cp -R src dst copies src’s
contents into a new dst only if dst doesn’t
already exist; if dst already exists (Step 1 already ran),
it nests src inside dst instead,
silently doubling real disk usage with no error. Check
ls -d "$DST" first — if it already exists, skip straight to
Step 2, or rm -rf it before re-copying.
Why --skip-lm-head-gptq specifically, here — not
--reuse-cached-fallback, not neither: this command’s
--output name is new, so
<output>.yaqa_resume is a brand-new folder (even
seeded via Step 1, nothing in it is tagged for lm_head
specifically — the old cache never ran GPTQ against this exact plan).
There is nothing cached for lm_head to reuse.
--reuse-cached-fallback only skips a repeat of the
damping search when a matching prior attempt is sitting in the
same resume-dir already — pointed at a folder with nothing
cached for this tensor, it finds nothing and the real, full,
~15-20-minute damping search would run anyway, for a tensor already
proven (real manifest evidence, both the old build and this project’s
own real testing) to always end in naive_fallback
regardless. --skip-lm-head-gptq is the only one of the
three real options that guarantees zero wasted time here, because it
never calls 06_lm_head_gptq.py at all — see section 9 below
for the complete decision guide covering all three real situations,
spelled out, not summarized.
None of the three examples above pass every real flag — every flag not mentioned silently uses its own real default, and those defaults are never “off”/“nothing”, they are specific, real values:
--subset → full real plan (all 363 tensors), not a
smoke-test subset.--source / --plan → this project’s own
real model and plan JSON (see the flags table above) — only override
these for a genuinely different model.--seq-len → 128. --batch-size → 1
(sequences per forward+backward call).--fallback-bits / --fallback-group-size →
6 / 64 (for any leaf module the plan doesn’t cover — none exist in this
project’s own plan, but the default exists for a different
--plan).--hessian-batch-budget-gb → 8.0 (GB of
H_I+H_O per batch).--sigma-reg → 1.0 (damping constant).--incoherence → none if omitted from your
own command, even though every example above passes it explicitly —
passing it explicitly is a habit worth keeping since
hadamard needs a loader that doesn’t exist yet, but
none is what you’d get either way.--correct-mtp → off. Omitting it
(Scenarios A and B) means the MTP sidecar stays naive-quantized — this
is not a bug, it’s the real default, unchanged since before 2026-09-08.
Only Scenario C turns it on.--mtp-bits / --mtp-group-size → the plan’s
own native MTP bit-width/group-size (whatever the naive sidecar would
already use) — only relevant when --correct-mtp is also
on.--reuse-cached-fallback → off.
Omitting it means a cached naive_fallback for
lm_head is NOT reused — the real damping search reruns from
scratch every time, which is always correct behavior, just potentially
slower (see section 9).--skip-lm-head-gptq → off
(run_full_yaqa.sh only). Omitting it means
06_lm_head_gptq.py always runs, exactly as before this flag
existed — zero behavior change for anyone not using it. See section 9
for the full three-way decision guide against
--reuse-cached-fallback.--godmode-multi-bit-checkpoint → off.
Omitting it means the run is exactly the same single-bit build as before
this flag existed — zero behavior change (see Scenario D).--godmode-candidate-bits → 2,3,4,5,6,8
(only read when the flag above is on).--godmode-checkpoint-path →
{resume_dir}/godmode_sensitivity_checkpoint.jsonl.The complete, exhaustive list (every flag on both scripts,
cross-checked directly against real --help output):
YAQA_COMMAND_REFERENCE.md’s “Every real flag” section.
Recommended (added 2026-09-16): just run
watch.sh with no arguments — it lists every real
*.yaqa_resume job it finds under
/Users/hghelab/.mtplx/models/, numbered, and prompts for
which one to watch. No path to remember, no hardcoded wrapper script per
job (that pattern — one watch_<job>.sh per job — was
tried and replaced; it doesn’t scale past a couple of jobs).
./watch.shOr skip the prompt if you already know which job:
./watch.sh HESSIAN-PROBE # by name fragment
./watch.sh 2 # by number from the listUnder the hood this still calls monitor.sh (see below) —
watch.sh is just the real, generic entry point so you never
have to construct that path by hand.
monitor.sh
directly, if you already have the exact path./monitor.sh <output_dir>.yaqa_resume/logsPoint it at the logs folder (not a specific file) — it
auto-detects and follows whichever batch is currently live, switching
automatically as the run progresses. Color-coded: green = running, red =
stopped/crashed, yellow = memory getting low. Shows real ps
stats, real system memory/swap, and the last 15 real log lines.
Every time monitor.sh refreshes (every 4s), it also
regenerates a real, single-file HTML dashboard at:
<output_dir>.yaqa_resume/dashboard.html
Open that file once in a browser — it auto-refreshes itself every 8s via a meta-refresh tag, so it stays live without you re-opening it. Real example (this project’s own model):
file:///Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw.yaqa_resume/dashboard.html
Everything on it is computed from real data each refresh — no simulated numbers:
Overall progress: a real progress ring (tensors done / 363) plus a real ETA, computed from the actual measured completion rate over the last 20 real tensors (not a fixed estimate).
Live process: real CPU/memory/elapsed stats for the currently-running batch.
Real completion velocity: a real chart of tensors completed per 15-minute window, over the run’s actual full history — including any real gaps (e.g. an overnight pause), never trimmed away, since hiding a real gap would make the chart quietly dishonest.
Real quality result: how much error YAQA
actually removed vs. naive rounding — win/loss counts, a real histogram
of per-tensor error reduction, broken down by bit-width, plus a
live-computed interpretation that names the actual real
best/worst tensor by name each refresh (not a fixed paragraph — this is
what first surfaced the real
layers.14.linear_attn.in_proj_qkv outlier, see
PORT_LEDGER.md’s RCA section).
Generated by generate_dashboard.py (called automatically
from inside monitor.sh – you never run it directly). Reuses
the exact same real interpretation logic as
interpret_manifest.py (section 8 below), just rendered
visually instead of as text. As of 2026-09-08, monitor.sh
also serves two live, no-flicker versions over a local HTTP server
(file:// pages can’t auto-refresh via fetch()
in Chrome): the same internal dashboard, and a separate customer-facing
view — both show the same real milestones, including a dedicated
YAQA-correcting MTP sidecar step when
--correct-mtp is in use.
Real fix, 2026-09-16 — the ports are no longer
fixed. They used to always be
8765/8766, which meant two real jobs being
watched at the same time fought over the same two ports and kept
silently swapping which job’s dashboard you were actually looking at
(caught live, mid-session, from the real self-heal log line). The
internal dashboard’s port is now derived from a hash of the specific
job’s own resume-dir path, so two jobs never collide —
monitor.sh prints the real, resolved port for the job
you’re watching on every refresh cycle; read it from there rather than
assuming 8765. The customer-facing port stays fixed at 8766
(it always serves the same shared brand/ folder regardless
of job, so there’s no real collision to avoid there), but the file it
writes is now named per-job
(live-status-<job-name>.html) instead of a single
shared file every job used to overwrite.
"$PYBIN" interpret_manifest.py <output_dir>.yaqa_resumeReal example (this project’s own model):
"$PYBIN" interpret_manifest.py /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw.yaqa_resumeExplains Metric A (err_yaqa/err_naive —
real Hessian-weighted reconstruction error, not KL
divergence) vs. Metric B (plain Frobenius distance) in plain terms, and
breaks real progress down by tensor kind and bit-width.
lm_head — what it gets, why it’s separate, and exactly
which flag to uselm_head is and why it needs its own pathlm_head is the model’s final output-vocabulary
projection (1.27B parameters). YAQA’s normal two-sided correction cannot
reach it: its output-side Hessian alone (tied to the 248,320-word
vocabulary) needs 246.7GB — a real memory limit, not a tunable one.
06_lm_head_gptq.py exists to give it something
better than nothing: a real, scoped, one-sided GPTQ correction instead
(GPTQ only ever needs the small input-side Hessian, which is
affordable). This is a genuine, intended, standard part of the pipeline
— run_full_yaqa.sh calls it automatically, every time,
right after all the other batches finish — not an optional extra step
you’re expected to skip by default.
If GPTQ’s real search finds a damping ratio that passes its safety
gate, the manifest entry is tagged "method": "gptq". If
every real damping ratio it tries fails that gate, it falls back and
tags the entry "method": "naive_fallback" — plain native
mx.quantize, the same thing
05_full_model_quantize.py’s own Part 5c would have done for
lm_head anyway if GPTQ were skipped entirely. Either way
lm_head still ends up quantized at the plan’s own
bits/group_size (6-bit/G64 for this project’s corrected plans) — it is
never left at BF16. Only the correction method differs, not the
bit-width.
On this project’s own real tensor, GPTQ’s search has been run for
real, multiple times, and has recorded naive_fallback every
single time — every real damping ratio it tried failed the safety gate.
That is not a hypothesis; it is the observed, repeated outcome on this
specific tensor with this specific Hessian.
Situation A — you’re building a genuinely new corrected plan
for the first time, and lm_head’s known-futile outcome is
not going to change. This is the common case: a new plan, a
brand-new --output name, therefore a brand-new, empty
<output>.yaqa_resume folder. Nothing is cached in
that folder yet for anything, lm_head included. In this
situation, letting 06_lm_head_gptq.py run means it
re-derives the exact same naive_fallback result it always
gets on this tensor — burning ~15-20 real minutes to arrive at a result
you already know in advance. Use
--skip-lm-head-gptq. This skips calling
06_lm_head_gptq.py at all; lm_head still gets
set correctly, at the plan’s real bits, via
05_full_model_quantize.py’s own built-in Part 5c
native-quantize fallback — numerically identical to what GPTQ’s
naive_fallback path would have produced anyway, just
without wasting the real time re-deriving it.
Situation B — you’re re-running against a resume-dir that
genuinely already has a cached lm_head GPTQ attempt from a
previous run against the exact same tensor and Hessian (for
example: a prior run on this same output crashed partway through, after
lm_head’s GPTQ step had already completed and been written
to the manifest, and you’re resuming it). Here, and only here,
--reuse-cached-fallback does real work: it checks the
manifest for a matching prior lm_head entry and, if found,
skips straight past re-running the damping-search loop instead of
repeating it. Use --reuse-cached-fallback
in this situation specifically. Passing it against an empty
resume-dir (Situation A) is not wrong, exactly, but it is a no-op —
there is nothing cached yet for it to find, so
06_lm_head_gptq.py still runs its full real search from
scratch regardless of the flag.
Situation C — you have a specific, real reason to believe
GPTQ might behave differently this time (a materially different
plan’s lm_head bit-width, a changed calibration set, a code
change to the damping search itself) and you want a genuine,
un-skipped attempt. Pass neither flag.
06_lm_head_gptq.py runs its real search fresh, with no
shortcut and no cached result to fall back on, and whatever it finds
(gptq or naive_fallback) is recorded as a new,
real manifest entry.
The three flags are not interchangeable and do not stack usefully:
--skip-lm-head-gptq prevents
06_lm_head_gptq.py from running at all, so
--reuse-cached-fallback has nothing left to apply to if
both are passed together — pick the one that matches your actual
situation, not both.
run_full_yaqa.sh always derives its resume folder as
"${OUTPUT_DIR}.yaqa_resume" — directly from whatever
--output/first-positional-arg name you give it. There is no
flag to point it at a different existing resume-dir. This
matters because YAQA’s per-tensor correction is independent of which
plan produced the target bit-width (checked by
05_full_model_quantize.py‘s resume-skip logic, real line
numbers ~733-758: it matches on bits,
group_size, and curvature_version per tensor
name) — so a new plan that only changes a handful of tensors’ bits (this
project’s real lm_head-floor fix changed 12 of 497, 2.4%) should be able
to reuse almost everything already computed under a previous,
closely-related plan. But the orchestrator has no built-in way to know
that a prior resume-dir it should look at even exists — a brand-new
output name always starts from a brand-new, empty
.yaqa_resume folder, full stop.
The real fix, done the same way this project already handles reusing a prior cascade checkpoint: physically copy the old resume-dir to the exact path the new run will look for, before running it.
SRC=~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume
DST=~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6.yaqa_resume
cp -R "$SRC" "$DST"After the copy, running
05_full_model_quantize.py --print-num-batches (or just
letting run_full_yaqa.sh start normally) against the new
plan will report real resume-skip hits for every tensor whose
bits/group_size/curvature_version match — verified directly on this
project’s own real P1-vs-P1-floored plan pair: 236 of 343 real
targets matched and were skipped, leaving only 107 tensors (8
real batches) actually needing recomputation. Check real disk headroom
first (df -h) — the copy duplicates the full resume-dir
size (16GB in this project’s case) before any of it can be deduplicated
away.
The real output model is a complete, standard directory at
<output_dir> — safetensors shards,
config.json, vision/audio sidecars, MTP head, all
reattached automatically via the same real mechanism
08_gptq_apply_plan.py already uses in production. Loadable
with plain mlx_lm.load(). Verify it:
"$PYBIN" -c "
from mlx_lm.utils import load
import mlx.core as mx
model, tok, cfg = load('<output_dir>')
ids = tok.encode('The capital of France is')
logits = model(mx.array([ids]))
mx.eval(logits)
print(tok.decode([int(mx.argmax(logits[0, -1]))]))
"Should print a coherent next token. This exact check has been run and passed on this project’s own real saved output already.
This is a real, separate addition to
02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py — the MILP
that turns real cascaded sensitivity data into a bits-per-tensor plan.
It answers a real, specific question: the solver’s own protection for
“important” tensors (lm_head, first/last-layer boundary
tensors) was, until now, a single flat 100× KL-weight multiplier applied
to a hand-picked category — a real, working patch, but blind to whether
any given tensor’s own real curvature actually justified that
multiplier. This section documents the real fix: a second, genuinely
independent signal (real Hessian geometry, not another KL measurement)
that can replace the flat multiplier with a continuous, per-tensor
one.
--hessian-score-file HESSIAN_SCORE_FILE
Optional path to a real JSON {tensor_name: hess_score} map (e.g. exported from
hessian_story_lib.load_hessian()). When given, multiplies each named tensor's
structural_weight by a real, bounded factor derived from its score's percentile
rank among all scored tensors -- the most dangerous tensor (lowest score) gets
the largest boost, the safest (highest score) gets none. Tensors with no real
score in the file are left at factor 1.0 (unchanged). Default: no file, no effect.
--hessian-weight-strength HESSIAN_WEIGHT_STRENGTH
Real multiplier range applied via --hessian-score-file: the single most
dangerous tensor's structural_weight is scaled by (1 + this), the safest by
1.0, linearly interpolated by real percentile rank in between. Only read when
--hessian-score-file is given. (default: 2.0)
Both are off by default — omit
--hessian-score-file entirely and the solver’s output is
byte-identical to before this feature existed (verified directly: same
real plan, zero diff, with the flag omitted).
hess_score(tensor) = mean(effective_rank(H_I)/dim_in, effective_rank(H_O)/dim_out),
where effective_rank(H) = trace(H)² / Σ(H²) — a real
spectral participation-ratio diagnostic computed from each tensor’s own
real H_I/H_O (the same Hessians YAQA’s own two-sided correction already
builds; nothing new is measured). Low score = the Hessian’s energy is
concentrated in very few directions (narrow, dangerous). High score =
spread broadly (forgiving). This is parsed directly from the real
"... real effective rank -- H_I=.../..., H_O=.../..." log
lines every real YAQA build batch already prints —
hessian_story_lib.load_hessian() already does this parsing;
nothing new to write.
Real automation (added 2026-09-14, direct user request — “we might need to streamline an orchestrator so the Hessian score is always there”): the raw per-tensor data was already being printed for free by every real correction pass — the only real gap was that nobody ever consolidated it automatically. That gap is now closed:
05_full_model_quantize.py itself now extracts
it automatically, real and non-blocking, as the last real step
of its own --assemble-from-resume code path — writes
${RESUME_DIR}/hessian_scores.json. A failure here never
invalidates the real output model that was just saved.run_full_yaqa.sh, as a
separate step after its own call to assembly. That’s not actually the
only real way this project invokes --assemble-from-resume —
MAIN_RESUME_BUILD.sh (a real, existing script in this same
directory) calls
05_full_model_quantize.py --assemble-from-resume
directly, bypassing run_full_yaqa.sh
entirely, and would have silently skipped the old hook. Moved the
extraction into 05_full_model_quantize.py’s own
main() instead — the one real code path every real
assembly, from either script, actually runs through.
run_full_yaqa.sh no longer has (or needs) its own separate
extraction step; the assembly call it already makes now produces the
score file as a side effect.08_extract_hessian_scores.py, a real,
tested, documented script — not a one-off Python snippet anymore, and
reused (not reimplemented) by the automatic hook above.Exact real command (this is what run_full_yaqa.sh now
runs automatically, and what you’d run by hand against any older,
already-completed build):
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
extract \
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume \
--output /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.jsonReal fix, 2026-09-16: this script is now the
consolidated toolkit (§11.9) — the extract subcommand above
is required (the bare form without it now errors).
The real, full lifecycle of this exact command — same command, two different moments, two different answers:
| Moment | Is this command correct? | Why |
|---|---|---|
| Day zero — no master score file exists yet | Yes — this exact command | This is precisely what creates the file for the first time, from one finished build’s real logs. |
Any time after — the file already has real data
(e.g. from a later watch run against a different
build) |
No — do not re-run this | extract reads only the one resume-dir you point it at
and overwrites the output. Real tensors from any other
build (only watch adds those) aren’t in what it just read,
so they’re gone from what it just wrote. Use watch instead
(§11.9) — it reads the existing file first and only ever adds. |
This project’s own score file already had its real day zero before
this section was written — so today, the correct command for it is
always watch, never this one.
Verified byte-for-byte identical to the original manual extraction
this session. Real, already-generated output, kept permanently in the
repo (not a scratch file):
research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json
— 360 real tensors scored (every real YAQA-corrected
trunk tensor from that build; lm_head and any bits=16
tensor are absent from this file by real construction, not by omission —
see 11.4). Real scope note: this score is a property of the model’s own
real weights and calibration data, not of which plan built it — one real
extraction is reusable for any future plan against the same model, it
does not need to be regenerated per plan.
Honesty note first, direct user question, direct
answer: these four runs are NOT a systematic strength sweep
(never ran strength = 1, 3, 4, 5, 6, 7 — only 2.0 and 8.0 were ever
actually run). 2.0 is the flag’s own coded default
(--hessian-weight-strength, default: 2.0) — a reasonable
starting value, never empirically tuned or optimized against any real
eval. 8.0 was one deliberate second data point, chosen to
be “notably higher” specifically to see how the effect scales — not a
search. Real summary table, all four:
| # | Real config (beyond target-bpw/group-size/candidates, held constant) | Strength | Compared against | Real result |
|---|---|---|---|---|
| A | --late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6
(no Hessian file at all) |
— | real shipped production plan | 69/497 differ. lm_head self-promotes
6→16, zero Hessian involved. |
| B | same floors as A +
--hessian-score-file ... --hessian-weight-strength 2.0 |
2.0 | plan A | Only 3/497 differ from A. lm_head stays at 16
(unchanged from A). Real, meaningful move:
layers.55.self_attn.k_proj: 16→4 — matches the original
role-based “safe to compress” recommendation cascaded-KL alone had
refused. |
| C | same floors as A +
--hessian-score-file ... --hessian-weight-strength 8.0 |
8.0 | plan A | 40/497 differ from A. Broad promotion of layers 34–62 to full
precision. lm_head drops back down from 16 to
exactly 6 — cranking the signal too far dilutes its own
relative specialness. |
| D | --late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 --no-protect-first-last
+
--hessian-score-file ... --hessian-weight-strength 2.0 |
2.0 | real shipped production plan | 34/497 differ. lm_head self-promotes
again, 6→16, this time with real regional floors still
active. |
Experiment A — pure cascaded KL, no floors at all, no Hessian signal. The real zero-Hessian baseline:
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
--target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \
--late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \
--output 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/plan_v4_round1_NOFLOORS.jsonExperiment B — add the real Hessian signal at the default strength (2.0), floors still off:
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
--target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \
--late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \
--hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--hessian-weight-strength 2.0 \
--output /tmp/plan_v4_round1_NOFLOORS_HYBRID.jsonExperiment C — same, but strength cranked to 8.0 (the real cautionary result):
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
--target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \
--late-attn-min-bits 0 --early-qkv-min-bits 0 --lm-head-min-bits 6 \
--hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--hessian-weight-strength 8.0 \
--output /tmp/plan_v4_round1_NOFLOORS_HYBRID_STRONG.jsonExperiment D — the actual “hybrid” configuration: production floors kept ON, only the flat 100× boundary category replaced by the real Hessian score:
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
--target-bpw 5.028075122774708 --group-size 64 --candidates 4,5,6,8,16 \
--late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 \
--no-protect-first-last \
--hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--hessian-weight-strength 2.0 \
--output 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/plan_v4_round1_HESSIAN_HYBRID.json--no-protect-first-last is the real, pre-existing flag
that disables the flat 100× first/ last-layer boundary category
(identify_boundary_layers());
--hessian-score-file supplies the real, continuous
replacement. --late-attn-min-bits 5 --early-qkv-min-bits 6
are left exactly at the real production values — not
disabled here, see 11.4 for why.
Real result: 34 of 497 tensors differ from the real shipped
production plan. L0/L63 MLP tensors
(previously locked at 16-bit purely by the flat boundary rule) drop to
their real curvature-justified bits (mostly 4-bit); lm_head
again self-promotes, this time all the way to 16-bit; several mid-depth
self_attn.k_proj/v_proj tensors that the
boundary rule never touched settle exactly at the real late-attention
floor (5-bit) instead of being accidentally over-protected. Same real
BPW budget throughout — a reallocation, not “spend more.”
Keep --late-attn-min-bits and
--early-qkv-min-bits at their real production defaults (5
and 6). Do not disable them. Real reasoning, not a guess:
identify_output_layers’s output/embedding matches plus
identify_boundary_layers’s first/last-layer catch-all) is a
genuinely different, weaker case than the region floors — it was never
motivated by a specific measured regression the way early-QKV and
late-attention were; it was “protect the edges” on general principle.
That is exactly the kind of coarse, unfounded-in-curvature category this
Hessian signal is built to replace, and Experiment D shows it doing so
with a legible, explicable result. This one is the real,
justified swap — real curvature data replacing a heuristic that
was never backed by a specific incident.lm_head’s hard floor
(--lm-head-min-bits 6) is never touched by
any of this, regardless of Hessian settings. It is structurally exempt:
lm_head’s own two-sided Hessian would need 246.7GB to
compute and is never built, so it has no real score in the file at all
(confirmed: 0 entries for language_model.lm_head in the
real score file). A hard floor is the only thing that has ever reliably
protected it (see section 9) — there is no real curvature signal here
for the Hessian mechanism to use even if you wanted it to.Concrete, current recommendation: run production
builds with the region floors ON (unchanged),
--lm-head-min-bits 6 ON (unchanged), and
--no-protect-first-last --hessian-score-file ... --hessian-weight-strength 2.0
as the one real, justified substitution — replacing only the boundary
category’s flat 100×, nothing else. This has not yet been used to build
a real shipped model; do that (and run the real IFEval/MMLU regression
checks against it) before making it any script’s own default.
The real formula, in three steps, exactly as implemented in
apply_hessian_weighting():
hess_score, most dangerous (lowest score) to safest
(highest score).r from 0 (most dangerous) to 1 (safest):
r = rank_index / (360 - 1).factor = 1.0 + strength × (1 − r).That is the entire mechanism. Real numbers, pulled directly from the
real 360-tensor score file, at strength=2.0:
| Real tensor | Real hess_score | Percentile rank r |
Real weight factor |
|---|---|---|---|
layers.63.mlp.up_proj (most dangerous of all 360) |
0.000161 | 0.000 | 3.000× |
layers.25.mlp.up_proj |
0.000499 | 0.251 | 2.499× |
layers.15.self_attn.o_proj (dead middle) |
0.001130 | 0.501 | 1.997× |
layers.36.mlp.up_proj |
0.002000 | 0.752 | 1.496× |
layers.55.self_attn.k_proj (safest of all 360) |
0.012979 | 1.000 | 1.000× |
Five genuinely different real weights, out of 360 total — every
scored tensor gets its own precise number. The old flat system would
have given all five of these 1.0 (none of them are named
lm_head or a boundary layer), blind to the fact that
layers.63.mlp.up_proj is objectively far more dangerous
than layers.36.mlp.up_proj by real measured curvature.
| Flat 100× category weight | Real Hessian percentile weight | |
|---|---|---|
| What decides the multiplier | A binary category match (name pattern / depth position) | This tensor’s own real measured curvature |
| Granularity | Same multiplier for every tensor in the category | Continuous, unique per tensor |
| Can it protect a tensor the category misses? | No — if it’s not in the category, it gets 1.0× regardless of real risk | Yes — any of the 360 real scored tensors, category or not |
| Can it over-protect a tensor that doesn’t need it? | Yes, systematically — every category member gets full weight even if
genuinely safe (real proof: L0/L63 MLP tensors
sat at 16-bit for no curvature reason) |
No — the safest real tensor gets exactly 1.0×, no boost at all |
Works for lm_head? |
Yes, but insufficient alone (real, measured: still lands at Q4 despite 100×) | No real data exists — structurally exempt from this whole mechanism |
| Real, current status | Still governs the region floors’ own logic (untouched) | Opt-in replacement, boundary category only, not yet a shipped default |
The real gap, confirmed by direct audit: of the 497
real solver targets, only 360 have a real Hessian score. 137 do not —
136 of those (lm_head is the 137th, and stays permanently
excluded, see below) simply because the build that produced
hess_scores.json held them at bits=16 in its
own plan, so YAQA’s real correction pass — the thing that prints the
effective-rank data this whole mechanism is built from — never ran on
them. This is a real, structural chicken-and-egg problem: the tensors an
old protection rule kept at full precision are exactly the tensors with
zero real data on whether that protection was ever justified.
Real breakdown of the 136 missing (excluding lm_head),
by tensor family:
| Family | Count |
|---|---|
in_proj_a |
48 |
in_proj_b |
48 |
v_proj |
10 |
k_proj |
9 |
out_proj |
6 |
in_proj_qkv |
4 |
in_proj_z |
4 |
down_proj |
2 |
up_proj |
2 |
o_proj |
2 |
gate_proj |
1 |
Why lm_head stays excluded, permanently, by
direct decision (not a gap to fix): a real one-sided GPTQ
correction was already tried on it — it never rounded better than the
naive fallback, needed heavy real padding, and lm_head is
already known-sensitive from the literature. The real decision: keep it
at its naive floor (6-bit) or BF16, never attempt YAQA/GPTQ correction
on it. This script never touches lm_head, under any
circumstance.
Why a real build is needed even though the Hessian math
itself doesn’t depend on bit-width:
H_I/H_O are built from real calibration
activations, computed before any rounding decision — they don’t know or
care what bit-width you’re about to pick. But
05_full_model_quantize.py’s own real tensor-selection gate,
load_plan_targets(), skips any tensor a plan marks
bits=16 — real correction (and the effective-rank print
that is its free byproduct) never runs on it. That’s a real, deliberate
pipeline shortcut (why spend real GPU time on a tensor you’re not
touching), not a mathematical requirement — but it’s the gate that
exists, so it has to be satisfied. Confirmed directly: even
--godmode-multi-bit-checkpoint (the real, already-existing
multi-bit sweep that reuses one real collected Hessian across every
candidate in {2,3,4,5,6,8} without re-running calibration)
inherits this same gate — it does not bypass it.
Step 1 — build a real “probe” plan. New, tested
script: scripts/yaqa_port/10_build_hessian_probe_plan.py.
Takes a real, already-built plan and forces only the currently-unscored,
non-lm_head tensors to a single real candidate bit-width
(default 8 — the safest real choice; its only job is to be
!= 16 so the tensor clears the pipeline’s selection gate).
Every other tensor, including lm_head, is left completely
untouched.
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/yaqa_port/10_build_hessian_probe_plan.py \
scripts/yaqa_port/research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/plans_P1_both_off/plan_baseline_noHessian.json \
--hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--probe-bits 8 \
--output 04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.jsonReal, verified output (ran this exact command):
136 real tensor(s) overridden to 8-bit (everything else, including lm_head, left untouched).
Step 2 — run the real build with godmode multi-bit sweeping
enabled, so every probed tensor’s real effective rank is
captured at every real candidate in {3,4,5,6,8} in one real
pass (not just at the single forced probe bit):
Corrected 2026-09-15 — this originally omitted the
required cd and used a plan path only correct from the
project root; ./run_full_yaqa.sh only runs from inside
scripts/yaqa_port/, two real directories deeper than the
project root. That relative-path version then broke a second way in real
use: pasted as a multi-line command, the \
line-continuations kept getting lost in transit, silently splitting one
command into several. Real, permanent fix: a saved script with every
path absolute, nothing left to paste or retype —
scripts/yaqa_port/run_hessian_probe.sh:
/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/run_hessian_probe.shThat’s the whole command — no flags, no cd first, run
from anywhere. Its real, fixed contents (everything below already inside
the script):
/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/run_full_yaqa.sh \
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE \
--skip-lm-head-gptq \
--plan /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.json \
--n-calibration 4 \
--incoherence none \
--correct-mtp \
--godmode-multi-bit-checkpoint \
--godmode-candidate-bits 3,4,5,6,8This is a real, multi-hour GPU job — only 136 of the 497 real tensors need correction here (everything already scored keeps its cached real result if run against the same resume-dir lineage; confirm real resume-dir semantics before assuming full reuse), but real godmode sweeping across 5 candidates per tensor is real, additional compute per tensor, not free.
Real, verified fix required before this step works
(2026-09-14, found and fixed the same day): godmode’s own
per-candidate correction calls originally passed
verbose=False (real code,
05_full_model_quantize.py), which silently suppressed the
effective-rank print this whole mechanism depends on — its richer data
only ever reached the separate
godmode_sensitivity_checkpoint.jsonl file, which
08_extract_hessian_scores.py never reads. Fixed to
verbose=True. That alone was not enough: the real label
passed to each godmode call was
f"{name}[godmode bits={cand_bits}]" — checked directly
against hessian_story_lib.RANK_RE, and the space inside
"[godmode bits=" breaks the regex’s name capture entirely
(confirmed: zero match, not a garbled one). Real fix: the label is now
the clean tensor name alone. This loses nothing real —
effective_rank(H_I)/(H_O) does not depend on
candidate bit-width (established earlier: the Hessian comes from real
calibration activations, before any rounding decision), so every real
godmode candidate for the same tensor prints the identical real number
under the same real name; load_hessian()’s
overwrite-on-duplicate-key behavior is a correct no-op here, not data
loss. cand_bits itself is preserved separately, as
godmode_candidates’s own real dict key. Both fixes are
already in 05_full_model_quantize.py — verified via a
direct regex test against the real corrected print format before this
section was written.
Step 3 — extract and merge the new real scores. As
of 2026-09-16 this is one real command, not a manual two-step (extract,
then hand-merge) — the watch subcommand does both, safe to
run live while this build is still in progress:
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
watch \
<this build's real resume-dir> \
/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--interval 30Only ever adds real, newly-available tensors (or updates changed ones
with --overwrite) — never removes or silently clobbers an
existing real entry. Real, permanent output either way:
scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json.
Real end state once this completes: 496 of 497 real
tensors scored (only lm_head permanently excluded) — the
region floors (--late-attn-min-bits,
--early-qkv-min-bits) can then be turned off for real, with
every tensor they used to blanket-protect instead getting real,
individual, Hessian-informed protection — the actual, original goal this
whole investigation was aimed at.
Real problem this closes: §11.5/§11.6 above showed
the soft percentile weight is a real, continuous, per-tensor signal — a
genuine improvement over the flat 100× category weight. But a direct
check found it is still not a guarantee: even the single most
dangerous real tensor in the whole model
(layers.63.mlp.up_proj) only reached 8-bit in the real
production plan, and pushing the multiplier to
strength=1000 (500× stronger than the tested default)
changed zero outcomes in the top 20 most dangerous real
tensors. The reason: these tensors’ real, raw KL cost is close enough to
flat across candidate bit-widths that no multiplier of an already-small
number wins against tensors with genuinely large KL cost, under the same
fixed BPW budget. A soft weight cannot fix a signal that’s fundamentally
too small to matter in the objective it’s added to.
The fix is two new, independent, opt-in flags — neither changes existing behavior when omitted (verified directly: identical output, 0/497 tensors differ, with both flags left out):
--hessian-floor-tiers "0.05:8,0.15:6" —
a real, per-tensor MILP hard floor, computed from each tensor’s
real Hessian percentile rank (same ranking as §11.5). Any real tensor
whose rank falls under 0.05 gets a genuine minimum of 8-bit; under 0.15
gets a minimum of 6-bit — enforced the exact same way
--lm-head-min-bits/--boundary-min-bits already
work (add_min_bits_constraint, a real
LinearConstraint, not a weight). The solver remains free to
go above the floor — real, confirmed: of the 54 tensors floored
in the tested run, 29 ended up at full 16-bit anyway. The real
feasibility ceiling for this exact production budget (5.028075 BPW) was
found by bisection: coverage beyond ~18.7% of scored tensors (67 of 360)
makes the solve genuinely infeasible; the tested two-tier spec (54
tensors) sits safely under that.
--hessian-primary (+
--hessian-primary-gap, default 1e-7) — a real,
lexicographic two-phase solve. Phase 1 solves a new MILP minimizing
danger-weighted bit deficit, KL playing no role at all; its result is
pinned (to a provably tight gap, not an accepted approximation —
confirmed gap=0.0 on the tested run) before the existing
weighted-KL phase runs, now constrained to only Phase-1-optimal
solutions. This is the same two-phase mechanism this file already uses
for --raw-tiebreak, just with Hessian danger promoted to
the primary objective and KL demoted to tie-break. Real, measured
consequence: KL ends up deciding 0 of the 360 real
Hessian-scored tensors — Phase 1’s ranking is precise enough that ties
among them are effectively never real — and only 3 of the 497 total
(both unscored, no Hessian data at all).
Floor alone and lexicographic alone each fix half the real gap;
combined, they dominate both individually on every real metric checked
(kept-at-16-bit and cut-to-≤5-bit, across every danger tier from top-10
to top-108) — see research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/MILP_AWARE_HESSIAN_GUIDE.html
for the full real comparison.
The exact real command, everything else matching the
already-benchmarked production checkpoint
(--max-low-bit-run 999 is required alongside
--hessian-primary — without it, the separate run-guard
mechanism piles extra constraints on top each iteration and the combined
solve goes infeasible; unrelated to the Hessian mechanism itself):
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
--target-bpw 5.028075122774708 \
--candidates 4,5,6,8,16 \
--pareto none \
--group-size 64 \
--late-attn-min-bits 0 \
--early-qkv-min-bits 0 \
--lm-head-min-bits 6 \
--boundary-min-bits 6 \
--hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--hessian-weight-strength 0 \
--hessian-floor-tiers "0.05:8,0.15:6" \
--hessian-primary \
--max-low-bit-run 999 \
--output <your_output_path>.json
How these flags actually pick bit-widths (added
2026-09-19). The rank, the danger weight, the floor tiers and
the two-phase solve are explained step by step, with diagrams and
checked numbers, in How
the Solver Turns a Hessian Score into a Bit-Width. The one-line
version: under --hessian-primary, spare bits above the
floors go to the tensors with the most danger per million parameters, so
a tiny nearly-safe tensor can reach 16-bit while the single most
dangerous (large) tensor stays at its floor, and tensors with no score
are pushed down to their floors.
--hessian-unscored-fallback {zero,kl} (added
2026-09-19, default zero = unchanged behaviour).
Under --hessian-primary, a tensor with no score in the
Hessian score file gets danger weight 0, and on that scale 0 is the
safest end, so it is under-protected: on the 407-scored
snapshot 64 of the 90 unscored tensors ended at 4-bit where a KL-driven
plan keeps 89 of them at 16-bit.
--hessian-unscored-fallback kl gives each unscored tensor
the percentile rank of its own isolated KL at the lowest candidate bit
(over all tensors, 1.0 = most dangerous) as its danger weight instead.
Scored tensors and every floor are unchanged; it fades out as the sweep
scores more tensors. Measured: 84 of 90 unscored tensors at 16-bit, 8
scored tensors move, 3 fewer scored tensors reach 16-bit, raw summed KL
50.09 to 39.41 (production 37.45). Full explanation and evidence: How
the Solver Turns a Hessian Score into a Bit-Width, section 5.6. The
complete, exact command with the fallback (recommended while coverage is
incomplete; copy and paste as is, it writes one new plan file and builds
nothing):
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
--target-bpw 5.028075122774708 \
--candidates 4,5,6,8,16 \
--pareto none \
--group-size 64 \
--late-attn-min-bits 0 \
--early-qkv-min-bits 0 \
--lm-head-min-bits 6 \
--boundary-min-bits 6 \
--hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--hessian-weight-strength 0 \
--hessian-floor-tiers "0.05:8,0.15:6" \
--hessian-primary \
--hessian-unscored-fallback kl \
--max-low-bit-run 999 \
--output 04_LANGUAGE_PLANS/HESSIAN_TEST/plan_floor_primary_klfallback_2026-09-19.json
Independent code review (2026-09-15) confirmed the core design
correct and properly gated, and found two real, since-fixed issues: the
floor’s own audit field originally reported requested tiers rather than
actually-shipped bits (fixed — now reports real shipped bits plus an
explicit violations list, [] on the tested
run), and Phase 1’s gap tolerance was originally looser than the pin’s
own framing implied (fixed via --hessian-primary-gap,
default 1e-7, independent of the general
--quality-mip-rel-gap).
--pareto none is
required whenever a Hessian file drives allocation
--pareto defaults to meaningful, which removes
any candidate bit-width a tensor’s own isolated-KL data says
isn’t a meaningful improvement over a cheaper one — a real, useful
filter when KL is the only signal deciding allocation, but it has no
idea a Hessian floor is about to demand a specific bit-width for a
completely different, KL-independent reason. Real, checked at 405/497
coverage: 55 real tensors had their
Hessian-floor-required bit filtered out by exactly this mechanism,
forcing several to jump all the way to 16-bit instead of the tier they
actually needed, and in some real runs this combined with
--hessian-primary’s own tight re-pinning to produce genuine
MILP infeasibility (HiGHS Status 8: Infeasible) — worse as
real coverage grows, not better. --pareto none fixes it:
confirmed real re-runs of the same production budget, previously-failing
configurations solved cleanly once added.
Separately, --hessian-weight-strength was directly,
isolatedly tested (2026-09-18): re-running the identical command with
only that flag toggled between its default (2.0) and
0, --pareto none held constant, produced
bit-for-bit identical results — the soft weight changes
nothing once the floor and/or --hessian-primary are active.
It was only ever inflating the reported OptiQ-weighted ΣKL
statistic (by up to 3× per tensor), never the real allocation. Full real
evidence: MILP_AWARE_HESSIAN_GUIDE.html.
The toolkit used to be three separate scripts in two
different directories. That was a real, reported problem
(“scripts all over the place… nightmare to keep track”) — now
consolidated into one file, one location, three subcommands. The two old
scripts (08b_merge_live_hessian_scores.py,
08c_build_hessian_hybrid_checkpoint.py) no longer exist;
everything they did now lives here:
# One-shot: extract scores from a finished build's own logs
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
extract \
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume \
--output /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json
# Keep growing that same file live, from any running job (Ctrl+C to stop; add --no-loop for one pass)
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
watch \
<resume_dir> \
/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--interval 30
# Merge a real KL checkpoint with a real godmode sweep into one file the
# UNMODIFIED MILP solver (section 3.3 below is unaffected) already reads natively
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
hybrid \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
<resume_dir>/godmode_sensitivity_checkpoint.jsonl \
--output /tmp/hessian_hybrid_checkpoint.jsonThis exact filename and location (scripts/yaqa_port/,
not the research_hadamard_blowup/ subfolder) is
load-bearing — 05_full_model_quantize.py imports it by file
path for its own automatic end-of-build hook. Do not rename or relocate
it without also updating that hook.
The real question this whole toolkit exists to answer: does
the free Hessian score agree closely enough with the expensive
cascaded-KL measurement to replace some of it? Investigated
directly tonight, real data throughout. Honest summary — full trail,
diagrams, and every command:
research_hadamard_blowup/exports/HESSIAN_MASTER_REFERENCE.html.
pct_reduction in the manifest)
— much stronger than anything found against KL. Current read: Hessian
looks like a strong signal for “how much does curvature-aware correction
help this tensor,” not yet a strong signal for “how much does this
tensor’s quantization hurt the model’s output.”PORT_LEDGER.md)Real per-layer streaming architecture (avoiding repeated full-model passes per batch) — a real efficiency opportunity, not a correctness gap.
The real loader for YaqaQuantizedLinear (needed
before --incoherence hadamard is actually
runnable).
Real end-to-end model quality evaluation (perplexity/KL vs. GPTQ/stock/HybridPareto).
See PORT_LEDGER.md for the complete, detailed
engineering log, and CHANGELOG.md for the short
chronological list of every real change.
© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.