2026-10-01
What this is: the real, exact steps to build the
first model from the “pure sensitivity” plan — bit allocation driven
entirely by the GODMODE Hessian sweep’s real per-bit-width error curves,
no --hessian-primary aggregate layered on top (see this
project’s own research: the aggregate mechanism’s linear deficit model
can’t see the diminishing returns a real curve shows, and empirically
produced a worse real KL-divergence sum than the plain
pre-Hessian baseline — 40.62 vs 37.45 raw Σ
isolated KL, same source data, directly comparable).
Already built:
04_LANGUAGE_PLANS/HESSIAN_TEST/plan_godmode_ONLY.json — 497
tensors, built from hessian_hybrid_checkpoint.json (the
full GODMODE sweep, 496 real per-tensor curves + 1 KL-fallback entry for
lm_head, which has no real Hessian coverage at all — see
FLAT_MODE.html for why that specific tensor can’t be
measured). --lm-head-min-bits 8 and
--boundary-min-bits 6 are the only manual floors;
everything else is driven by the real curve data directly.
A resume seed reuses already-corrected tensors from your prior builds that happen to match this plan’s bit assignments, so the real, slow per-tensor Hessian correction only has to run on tensors that aren’t already done at the right width.
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
# Check real coverage first, copies nothing:
python3 build_resume_seed.py \
--plan ../../04_LANGUAGE_PLANS/HESSIAN_TEST/plan_godmode_ONLY.json \
--output /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-GODMODE-v1.yaqa_resume \
--dry-run
# Actually build it:
python3 build_resume_seed.py \
--plan ../../04_LANGUAGE_PLANS/HESSIAN_TEST/plan_godmode_ONLY.json \
--output /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-GODMODE-v1.yaqa_resumeAlready done once (2026-10-01): 287 of 497 tensors (57.7%) reused
from three real prior builds, scored by real measured correction quality
(err_yaqa), not just whichever folder was scanned first.
11.61 GB.
The script refuses to run (outside –dry-run) if
–output already exists, specifically so it never silently
merges into or overwrites a real resume directory. To test it fresh,
point –output at a new path — the existing
GODMODE-v1.yaqa_resume isn’t in the way, and deleting it
would just throw away 11.61 GB of real, reusable, already-corrected
tensors for no benefit.
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
./run_full_yaqa.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-GODMODE-v1 \
--source /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara \
--plan ../../04_LANGUAGE_PLANS/HESSIAN_TEST/plan_godmode_ONLY.json \
--n-calibration 4 --incoherence nonerun_full_yaqa.sh’s own real default
--source is already this exact path — confirmed directly in
its source
(SOURCE_DIR="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara")
— passing it explicitly here is just for clarity, not required.
It auto-detects <output>.yaqa_resume next to the
output path you give it — never pass --resume-dir yourself
for a normal run.
Never call 05_full_model_quantize.py directly for a real
multi-batch run — only run_full_yaqa.sh, which runs each
memory-bounded batch as its own fresh process. A direct call risks a
real, confirmed Metal resource-handle crash.
./monitor.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-GODMODE-v1.yaqa_resume/logs/batch_1.logAlso serves a live HTML dashboard
(generate_dashboard.py) with real per-tensor progress and
real CPU/memory stats.
run_full_yaqa.sh finds the resume folder
automatically and continues from there.MODEL="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-GODMODE-v1"
mtplx inspect "$MODEL" --json
mtplx forge verify "$MODEL" --stamp --json