2026-09-04
Real path:
/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/
Real repo:
github.com/hakim2206/YAQA_MLX_Vegalaboratories_LTD
PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port"$PYBIN" test_regression.pymlx_lm-loadable)./run_full_yaqa.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw \
--n-calibration 4 --incoherence noneThis is the command to actually produce a real, deployable, standard-loadable YAQA model. It runs every real plan tensor in memory-bounded batches, each as its own separate process (required — a real Metal resource-handle limit crashes if too many batches run inside one long-lived process), then assembles the final real save.
./monitor.sh <output_dir>.yaqa_resume/logs/batch_<N>.logAuto-refreshing, color-coded (green = running, red = stopped/crashed,
yellow = low memory), real ps/log data only — no simulated
progress.
Real command, exact, currently running as of this document:
./run_full_yaqa.sh /Users/hghelab/.mtplx/models/yaqa_orchestrator_test \
--subset default --hessian-batch-budget-gb 0.5Purpose: forces the 3-tensor smoke-test subset into 3 separate batches (tiny budget) to prove the separate-process fix for the Metal resource-limit crash actually works, before trusting it for the real multi-hour full run above.
The trunk-quantization run above does not touch the MTP
speculative-decode sidecar’s own correction quality — by default it’s
naive-quantized (mx.quantize, no Hessian correction).
--correct-mtp fixes that. If a trunk build already
completed (its .yaqa_resume cache still on disk), this
reuses that cache instead of recomputing the whole multi-hour run — only
reassembly + the new MTP step actually run.
Run these yourself, one at a time — this is a real production launch, always user-triggered, never run automatically on your behalf.
Step 1 — copy the existing resume cache under a new output name (non-destructive; does not touch the existing completed model at all):
OLD="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32"
NEW="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected"
cp -a "${OLD}.yaqa_resume" "${NEW}.yaqa_resume"Real pitfall, hit directly (2026-09-08): only run this
once. cp -a src dst behaves differently depending
on whether dst already exists – if it does (e.g. this was
already run once before), cp nests src
inside dst as a subdirectory instead of refreshing
dst’s contents, silently doubling real disk usage
(confirmed: 16GB -> 32GB, the entire original resume dir duplicated
one level down inside the new one) with no error or warning. Before
re-running Step 1, check whether ${NEW}.yaqa_resume already
exists (ls -d "${NEW}.yaqa_resume") – if it does, either
reuse it directly (skip Step 1, go straight to Step 2) or
rm -rf it first and only then re-run the copy.
Step 2 — launch the build (long-running; run in its own terminal tab so you can watch it directly and stop it whenever you want):
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
./run_full_yaqa.sh "$NEW" --n-calibration 4 --incoherence none --correct-mtp --reuse-cached-fallbackSkips every already-cached trunk tensor and the already-cached
lm_head correction (both resume-skip off the same manifest), goes
straight to reassembling the model, then runs the new
correct_mtp_sidecar() pass and overwrites the naive MTP
sidecar with a real YAQA-corrected one.
--reuse-cached-fallback (added 2026-09-08, real user
request): lm_head’s one-sided real GPTQ correction is a deterministic
search over a fixed damping-ratio schedule against a fixed real Hessian
– if a prior run already exhausted it and recorded
method: "naive_fallback" in the manifest (every real
damping ratio failed the safety gate), retrying it will fail identically
again, just burning ~15-20 real minutes (24 Hessian chunks + damping
trials) to reach the same known result. This flag makes the existing
resume-skip in 06_lm_head_gptq.py also accept a cached
naive_fallback result (still checked against the current
plan’s bits/group_size, same as it already does for a successful
"gptq" result) and skip straight past it. Off by default –
opt in explicitly, since a genuinely different
--n-calibration/--source could change the real
outcome and this flag would then hide that. lm_head’s real fallback
quantizes at the plan’s own bits/group_size regardless of method
(6-bit/G64 for this project’s plan) – naive_fallback only
means “no Hessian-based correction,” never “left at BF16.”
Step 3 — watch it live, in a second terminal tab:
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
./monitor.sh "${NEW}.yaqa_resume"Also serves two live dashboards (auto-refresh, no flicker): the
internal one at http://127.0.0.1:8765/dashboard.html and
the customer-facing one at
http://127.0.0.1:8766/live-status.html — both show a
YAQA-correcting MTP sidecar milestone step (added
2026-09-08; see CHANGELOG.md).
--mtp-bits/--mtp-group-size are
deliberately omitted from Step 2 — they default to the plan’s own native
MTP bit-width, so the first real run is bit-for-bit naive-vs-corrected,
not confounded by also changing the target bit-width.
--help directly)05_full_model_quantize.py| Flag | Default | Purpose |
|---|---|---|
--source |
this project’s real model | Real BF16 source model path or HF repo id. |
--plan |
this project’s real plan JSON | Real per-tensor plan JSON path — a different --source
almost certainly needs its own plan too. |
--lm-head-name |
auto-detect | Real dotted name of the vocab-output tensor to exclude from YAQA’s two-sided correction. Only pass explicitly if auto-detection raises. |
--subset |
full plan (363 tensors) | default = 3-tensor smoke test, or a comma-separated
list of real tensor names. |
--output |
none (comparison-only) | Real output directory. If given, Part 5c runs and actually saves a real model; if omitted, only the rotated-vs-unrotated comparison metrics print, no save. |
--n-calibration |
1 | Real stratified sequences PER DOMAIN. GPTQ’s own project default is 4. |
--seq-len |
128 | — |
--batch-size |
1 | Real sequences per forward+backward call during Hessian collection;
raising --n-calibration does not require raising this. |
--fallback-bits |
6 | Bit-width for any leaf module the plan doesn’t cover. |
--fallback-group-size |
64 | Paired with --fallback-bits. |
--resume-dir |
none | Real per-tensor result cache dir. Required for
--batch-only/--assemble-from-resume. |
--batch-only N |
none | Process only the Nth batch (0-indexed, from the current remaining
set), write to --resume-dir, exit. |
--print-num-batches |
off | Print the real remaining batch count and exit. |
--assemble-from-resume |
off | Skip Hessian collection — load every cached tensor result and run the real final save. |
--hessian-batch-budget-gb |
8.0 | Real memory budget (GB) for H_I+H_O held per batch — all 363 targets at once need 557GB, doesn’t fit in 128GB. |
--incoherence {none,hadamard} |
none |
none = production default, mlx_lm-loadable
today. hadamard = research branch, needs a loader that
doesn’t exist yet. |
--sigma-reg |
1.0 | Damping constant, empirically tuned at the old single-sequence calibration scale — may need re-tuning at a richer scale. |
--correct-mtp |
off | Real YAQA correction for the MTP sidecar, run after the naive sidecar is written. No-op if the source model has no MTP tensors. |
--mtp-bits {2,3,4,5,6,8} |
plan’s own native rule | Target bit-width for the corrected MTP sidecar. |
--mtp-group-size |
plan’s own group_size | Paired with --mtp-bits. |
--reuse-cached-fallback |
off | Real no-op here — exists only so
run_full_yaqa.sh forwarding this flag to every invocation
doesn’t crash this script with “unrecognized arguments.” Its real effect
is in 06_lm_head_gptq.py below. |
06_lm_head_gptq.py| Flag | Default | Purpose |
|---|---|---|
--source |
this project’s real model | Must match the main run’s --source —
run_full_yaqa.sh forwards it automatically. |
--plan |
this project’s real plan JSON | Must match the main run’s --plan. |
--lm-head-name |
auto-detect | Same real dual-signal detection as the main script; deterministic, so both scripts agree independently even without an explicit override. |
--resume-dir |
required | Same .yaqa_resume dir the main run’s manifest lives in
— adds exactly one entry, never touches the other 362. |
--n-calibration |
4 | Should match the main run’s value for a fair, consistent calibration scale. |
--seq-len |
128 | — |
--batch-size |
1 | — |
--reuse-cached-fallback |
off | See the dedicated section above — accepts a cached
naive_fallback result (not just a successful
gptq one) and skips the real damping-search retry. |
02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py
— MILP-aware Hessian flags (2026-09-15)Full real mechanism, evidence, and worked comparison:
RUNNING_GUIDE.md §11.8. Both flags are opt-in and change
nothing when omitted (verified: identical output, 0/497 tensors
differ).
| Flag | Default | Purpose |
|---|---|---|
--hessian-floor-tiers |
none (disabled) | Real, per-tensor MILP hard floor, e.g. "0.05:8,0.15:6"
— any tensor in the real most-dangerous 5% (by Hessian percentile rank)
gets a genuine minimum of 8-bit, under 15% gets a minimum of 6-bit. A
real constraint, not a weight — the solver cannot trade it away. Solver
remains free to go higher (real, confirmed: 29 of 54 floored tensors
reached 16-bit anyway). |
--hessian-primary |
off | Real, lexicographic two-phase solve: Phase 1 solves on Hessian danger alone (KL plays no role), pins that result, then Phase 2 (the existing weighted-KL objective) only picks among Phase-1-tied solutions. Real, measured: KL ends up deciding 0 of 360 scored tensors, and only 3 of 497 total. |
--hessian-primary-gap |
1e-7 |
How tight Phase 1 must solve before its result is pinned —
deliberately far tighter than --quality-mip-rel-gap, so the
pin is a provably exact answer, not an accepted approximation. |
--hessian-unscored-fallback |
zero |
Added 2026-09-19. What --hessian-primary does with
tensors that have no score in the Hessian score file. zero
(unchanged behaviour): danger weight 0.0, the SAFEST end of the scale,
so they are pushed to their floors. kl: each unscored
tensor gets the percentile rank of its own isolated KL at the lowest
candidate bit, over all tensors (1.0 = most dangerous), so silence does
not mean safest. No effect on floors or scored tensors; fades out as
coverage grows. |
Real command, everything else matching the already-benchmarked
production checkpoint (note: --max-low-bit-run 999 is
required alongside --hessian-primary, or the separate
run-guard mechanism piles extra constraints on top each iteration and
the combined solve goes infeasible):
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
--target-bpw 5.028075122774708 \
--candidates 4,5,6,8,16 \
--pareto none \
--group-size 64 \
--late-attn-min-bits 0 \
--early-qkv-min-bits 0 \
--lm-head-min-bits 6 \
--boundary-min-bits 6 \
--hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--hessian-weight-strength 0 \
--hessian-floor-tiers "0.05:8,0.15:6" \
--hessian-primary \
--max-low-bit-run 999 \
--output <your_output_path>.jsonSame command with the unscored-tensor fallback (added 2026-09-19, recommended while the Hessian sweep is incomplete), complete and exact:
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
--target-bpw 5.028075122774708 \
--candidates 4,5,6,8,16 \
--pareto none \
--group-size 64 \
--late-attn-min-bits 0 \
--early-qkv-min-bits 0 \
--lm-head-min-bits 6 \
--boundary-min-bits 6 \
--hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
--hessian-weight-strength 0 \
--hessian-floor-tiers "0.05:8,0.15:6" \
--hessian-primary \
--hessian-unscored-fallback kl \
--max-low-bit-run 999 \
--output 04_LANGUAGE_PLANS/HESSIAN_TEST/plan_floor_primary_klfallback_2026-09-19.json
--pareto none is
required whenever --hessian-score-file is given
The real default (--pareto meaningful) filters candidate
bits using isolated-KL dominance, which can silently remove the exact
bit-width a Hessian floor needs — real, checked: 55 of 497 tensors
affected at 405/497 coverage, sometimes producing genuine MILP
infeasibility (HiGHS Status 8: Infeasible) together with
--hessian-primary’s own re-pinning, and it gets
worse as real coverage grows, not better.
--hessian-weight-strength was separately, directly
isolated-tested (identical command, only that flag toggled) and
confirmed to change zero real output once the floor/primary are active —
set it to 0, it only inflates the reported weighted-KL
statistic otherwise. Full evidence: MILP_AWARE_HESSIAN_GUIDE.html.
How these flags pick bit-widths (added 2026-09-19). Rank, danger weight, floor tiers and the two-phase solve, with diagrams and checked numbers: How the Solver Turns a Hessian Score into a Bit-Width.
mx.quantize non-idempotency bug: found and fixed
(2026-09-04).mx.clear_cache()
tried and confirmed NOT to fix it; separate-process-per-batch fix built,
validation in progress.See PORT_LEDGER.md in this same directory for the
complete, detailed engineering log.
© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.