← Back to index

YAQA MLX Port — Exact Command Reference

Hakim Ghelab, VegaLaboratories LTD

2026-09-04

Real path: /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/ Real repo: github.com/hakim2206/YAQA_MLX_Vegalaboratories_LTD

Environment

PYBIN="/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port

Regression suite (fast, no real model, ~1-2 min)

"$PYBIN" test_regression.py

The real full quantization run (production, mlx_lm-loadable)

./run_full_yaqa.sh /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw \
  --n-calibration 4 --incoherence none

This is the command to actually produce a real, deployable, standard-loadable YAQA model. It runs every real plan tensor in memory-bounded batches, each as its own separate process (required — a real Metal resource-handle limit crashes if too many batches run inside one long-lived process), then assembles the final real save.

Watching a run live

./monitor.sh <output_dir>.yaqa_resume/logs/batch_<N>.log

Auto-refreshing, color-coded (green = running, red = stopped/crashed, yellow = low memory), real ps/log data only — no simulated progress.

Validation test currently in progress (2026-09-04)

Real command, exact, currently running as of this document:

./run_full_yaqa.sh /Users/hghelab/.mtplx/models/yaqa_orchestrator_test \
  --subset default --hessian-batch-budget-gb 0.5

Purpose: forces the 3-tensor smoke-test subset into 3 separate batches (tiny budget) to prove the separate-process fix for the Metal resource-limit crash actually works, before trusting it for the real multi-hour full run above.

Adding YAQA-corrected MTP sidecar to an already-completed build (2026-09-08)

The trunk-quantization run above does not touch the MTP speculative-decode sidecar’s own correction quality — by default it’s naive-quantized (mx.quantize, no Hessian correction). --correct-mtp fixes that. If a trunk build already completed (its .yaqa_resume cache still on disk), this reuses that cache instead of recomputing the whole multi-hour run — only reassembly + the new MTP step actually run.

Run these yourself, one at a time — this is a real production launch, always user-triggered, never run automatically on your behalf.

Step 1 — copy the existing resume cache under a new output name (non-destructive; does not touch the existing completed model at all):

OLD="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32"
NEW="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected"
cp -a "${OLD}.yaqa_resume" "${NEW}.yaqa_resume"

Real pitfall, hit directly (2026-09-08): only run this once. cp -a src dst behaves differently depending on whether dst already exists – if it does (e.g. this was already run once before), cp nests src inside dst as a subdirectory instead of refreshing dst’s contents, silently doubling real disk usage (confirmed: 16GB -> 32GB, the entire original resume dir duplicated one level down inside the new one) with no error or warning. Before re-running Step 1, check whether ${NEW}.yaqa_resume already exists (ls -d "${NEW}.yaqa_resume") – if it does, either reuse it directly (skip Step 1, go straight to Step 2) or rm -rf it first and only then re-run the copy.

Step 2 — launch the build (long-running; run in its own terminal tab so you can watch it directly and stop it whenever you want):

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
./run_full_yaqa.sh "$NEW" --n-calibration 4 --incoherence none --correct-mtp --reuse-cached-fallback

Skips every already-cached trunk tensor and the already-cached lm_head correction (both resume-skip off the same manifest), goes straight to reassembling the model, then runs the new correct_mtp_sidecar() pass and overwrites the naive MTP sidecar with a real YAQA-corrected one.

--reuse-cached-fallback (added 2026-09-08, real user request): lm_head’s one-sided real GPTQ correction is a deterministic search over a fixed damping-ratio schedule against a fixed real Hessian – if a prior run already exhausted it and recorded method: "naive_fallback" in the manifest (every real damping ratio failed the safety gate), retrying it will fail identically again, just burning ~15-20 real minutes (24 Hessian chunks + damping trials) to reach the same known result. This flag makes the existing resume-skip in 06_lm_head_gptq.py also accept a cached naive_fallback result (still checked against the current plan’s bits/group_size, same as it already does for a successful "gptq" result) and skip straight past it. Off by default – opt in explicitly, since a genuinely different --n-calibration/--source could change the real outcome and this flag would then hide that. lm_head’s real fallback quantizes at the plan’s own bits/group_size regardless of method (6-bit/G64 for this project’s plan) – naive_fallback only means “no Hessian-based correction,” never “left at BF16.”

Step 3 — watch it live, in a second terminal tab:

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
./monitor.sh "${NEW}.yaqa_resume"

Also serves two live dashboards (auto-refresh, no flicker): the internal one at http://127.0.0.1:8765/dashboard.html and the customer-facing one at http://127.0.0.1:8766/live-status.html — both show a YAQA-correcting MTP sidecar milestone step (added 2026-09-08; see CHANGELOG.md).

--mtp-bits/--mtp-group-size are deliberately omitted from Step 2 — they default to the plan’s own native MTP bit-width, so the first real run is bit-for-bit naive-vs-corrected, not confounded by also changing the target bit-width.

Every real flag (2026-09-08 — complete, cross-checked against --help directly)

05_full_model_quantize.py

Flag Default Purpose
--source this project’s real model Real BF16 source model path or HF repo id.
--plan this project’s real plan JSON Real per-tensor plan JSON path — a different --source almost certainly needs its own plan too.
--lm-head-name auto-detect Real dotted name of the vocab-output tensor to exclude from YAQA’s two-sided correction. Only pass explicitly if auto-detection raises.
--subset full plan (363 tensors) default = 3-tensor smoke test, or a comma-separated list of real tensor names.
--output none (comparison-only) Real output directory. If given, Part 5c runs and actually saves a real model; if omitted, only the rotated-vs-unrotated comparison metrics print, no save.
--n-calibration 1 Real stratified sequences PER DOMAIN. GPTQ’s own project default is 4.
--seq-len 128 —
--batch-size 1 Real sequences per forward+backward call during Hessian collection; raising --n-calibration does not require raising this.
--fallback-bits 6 Bit-width for any leaf module the plan doesn’t cover.
--fallback-group-size 64 Paired with --fallback-bits.
--resume-dir none Real per-tensor result cache dir. Required for --batch-only/--assemble-from-resume.
--batch-only N none Process only the Nth batch (0-indexed, from the current remaining set), write to --resume-dir, exit.
--print-num-batches off Print the real remaining batch count and exit.
--assemble-from-resume off Skip Hessian collection — load every cached tensor result and run the real final save.
--hessian-batch-budget-gb 8.0 Real memory budget (GB) for H_I+H_O held per batch — all 363 targets at once need 557GB, doesn’t fit in 128GB.
--incoherence {none,hadamard} none none = production default, mlx_lm-loadable today. hadamard = research branch, needs a loader that doesn’t exist yet.
--sigma-reg 1.0 Damping constant, empirically tuned at the old single-sequence calibration scale — may need re-tuning at a richer scale.
--correct-mtp off Real YAQA correction for the MTP sidecar, run after the naive sidecar is written. No-op if the source model has no MTP tensors.
--mtp-bits {2,3,4,5,6,8} plan’s own native rule Target bit-width for the corrected MTP sidecar.
--mtp-group-size plan’s own group_size Paired with --mtp-bits.
--reuse-cached-fallback off Real no-op here — exists only so run_full_yaqa.sh forwarding this flag to every invocation doesn’t crash this script with “unrecognized arguments.” Its real effect is in 06_lm_head_gptq.py below.

06_lm_head_gptq.py

Flag Default Purpose
--source this project’s real model Must match the main run’s --source — run_full_yaqa.sh forwards it automatically.
--plan this project’s real plan JSON Must match the main run’s --plan.
--lm-head-name auto-detect Same real dual-signal detection as the main script; deterministic, so both scripts agree independently even without an explicit override.
--resume-dir required Same .yaqa_resume dir the main run’s manifest lives in — adds exactly one entry, never touches the other 362.
--n-calibration 4 Should match the main run’s value for a fair, consistent calibration scale.
--seq-len 128 —
--batch-size 1 —
--reuse-cached-fallback off See the dedicated section above — accepts a cached naive_fallback result (not just a successful gptq one) and skips the real damping-search retry.

02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py — MILP-aware Hessian flags (2026-09-15)

Full real mechanism, evidence, and worked comparison: RUNNING_GUIDE.md §11.8. Both flags are opt-in and change nothing when omitted (verified: identical output, 0/497 tensors differ).

Flag Default Purpose
--hessian-floor-tiers none (disabled) Real, per-tensor MILP hard floor, e.g. "0.05:8,0.15:6" — any tensor in the real most-dangerous 5% (by Hessian percentile rank) gets a genuine minimum of 8-bit, under 15% gets a minimum of 6-bit. A real constraint, not a weight — the solver cannot trade it away. Solver remains free to go higher (real, confirmed: 29 of 54 floored tensors reached 16-bit anyway).
--hessian-primary off Real, lexicographic two-phase solve: Phase 1 solves on Hessian danger alone (KL plays no role), pins that result, then Phase 2 (the existing weighted-KL objective) only picks among Phase-1-tied solutions. Real, measured: KL ends up deciding 0 of 360 scored tensors, and only 3 of 497 total.
--hessian-primary-gap 1e-7 How tight Phase 1 must solve before its result is pinned — deliberately far tighter than --quality-mip-rel-gap, so the pin is a provably exact answer, not an accepted approximation.
--hessian-unscored-fallback zero Added 2026-09-19. What --hessian-primary does with tensors that have no score in the Hessian score file. zero (unchanged behaviour): danger weight 0.0, the SAFEST end of the scale, so they are pushed to their floors. kl: each unscored tensor gets the percentile rank of its own isolated KL at the lowest candidate bit, over all tensors (1.0 = most dangerous), so silence does not mean safest. No effect on floors or scored tensors; fades out as coverage grows.

Real command, everything else matching the already-benchmarked production checkpoint (note: --max-low-bit-run 999 is required alongside --hessian-primary, or the separate run-guard mechanism piles extra constraints on top each iteration and the combined solve goes infeasible):

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 \
  --candidates 4,5,6,8,16 \
  --pareto none \
  --group-size 64 \
  --late-attn-min-bits 0 \
  --early-qkv-min-bits 0 \
  --lm-head-min-bits 6 \
  --boundary-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 0 \
  --hessian-floor-tiers "0.05:8,0.15:6" \
  --hessian-primary \
  --max-low-bit-run 999 \
  --output <your_output_path>.json

Same command with the unscored-tensor fallback (added 2026-09-19, recommended while the Hessian sweep is incomplete), complete and exact:

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara

python3 scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  --target-bpw 5.028075122774708 \
  --candidates 4,5,6,8,16 \
  --pareto none \
  --group-size 64 \
  --late-attn-min-bits 0 \
  --early-qkv-min-bits 0 \
  --lm-head-min-bits 6 \
  --boundary-min-bits 6 \
  --hessian-score-file scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --hessian-weight-strength 0 \
  --hessian-floor-tiers "0.05:8,0.15:6" \
  --hessian-primary \
  --hessian-unscored-fallback kl \
  --max-low-bit-run 999 \
  --output 04_LANGUAGE_PLANS/HESSIAN_TEST/plan_floor_primary_klfallback_2026-09-19.json
Warning, real and confirmed (2026-09-18) — --pareto none is required whenever --hessian-score-file is given

The real default (--pareto meaningful) filters candidate bits using isolated-KL dominance, which can silently remove the exact bit-width a Hessian floor needs — real, checked: 55 of 497 tensors affected at 405/497 coverage, sometimes producing genuine MILP infeasibility (HiGHS Status 8: Infeasible) together with --hessian-primary’s own re-pinning, and it gets worse as real coverage grows, not better. --hessian-weight-strength was separately, directly isolated-tested (identical command, only that flag toggled) and confirmed to change zero real output once the floor/primary are active — set it to 0, it only inflates the reported weighted-KL statistic otherwise. Full evidence: MILP_AWARE_HESSIAN_GUIDE.html.

How these flags pick bit-widths (added 2026-09-19). Rank, danger weight, floor tiers and the two-phase solve, with diagrams and checked numbers: How the Solver Turns a Hessian Score into a Bit-Width.

Real, honest status as of this document

See PORT_LEDGER.md in this same directory for the complete, detailed engineering log.


© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.