Companion to Benchmark_Master_Recap.html and MANUAL_RUN_GUIDE.md in this same project. This page answers two questions: "how do I re-run the exact test that produced row X in the master table," and "how do I point that same protocol at a brand-new model." Every command below is real, taken directly from the actual scripts on disk, not paraphrased.
Every real quality-benchmark script in this project lives at the top of HybridOptiQ_FINAL_BENCHMARK/ and follows the same shape: MMLU (n=114), GSM8K/IFEval/BFCL/HumanEval (n=100), always with --reasoning. Run any of them with bash <script> from anywhere — every path inside is absolute.
| Script | Targets | Output logs |
|---|---|---|
| repro_check_yaqa.sh | YAQA buildQwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32 (post scales/biases + embed_tokens fix rebuild) | quality/yaqa_<task>_n<N>.log |
| repro_check_v33floor.sh | v33floorHybridPareto5-V6_matched_v33floor — V3.3 early-QKV floor ON | quality/v33floor_<task>_n<N>.log |
| repro_check_matched_flat_v52.sh | v6matched, flat5naive, v52runs all three in sequence via a shared run_all_tasks() function |
quality/{v6matched,flat5naive,v52}_<task>_n<N>.log |
| repro_check_stock5_only.sh | stock5OptiQ-5bpw/optiq_mixed — rerun after the 2026-08-29 13:44 disk-full crash | quality/stock5_<task>_n<N>.log |
| repro_check.sh | v6, stock5original pass for both, before the disk-full crash required stock5's separate rerun above | quality/{v6,stock5}_<task>_n<N>.log |
# from anywhere — every script uses absolute paths internally
bash /Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/repro_check_yaqa.sh
# or, to watch it live instead of waiting for it to finish:
nohup /Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/repro_check_yaqa.sh \
> /Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/yaqa_benchmark_run.log 2>&1 &
tail -f /Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/yaqa_benchmark_run.log
Each of the 5 tasks in one script's run is a separate optiq eval invocation on the full model — expect this to take a while per model (not a smoke test). Reading the results back afterward: each .log file's printed summary is the real result; --output-json silently does nothing for single-task runs (see the tool quirks below), so don't look for a JSON file next to these — only --task all writes one.
Every script above is the same template with the model path(s) and a tag swapped in. This is the exact real template — copy it, change MODEL_PATH and TAG, nothing else needs to change to stay consistent with every other row in the master table.
#!/bin/bash set -uo pipefail OPTIQ=/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/optiq OUT=/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/quality MODEL_PATH=/path/to/the/new/model TAG=new_model_name echo "[$(date '+%H:%M:%S')] $TAG mmlu (n=114)" "$OPTIQ" eval "$MODEL_PATH" --task mmlu --n-samples 114 --reasoning \ > "$OUT/${TAG}_mmlu_n114.log" 2>&1 echo "[$(date '+%H:%M:%S')] $TAG gsm8k (n=100)" "$OPTIQ" eval "$MODEL_PATH" --task gsm8k --n-samples 100 --reasoning \ > "$OUT/${TAG}_gsm8k_n100.log" 2>&1 echo "[$(date '+%H:%M:%S')] $TAG ifeval (n=100)" "$OPTIQ" eval "$MODEL_PATH" --task ifeval --n-samples 100 --reasoning \ > "$OUT/${TAG}_ifeval_n100.log" 2>&1 echo "[$(date '+%H:%M:%S')] $TAG bfcl (n=100)" "$OPTIQ" eval "$MODEL_PATH" --task bfcl --n-samples 100 --reasoning \ > "$OUT/${TAG}_bfcl_n100.log" 2>&1 echo "[$(date '+%H:%M:%S')] $TAG humaneval (n=100)" "$OPTIQ" eval "$MODEL_PATH" --task humaneval --n-samples 100 --reasoning \ > "$OUT/${TAG}_humaneval_n100.log" 2>&1
Real, hard-won gotchas — each one caused a real wrong result or a silent failure the first time, before it was understood:
Always pass --reasoning. This is a thinking model — without it, MMLU scores by first-token logit and collapses to chance. Every real run in this project passes it, on every task.
Always pass --reference-model / --reference-mode bf16 explicitly on anything touching KL (smoketest, kl). Leaving it unset makes optiq guess a Hugging Face Hub repo id from the local folder name and 404 — none of these local artifacts exist on the Hub.
--output-json only genuinely works for --task all, mmlu, and bfcl. For kl, smoketest, gsm8k/gsm8k-50, ifeval, humaneval, hashhop it's a silent no-op (confirmed in the real mlx-optiq source) — every script above redirects stdout to a .log file for exactly this reason. Read the printed summary line, don't look for JSON.
smoketest can exit code 1 after printing a clean success line. Root cause not found in source. Treat the printed "Smoketest complete" summary as ground truth, not the shell exit code, for this one task specifically.
Avoid GPU contention between concurrent runs. repro_check_v33floor.sh polls kill -0 $WAIT_PID in a loop before starting, specifically to avoid overlapping with another still-running benchmark on the same machine. Don't launch two of these scripts at once.
Two separate installs, in this order — combining them into one uv tool install --with datasets --with git+... silently drops click (optiq's own CLI framework) via a resolution conflict:
# Step 1: datasets (needed by GSM8K) uv tool install --reinstall --with datasets mlx-optiq # Step 2: hashhop (GitHub-only, conflicts with click during full resolution -- # install with --no-deps straight into the same tool venv instead) uv pip install \ --python /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python \ 'git+https://github.com/codelion/hash-hop' --no-deps # confirm: /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python -c "import click, mlx.core, mlx_lm, datasets, hashhop; print('ok')" optiq --version
The quality scripts above never measure tok/s. Speed is a separate real tool, mtplx tune — this is what produced every AR/D1/D2/D3 number referenced in the RCA report and the master recap's "D3 speed" column.
# 3 real repetitions, same convention as every model in the master table for i in 1 2 3; do mtplx tune --model "/path/to/model" --depths 1,2,3 --retune --json \ --output "/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/speed_tune/TAG_run${i}.json" done # per-depth acceptance breakdown (what showed the D3 depth-3 acceptance collapse) mtplx mtp-depth-sweep --model "/path/to/model" --depths 1,2,3 --compare-ar \ --output "/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/speed_depthsweep/TAG.json" # head-to-head comparison against a reference model, real report + JSON # (this is what produced the AR/D1/D2/D3 tables referenced throughout the RCA) python3 forge_doctor_tune_compare.py \ --a "/path/to/reference/model" --a-label STRATIFIED \ --b "/path/to/candidate/model" --b-label YAQA
Not a benchmark script, but the command that produced the model the scripts above are currently pointed at. Full detail, root cause, and independent verification: research_hadamard_blowup/exports/RCA_scales_f32_bug.html in the yaqa_port project.
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 -u 05_full_model_quantize.py \ --assemble-from-resume \ --resume-dir "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume" \ --output "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32" \ --n-calibration 4 \ --incoherence none \ --fallback-bits 4
26h Hessian-correction sweep not re-run — --assemble-from-resume re-saves the already-corrected tensors from .yaqa_resume/ at the right dtype, with embed_tokens now left at full precision rather than quantized. See the CHANGELOG entry dated 2026-09-07 (later still) for the full real evidence trail.
All repro_check*.sh + run_full_benchmark.sh
One .log per task per model, real printed summaries
mtplx tune JSON, 3 reps per model
The living quality+speed table, all 6 prior models
Root cause, fix, and post-rebuild verification
Complete engineering log, every model, every fix