← Back to index Benchmark Run Guide
VEGA / MLX OPTIQ LAB · Benchmark Run Guide
Reference · every command, copy-pasteable, no placeholders

How to re-run an existing benchmark, or design a new one.

Companion to Benchmark_Master_Recap.html and MANUAL_RUN_GUIDE.md in this same project. This page answers two questions: "how do I re-run the exact test that produced row X in the master table," and "how do I point that same protocol at a brand-new model." Every command below is real, taken directly from the actual scripts on disk, not paraphrased.

See also: Benchmark_Master_Recap.html · MANUAL_RUN_GUIDE.md

ARe-run an existing script

Every real quality-benchmark script in this project lives at the top of HybridOptiQ_FINAL_BENCHMARK/ and follows the same shape: MMLU (n=114), GSM8K/IFEval/BFCL/HumanEval (n=100), always with --reasoning. Run any of them with bash <script> from anywhere — every path inside is absolute.

ScriptTargetsOutput logs
repro_check_yaqa.sh YAQA buildQwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32 (post scales/biases + embed_tokens fix rebuild) quality/yaqa_<task>_n<N>.log
repro_check_v33floor.sh v33floorHybridPareto5-V6_matched_v33floor — V3.3 early-QKV floor ON quality/v33floor_<task>_n<N>.log
repro_check_matched_flat_v52.sh v6matched, flat5naive, v52runs all three in sequence via a shared run_all_tasks() function quality/{v6matched,flat5naive,v52}_<task>_n<N>.log
repro_check_stock5_only.sh stock5OptiQ-5bpw/optiq_mixed — rerun after the 2026-08-29 13:44 disk-full crash quality/stock5_<task>_n<N>.log
repro_check.sh v6, stock5original pass for both, before the disk-full crash required stock5's separate rerun above quality/{v6,stock5}_<task>_n<N>.log
# from anywhere — every script uses absolute paths internally
bash /Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/repro_check_yaqa.sh

# or, to watch it live instead of waiting for it to finish:
nohup /Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/repro_check_yaqa.sh \
  > /Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/yaqa_benchmark_run.log 2>&1 &
tail -f /Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/yaqa_benchmark_run.log
Runtime

Each of the 5 tasks in one script's run is a separate optiq eval invocation on the full model — expect this to take a while per model (not a smoke test). Reading the results back afterward: each .log file's printed summary is the real result; --output-json silently does nothing for single-task runs (see the tool quirks below), so don't look for a JSON file next to these — only --task all writes one.

BDesign a new one, for a new model

Every script above is the same template with the model path(s) and a tag swapped in. This is the exact real template — copy it, change MODEL_PATH and TAG, nothing else needs to change to stay consistent with every other row in the master table.

#!/bin/bash
set -uo pipefail

OPTIQ=/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/optiq
OUT=/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/quality
MODEL_PATH=/path/to/the/new/model
TAG=new_model_name

echo "[$(date '+%H:%M:%S')] $TAG mmlu (n=114)"
"$OPTIQ" eval "$MODEL_PATH" --task mmlu --n-samples 114 --reasoning \
  > "$OUT/${TAG}_mmlu_n114.log" 2>&1

echo "[$(date '+%H:%M:%S')] $TAG gsm8k (n=100)"
"$OPTIQ" eval "$MODEL_PATH" --task gsm8k --n-samples 100 --reasoning \
  > "$OUT/${TAG}_gsm8k_n100.log" 2>&1

echo "[$(date '+%H:%M:%S')] $TAG ifeval (n=100)"
"$OPTIQ" eval "$MODEL_PATH" --task ifeval --n-samples 100 --reasoning \
  > "$OUT/${TAG}_ifeval_n100.log" 2>&1

echo "[$(date '+%H:%M:%S')] $TAG bfcl (n=100)"
"$OPTIQ" eval "$MODEL_PATH" --task bfcl --n-samples 100 --reasoning \
  > "$OUT/${TAG}_bfcl_n100.log" 2>&1

echo "[$(date '+%H:%M:%S')] $TAG humaneval (n=100)"
"$OPTIQ" eval "$MODEL_PATH" --task humaneval --n-samples 100 --reasoning \
  > "$OUT/${TAG}_humaneval_n100.log" 2>&1

Real, hard-won gotchas — each one caused a real wrong result or a silent failure the first time, before it was understood:

1

Always pass --reasoning. This is a thinking model — without it, MMLU scores by first-token logit and collapses to chance. Every real run in this project passes it, on every task.

2

Always pass --reference-model / --reference-mode bf16 explicitly on anything touching KL (smoketest, kl). Leaving it unset makes optiq guess a Hugging Face Hub repo id from the local folder name and 404 — none of these local artifacts exist on the Hub.

3

--output-json only genuinely works for --task all, mmlu, and bfcl. For kl, smoketest, gsm8k/gsm8k-50, ifeval, humaneval, hashhop it's a silent no-op (confirmed in the real mlx-optiq source) — every script above redirects stdout to a .log file for exactly this reason. Read the printed summary line, don't look for JSON.

4

smoketest can exit code 1 after printing a clean success line. Root cause not found in source. Treat the printed "Smoketest complete" summary as ground truth, not the shell exit code, for this one task specifically.

5

Avoid GPU contention between concurrent runs. repro_check_v33floor.sh polls kill -0 $WAIT_PID in a loop before starting, specifically to avoid overlapping with another still-running benchmark on the same machine. Don't launch two of these scripts at once.

One-time environment setup (already done on this machine, listed for the record)

Two separate installs, in this order — combining them into one uv tool install --with datasets --with git+... silently drops click (optiq's own CLI framework) via a resolution conflict:

# Step 1: datasets (needed by GSM8K)
uv tool install --reinstall --with datasets mlx-optiq

# Step 2: hashhop (GitHub-only, conflicts with click during full resolution --
# install with --no-deps straight into the same tool venv instead)
uv pip install \
  --python /Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python \
  'git+https://github.com/codelion/hash-hop' --no-deps

# confirm:
/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python -c "import click, mlx.core, mlx_lm, datasets, hashhop; print('ok')"
optiq --version

CSpeed, separately from quality

The quality scripts above never measure tok/s. Speed is a separate real tool, mtplx tune — this is what produced every AR/D1/D2/D3 number referenced in the RCA report and the master recap's "D3 speed" column.

# 3 real repetitions, same convention as every model in the master table
for i in 1 2 3; do
  mtplx tune --model "/path/to/model" --depths 1,2,3 --retune --json \
    --output "/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/speed_tune/TAG_run${i}.json"
done

# per-depth acceptance breakdown (what showed the D3 depth-3 acceptance collapse)
mtplx mtp-depth-sweep --model "/path/to/model" --depths 1,2,3 --compare-ar \
  --output "/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/speed_depthsweep/TAG.json"

# head-to-head comparison against a reference model, real report + JSON
# (this is what produced the AR/D1/D2/D3 tables referenced throughout the RCA)
python3 forge_doctor_tune_compare.py \
  --a "/path/to/reference/model" --a-label STRATIFIED \
  --b "/path/to/candidate/model" --b-label YAQA

DThe YAQA rebuild command, for reference

Not a benchmark script, but the command that produced the model the scripts above are currently pointed at. Full detail, root cause, and independent verification: research_hadamard_blowup/exports/RCA_scales_f32_bug.html in the yaqa_port project.

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port

/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 -u 05_full_model_quantize.py \
  --assemble-from-resume \
  --resume-dir "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume" \
  --output "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32" \
  --n-calibration 4 \
  --incoherence none \
  --fallback-bits 4

26h Hessian-correction sweep not re-run — --assemble-from-resume re-saves the already-corrected tensors from .yaqa_resume/ at the right dtype, with embed_tokens now left at full precision rather than quantized. See the CHANGELOG entry dated 2026-09-07 (later still) for the full real evidence trail.

Where these scripts and this guide live

Quality scripts

All repro_check*.sh + run_full_benchmark.sh

/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/

Result logs

One .log per task per model, real printed summaries

/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/quality/

Speed results

mtplx tune JSON, 3 reps per model

/Users/hghelab/HybridOptiQ_FINAL_BENCHMARK/speed_tune/

Master recap

The living quality+speed table, all 6 prior models

docs/HTML/Benchmark_Master_Recap.html

YAQA RCA

Root cause, fix, and post-rebuild verification

.../yaqa_port/research_hadamard_blowup/exports/RCA_scales_f32_bug.html

Full narrative

Complete engineering log, every model, every fix

LEDGER.md