The first known port of YAQA (Cornell-RelaxML) to Apple's MLX — and the real architecture built to make its two-sided Hessian correction fit on one Mac, where the original method assumes a multi-GPU cluster with CPU offloading.
YAQA's own two-sided correction needs an input-side and an output-side sensitivity map for every tensor it corrects. Holding all of them at once, for this model's real 363-tensor plan, needs more memory than any single Apple Silicon machine has.
The real reference's own Hessian-collection stage (hessian_llama/get_hess_llama.py,
read directly) holds every layer's data simultaneously too — it fits on their hardware via
FSDP sharding across multiple GPUs plus CPU offloading. A single machine's unified
memory has no equivalent to spread across. The real fix had to be architectural, not just a bigger box.
YAQA's own two-sided correction can't reach lm_head on this
machine (its output-side Hessian alone is 246.7GB). The obvious framing is "so we approximate
it instead." Reading the real reference's own save code directly shows something stronger.
hfize_llama.py)model.lm_head.weight.copy_(
lmhead_data['lm_head'].to(model.lm_head.weight.dtype))
Every other real layer in this same file goes through
utils.unpack_quip(...) — real trellis-codebook decompression. lm_head gets
none of that: a direct, uncompressed copy at the model's original full precision. The original
method doesn't approximate this tensor. It doesn't compress it at all.
06_lm_head_gptq.py)Real, scoped GPTQ compensation — column-by-column,
damped-Cholesky-corrected — to lm_head's actual real plan target, using Apple's own
first-party Catcher class unmodified (the same mechanism 08_gptq_apply_plan.py
already runs successfully on every layer, including this one, today).
GPTQ only ever needs the small input-side Hessian (bounded by hidden size) — the 246.7GB wall is entirely on the output side, so it was never in this method's way to begin with.
This isn't a fallback for a YAQA limitation. It's a real, verified case of this port doing
more to lm_head — the single largest tensor in the model — than the original
paper's own reference implementation does. YAQA-UMA's two-sided correction covers 362 of 363
real tensors; the 363rd gets a real, principled compression the reference never even attempts.
Each "run batch as its own fresh OS process" step is
load-bearing, not cosmetic: running multiple batches inside one long-lived process hits a real
Metal resource-handle limit (confirmed directly — mx.clear_cache() does not reset it).
A fresh process means a fresh Metal context, every time.
Every tensor's real result is written to disk the moment it's corrected — not batched up and written at the end. A kill at any point loses at most the one tensor in progress, never anything already committed.
A restart checks the real manifest by tensor name and real bits/group_size — it works even across a changed memory budget, because "batch 7" was never a stable identity, only "this specific tensor is done" is.
An earlier version of this orchestrator assumed a fixed batch count throughout a run — once resume-skip made the real remaining count shrink mid-run, the loop crashed trying to reach an index that no longer existed. No real tensor data was lost (every completed one stayed cached); only the loop's own counting was wrong. Fixed by always requesting "the next undone chunk" and re-checking the real remaining count every iteration, until it hits zero.
Added 2026-09-25 after a night-time crash cost hours. The batching above decides
how the work is split; this controller decides whether it is safe to start each batch, how big it
may be, and what to do when one fails. It lives inside run_full_yaqa.sh and changes nothing in
05_full_model_quantize.py. Every rule below is tied to evidence measured on this machine, and the
places where the evidence stops are stated plainly.
Report (in /Library/Logs/DiagnosticReports/) | When | What it says |
|---|---|---|
gpuEvent-python3.11-2026-09-25-051653.ips | 2026-09-25 05:16:53 | reason firmware-detected lockup, signature 579, restart_reason 4 |
gpuEvent-python3.11-2026-09-20-121024.ips | 2026-09-20 12:10:24 | identical: reason firmware-detected lockup, signature 579, restart_reason 4 |
JetsamEvent-2026-09-24-215729.ips | 2026-09-24 21:57:29 | a different process (python3.14) was by far the largest memory user when macOS ended work: a separate, memory failure |
The crashed batch (14 tensors, 7.76 GB of Hessians, about 22 TFLOP per chunk) was the same size as batches that had run for hours without incident. Its first six chunk losses (2.1989, 1.2555, 0.9013, 1.3176, 3.7357, 2.1532) match the batch that finished all 25 chunks to four decimals, so the data and the computation were identical. There was no memory event: the machine showed 95% free memory afterwards, and the report is a GPU firmware reset, not an out-of-memory kill. The trigger is unknown. So the controller never shrinks on a guess: it shrinks only when a real signal says to, and otherwise defends against the failure it can actually see, which is a GPU that is busy, a batch that silently stalls, and nobody being told.
Steps 1–3 repeat for every batch. Nothing is lost by a retry: results are written per tensor (see above), so a retried batch simply skips whatever already finished, even if its size changed.
| Class | How it is recognised | Retry 1 | Retry 2 | Why this rule |
|---|---|---|---|---|
gpu_lockup | log contains GPU Hang / kIOGPU…Hang | same budget, after the GPU is idle | budget ÷ 2 (saved as the learned cap) | a lockup is not deterministic: identical work passed before, so first try again cleanly |
stall | no new chunk N/24 line for 600 s while still collecting Hessians; the process is killed | same as lockup | same as lockup | a silent hang would otherwise cost the whole night; the check switches off when collection ends, so the hours-long correction phase is never killed |
memory | exit 137, or log says out of memory / Resource limit | min(budget ÷ 2, freshly sensed fit) | same again | the fit model was wrong, so shrink each time |
deterministic | anything else (assertion, NaN, code bug) | none: stop at once | the same input would fail the same way | |
interrupted | KeyboardInterrupt, exit 130 / 143 | none: exit 130, quiet | your Ctrl-C is a decision, not a fault | |
budget = max(1.4, min(default, learned cap, fit))
fit = (free RAM − model − 6) ÷ 1.5
model is the size of the BF16 weights on disk (55.6 GB here, measured). 6 GB is a fixed margin; 1.5 covers the
Hessian matrices plus the temporary matrices the accumulation creates. 1.4 GB is one largest tensor, the smallest batch
that exists. Learned cap starts at the default and only moves after a failure (down) or clean batches (up).
Idle machine, 122 GB free: fit = (122 − 56 − 6) ÷ 1.5 = 40, so the default 8 GB applies.
Another process holds memory, 69 GB free: fit = (69 − 56 − 6) ÷ 1.5 = 4.67, so the batch is sized 4.67 GB before it starts.
Cost of smaller batches is small: the collection pass is about 10 s per chunk, roughly 4 minutes per batch, against about 3.3 hours of per-tensor correction per batch.
File (inside <output>.yaqa_resume/) | What it holds |
|---|---|
adaptive_events.log | Every decision with its reason: sensed headroom, chosen budget, each failure's class, each retry rule applied |
adaptive_state.json | What it has learned (survives a restart): learned_cap_gb, clean_streak, last peak GPU memory, seconds per chunk, last failure |
logs/failed/ | Every failed attempt's full log, named batch_N_attemptK_<class>_<time>.log, never overwritten |
FAILED.txt | Written only when it gives up: class, detail, log tail, and the newest macOS GPU event report path |
Run the probe exactly as before (the retry, sensing and notification are automatic). To re-run after it stopped, use the same command; finished tensors are skipped:
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port && ./run_hessian_probe.sh
Self-test (no GPU, no real data, about 2 minutes; prints PASS or FAIL for each of ten scenarios):
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port && ./test_adaptive_controller.sh
| Claim | Status | Evidence |
|---|---|---|
| The two GPU events are firmware lockups with one signature | Verified | Both .ips reports read directly |
| The crash was not a size effect | Verified | Same data (chunk losses match to 4 decimals), same-size batches ran fine, no memory event |
| The sensors respond to real load | Verified | A 2 GB test allocation moved GPU utilization 10% → 100% and GPU memory 1.85 → 4.93 GB; vm_stat headroom read 122 GB idle |
| Every retry / stop / sensing rule above behaves as written | Verified | Ten fault-injection scenarios pass: one lockup, two, three; a genuine bug; a silent stall; memory kills twice; low RAM; a busy GPU; Ctrl-C with no orphaned process |
| The tests found real defects | Fixed | (1) grep -c exits 1 on zero matches, which under set -e would have silently killed the controller right after every batch started; (2) an empty argument array crashed macOS bash 3.2 with "unbound variable" (latent in the original script) |
| What triggers the firmware lockup | Unknown | The reports do not say which command stalled. Candidates: a long fp32 matmul command buffer, another GPU client, power or thermal state |
| That the controller prevents the next lockup | Not verified | It has not yet run on a real GPU batch. It is built to recover in seconds and to record enough (peak GPU memory, chunk seconds, failure class) that the next event can be diagnosed instead of guessed |
| Thresholds (20% idle, 600 s stall, ÷ 1.5, 6 GB margin, halving) | Engineering choices | Not fitted to data; the recorded telemetry is what lets them be tuned |
| Claim | How it was actually checked |
|---|---|
| Batch composition doesn't change a tensor's measured Hessian | Direct test: same tensor, wrapped alone vs. wrapped with another — 0.0 difference, exactly, not approximately |
| The saved model actually works | Real save → real mlx_lm.load() → real forward pass → correct prediction ("Paris" for "The capital of France is") |
| Calibration is genuinely diverse | Real per-domain token counts confirmed in every batch's own log: agent, code, instruct, prose, thought, tool — all six, every time |
lm_head isn't silently dropped | Real GPTQ one-sided correction, scoped to avoid the same memory wall, tagged distinctly in the manifest |
Every architecture-specific constant this pipeline used to hardcode
is now a real command-line override: --source (the BF16 model), --plan
(its per-tensor bit allocation), --lm-head-name (its vocabulary-output tensor's real
name). Point this at a different MLX model and its own plan, and the same batching, resume, and
correction machinery applies unchanged.
Alternatives considered, if a plainer name reads better: