← Back to index YAQA-UMA
MLX_OptiQ · Engineering Record · written 2026-09-05 · adaptive controller added 2026-09-25

Adapting YAQA to a Single Unified-Memory Machine

The first known port of YAQA (Cornell-RelaxML) to Apple's MLX — and the real architecture built to make its two-sided Hessian correction fit on one Mac, where the original method assumes a multi-GPU cluster with CPU offloading.

The real scale of the problem

YAQA's own two-sided correction needs an input-side and an output-side sensitivity map for every tensor it corrects. Holding all of them at once, for this model's real 363-tensor plan, needs more memory than any single Apple Silicon machine has.

557 GB
Real Hessian memory needed to hold all 363 tensors at once
128 GB
Real RAM on the machine this runs on
246.7 GB
One tensor's own Hessian alone (lm_head — vocabulary-sized output)
362 / 363
Real tensors YAQA's two-sided correction actually reaches
Why this isn't a porting mistake

The real reference's own Hessian-collection stage (hessian_llama/get_hess_llama.py, read directly) holds every layer's data simultaneously too — it fits on their hardware via FSDP sharding across multiple GPUs plus CPU offloading. A single machine's unified memory has no equivalent to spread across. The real fix had to be architectural, not just a bigger box.

lm_head: not a workaround — further than the original method goes

YAQA's own two-sided correction can't reach lm_head on this machine (its output-side Hessian alone is 246.7GB). The obvious framing is "so we approximate it instead." Reading the real reference's own save code directly shows something stronger.

The real reference (hfize_llama.py)

model.lm_head.weight.copy_(
    lmhead_data['lm_head'].to(model.lm_head.weight.dtype))

Every other real layer in this same file goes through utils.unpack_quip(...) — real trellis-codebook decompression. lm_head gets none of that: a direct, uncompressed copy at the model's original full precision. The original method doesn't approximate this tensor. It doesn't compress it at all.

This port (06_lm_head_gptq.py)

Real, scoped GPTQ compensation — column-by-column, damped-Cholesky-corrected — to lm_head's actual real plan target, using Apple's own first-party Catcher class unmodified (the same mechanism 08_gptq_apply_plan.py already runs successfully on every layer, including this one, today).

GPTQ only ever needs the small input-side Hessian (bounded by hidden size) — the 246.7GB wall is entirely on the output side, so it was never in this method's way to begin with.

16-bit
Real reference precision for lm_head — 2.543GB, uncompressed, copied straight through
6-bit
This port's real plan target for lm_head — 0.954GB, with a principled GPTQ correction
62.5%
Real size reduction on this tensor — the one thing the original method's own code never attempts
Why this matters

This isn't a fallback for a YAQA limitation. It's a real, verified case of this port doing more to lm_head — the single largest tensor in the model — than the original paper's own reference implementation does. YAQA-UMA's two-sided correction covers 362 of 363 real tensors; the 363rd gets a real, principled compression the reference never even attempts.

The batching + resume architecture, spatially

> 0= 0, all doneReal plan: 363 tensorsacross 64 real layersSort by real layer index(not arbitrary order)Greedy-pack into memory-boundedbatches (default 8GB/batch)Real remainingbatch count?Run batch as its OWNfresh OS processWrap this batch's tensors(input/output capture)Real forward+backward passover N=24 stratified calibrationCollect real H_I, H_Ofor every tensor in this batchPer-tensor LDLQ_2hesscorrection + real packWrite result to resume-cache(manifest.json + safetensors)Separate: lm_head viareal one-sided GPTQFinal assembly: read everyreal result, build complete modelReal, loadable model(mlx_lm.load() verified)

Each "run batch as its own fresh OS process" step is load-bearing, not cosmetic: running multiple batches inside one long-lived process hits a real Metal resource-handle limit (confirmed directly — mx.clear_cache() does not reset it). A fresh process means a fresh Metal context, every time.

Why nothing gets lost, ever

Per-tensor, not per-batch

Every tensor's real result is written to disk the moment it's corrected — not batched up and written at the end. A kill at any point loses at most the one tensor in progress, never anything already committed.

Resume by name, not by index

A restart checks the real manifest by tensor name and real bits/group_size — it works even across a changed memory budget, because "batch 7" was never a stable identity, only "this specific tensor is done" is.

Real, found and fixed this session

An earlier version of this orchestrator assumed a fixed batch count throughout a run — once resume-skip made the real remaining count shrink mid-run, the loop crashed trying to reach an index that no longer existed. No real tensor data was lost (every completed one stayed cached); only the loop's own counting was wrong. Fixed by always requesting "the next undone chunk" and re-checking the real remaining count every iteration, until it hits zero.

The adaptive controller: sense, classify, retry

Added 2026-09-25 after a night-time crash cost hours. The batching above decides how the work is split; this controller decides whether it is safe to start each batch, how big it may be, and what to do when one fails. It lives inside run_full_yaqa.sh and changes nothing in 05_full_model_quantize.py. Every rule below is tied to evidence measured on this machine, and the places where the evidence stops are stated plainly.

What actually failed, from macOS's own crash reports

Report (in /Library/Logs/DiagnosticReports/)WhenWhat it says
gpuEvent-python3.11-2026-09-25-051653.ips2026-09-25 05:16:53reason firmware-detected lockup, signature 579, restart_reason 4
gpuEvent-python3.11-2026-09-20-121024.ips2026-09-20 12:10:24identical: reason firmware-detected lockup, signature 579, restart_reason 4
JetsamEvent-2026-09-24-215729.ips2026-09-24 21:57:29a different process (python3.14) was by far the largest memory user when macOS ended work: a separate, memory failure
Why "make the batch smaller" was the wrong reflex

The crashed batch (14 tensors, 7.76 GB of Hessians, about 22 TFLOP per chunk) was the same size as batches that had run for hours without incident. Its first six chunk losses (2.1989, 1.2555, 0.9013, 1.3176, 3.7357, 2.1532) match the batch that finished all 25 chunks to four decimals, so the data and the computation were identical. There was no memory event: the machine showed 95% free memory afterwards, and the report is a GPU firmware reset, not an out-of-memory kill. The trigger is unknown. So the controller never shrinks on a guess: it shrinks only when a real signal says to, and otherwise defends against the failure it can actually see, which is a GPU that is busy, a batch that silently stalls, and nobody being told.

The decision flow, one batch

1 · SENSE GPU idle (≤ 20% for 3 samples) Real free RAM from vm_stat 2 · SIZE THE BATCH min(default, learned cap, (free RAM − model − 6) ÷ 1.5) 3 · RUN + MEASURE Batch = its own fresh process Polls GPU memory + chunk times Exit 0 ? yes 4 · RECORD, NEXT BATCH Peak GPU GB + seconds per chunk 2 clean in a row → cap × 1.5 · ↺ step 1 no 5 · CLASSIFY THE FAILURE from the log text and exit code GPU LOCKUP / SILENT STALL Retry 1: same size, once the GPU is idle again Retry 2: half the size then ↺ step 3 OUT OF MEMORY (exit 137) Retry 1 and 2: halve the size, capped by freshly sensed RAM then ↺ step 3 GENUINE BUG (assert, NaN…) No retry. Stop immediately: retrying only wastes hours After the 2nd failed retry (or a bug): write FAILED.txt + macOS banner and sound You find out in minutes, not after an 8-hour night. Ctrl-C is never retried.

Steps 1–3 repeat for every batch. Nothing is lost by a retry: results are written per tensor (see above), so a retried batch simply skips whatever already finished, even if its size changed.

The rules, exactly

ClassHow it is recognisedRetry 1Retry 2Why this rule
gpu_lockuplog contains GPU Hang / kIOGPU…Hangsame budget, after the GPU is idlebudget ÷ 2 (saved as the learned cap)a lockup is not deterministic: identical work passed before, so first try again cleanly
stallno new chunk N/24 line for 600 s while still collecting Hessians; the process is killedsame as lockupsame as lockupa silent hang would otherwise cost the whole night; the check switches off when collection ends, so the hours-long correction phase is never killed
memoryexit 137, or log says out of memory / Resource limitmin(budget ÷ 2, freshly sensed fit)same againthe fit model was wrong, so shrink each time
deterministicanything else (assertion, NaN, code bug)none: stop at oncethe same input would fail the same way
interruptedKeyboardInterrupt, exit 130 / 143none: exit 130, quietyour Ctrl-C is a decision, not a fault

How the batch size is decided, with real numbers

The formula

budget = max(1.4, min(default, learned cap, fit))
fit = (free RAM − model − 6) ÷ 1.5

model is the size of the BF16 weights on disk (55.6 GB here, measured). 6 GB is a fixed margin; 1.5 covers the Hessian matrices plus the temporary matrices the accumulation creates. 1.4 GB is one largest tensor, the smallest batch that exists. Learned cap starts at the default and only moves after a failure (down) or clean batches (up).

Worked examples

Idle machine, 122 GB free: fit = (122 − 56 − 6) ÷ 1.5 = 40, so the default 8 GB applies.

Another process holds memory, 69 GB free: fit = (69 − 56 − 6) ÷ 1.5 = 4.67, so the batch is sized 4.67 GB before it starts.

Cost of smaller batches is small: the collection pass is about 10 s per chunk, roughly 4 minutes per batch, against about 3.3 hours of per-tensor correction per batch.

Where to look, exactly

File (inside <output>.yaqa_resume/)What it holds
adaptive_events.logEvery decision with its reason: sensed headroom, chosen budget, each failure's class, each retry rule applied
adaptive_state.jsonWhat it has learned (survives a restart): learned_cap_gb, clean_streak, last peak GPU memory, seconds per chunk, last failure
logs/failed/Every failed attempt's full log, named batch_N_attemptK_<class>_<time>.log, never overwritten
FAILED.txtWritten only when it gives up: class, detail, log tail, and the newest macOS GPU event report path

Run the probe exactly as before (the retry, sensing and notification are automatic). To re-run after it stopped, use the same command; finished tensors are skipped:

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port && ./run_hessian_probe.sh

Self-test (no GPU, no real data, about 2 minutes; prints PASS or FAIL for each of ten scenarios):

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port && ./test_adaptive_controller.sh

What is verified, and what is not

ClaimStatusEvidence
The two GPU events are firmware lockups with one signatureVerifiedBoth .ips reports read directly
The crash was not a size effectVerifiedSame data (chunk losses match to 4 decimals), same-size batches ran fine, no memory event
The sensors respond to real loadVerifiedA 2 GB test allocation moved GPU utilization 10% → 100% and GPU memory 1.85 → 4.93 GB; vm_stat headroom read 122 GB idle
Every retry / stop / sensing rule above behaves as writtenVerifiedTen fault-injection scenarios pass: one lockup, two, three; a genuine bug; a silent stall; memory kills twice; low RAM; a busy GPU; Ctrl-C with no orphaned process
The tests found real defectsFixed(1) grep -c exits 1 on zero matches, which under set -e would have silently killed the controller right after every batch started; (2) an empty argument array crashed macOS bash 3.2 with "unbound variable" (latent in the original script)
What triggers the firmware lockupUnknownThe reports do not say which command stalled. Candidates: a long fp32 matmul command buffer, another GPU client, power or thermal state
That the controller prevents the next lockupNot verifiedIt has not yet run on a real GPU batch. It is built to recover in seconds and to record enough (peak GPU memory, chunk seconds, failure class) that the next event can be diagnosed instead of guessed
Thresholds (20% idle, 600 s stall, ÷ 1.5, 6 GB margin, halving)Engineering choicesNot fitted to data; the recorded telemetry is what lets them be tuned

Verified end to end, not assumed

ClaimHow it was actually checked
Batch composition doesn't change a tensor's measured HessianDirect test: same tensor, wrapped alone vs. wrapped with another — 0.0 difference, exactly, not approximately
The saved model actually worksReal save → real mlx_lm.load() → real forward pass → correct prediction ("Paris" for "The capital of France is")
Calibration is genuinely diverseReal per-domain token counts confirmed in every batch's own log: agent, code, instruct, prose, thought, tool — all six, every time
lm_head isn't silently droppedReal GPTQ one-sided correction, scoped to avoid the same memory wall, tagged distinctly in the manifest

Real universality, not just for this one model

Every architecture-specific constant this pipeline used to hardcode is now a real command-line override: --source (the BF16 model), --plan (its per-tensor bit allocation), --lm-head-name (its vocabulary-output tensor's real name). Point this at a different MLX model and its own plan, and the same batching, resume, and correction machinery applies unchanged.

A name for it

YAQA-UMA
YAQA for Unified Memory Architecture — real, not decorative: Apple Silicon's actual shared CPU/GPU memory design is the specific real constraint this whole batching and resume system was built to work within.

Alternatives considered, if a plainer name reads better:

mlx-yaqa Chunked YAQA (cYAQA) YAQA/Metal