A compression run on this Mac takes days. On the night of 25 September the GPU reset itself three hours into a run, and the run sat stopped until someone noticed. This page explains what the run does, what went wrong, what the evidence says caused it (and what it does not), and how the new controller senses the machine, sizes each step to fit, and recovers on its own.
Read this section and you have the point. The rest is the detail and the evidence.
The arithmetic that forces it.
The machine is a Mac with 128 GB of memory that the CPU and the GPU share (unified memory). The operating system lets the GPU use at most about 115 GB of it. The model’s own weights, kept loaded the whole time, take 56 GB. That leaves roughly 60 GB, for everything else.
For this plan, the sensitivity tables for all 487 tensors add up to 326 GB: more than five times the room available. So they cannot be built together. The run measures a group of tensors at a time, using a budget of 8 GB of tables per group by default, corrects them, saves them, and moves on.
mx.device_info(), the weight size from disk.A group of tensors whose sensitivity tables are built together in one pass of the model over sample text. After the pass, each tensor is corrected one at a time and its result is saved to disk straight away.
What happens inside a batch, and how long each part takes.
Two design choices matter for everything that follows:
What happened, in order, with the real timestamps.
The old script did exactly one thing on any failure: print the last lines of the log and exit. It had no way to tell a one-off GPU reset from a real bug, no retry, and no way of telling you. That, not the reset itself, is what turned a few minutes into a lost night.
Whenever the GPU resets, macOS saves a small report. There are two, five days apart, and they match:
| Report file (in /Library/Logs/DiagnosticReports/) | Time | Reason it records | Signature |
|---|---|---|---|
| gpuEvent-python3.11-2026-09-20-121024.ips | 20 Sep 12:10:24 | firmware-detected lockup | 579 · guilty_dm 3 · restart_reason 4 |
| gpuEvent-python3.11-2026-09-25-051653.ips | 25 Sep 05:16:53 | firmware-detected lockup | 579 · guilty_dm 3 · restart_reason 4 |
The GPU has its own small built-in program (firmware). It watches the work it is given, and if a piece of work stops making progress it declares a lockup and resets the GPU. That is different from running out of memory, which produces a different report and a different error.
The natural reflex is “make the batches smaller.” The evidence says that would not have prevented this.
The crashed batch and the batch that finished use the same sample text in the same order. The loss (how wrong the model is on each chunk) is the same number to four decimal places for the six chunks the crashed batch completed. Nothing about the input differed.
logs/batch_0.log and logs/batch_1.log. The rings sit exactly on the dots for chunks 1 to 6.The batches are filled to the 8 GB budget, so nearly all of them are the same size. The batch that crashed was 7.76 GB and about 22 trillion arithmetic operations (22 TFLOP) per chunk. Batches of that size ran for hours without incident on 19 and 20 September.
After the crash, memory showed 95% free, and there was no out-of-memory event. There was a memory event the night before (a different Python program was by far the largest memory user and macOS ended work), so memory contention is real. It is a separate failure with a separate rule below.
Reason on the report: firmware-detected lockup. Seen twice. Cause unknown. Not related to size in the evidence.
Seen on 24 September: a Python 3.14 process was by far the largest memory user and macOS ended work. Cause known; the fix is to size batches to what is really free.
For example a number that is not a number (NaN) or a failed check in the code. Retrying gives the same failure, so the right response is to stop and say so.
The first suggestion made during this investigation was to lower the batch budget to 4 GB. That was a guess and was withdrawn once the crash report and the identical-data evidence were read. A smaller batch does not address a firmware lockup that struck a batch the same size as ones that worked.
The crash reports say a lockup happened; they do not say which piece of GPU work stalled. Candidates: a very long single calculation, another program using the GPU at that moment, or power and heat conditions. This is why the controller is built to defend and to record, not to pretend it has removed the cause.
Not the batching. The lack of any measurement, judgement or alert around it.
It lives inside the same script (run_full_yaqa.sh) and changes nothing about the maths.
So today’s proven value is fast recovery and reporting, which works whatever the lockup trigger turns out to be. A census on 25 September showed 20 or more programs sharing the GPU (WindowServer, Chrome, Safari, Spotlight and more) that the busy-percentage gate cannot tell apart, which is why a black-box recorder of the GPU client list and memory figures now runs alongside the build.
Before starting (and before every retry) it checks that the GPU is quiet: the busy figure macOS reports, from every program, must be 20% or less on three checks in a row, ten seconds apart. It waits up to ten minutes; if the GPU is still busy it says so, sends a notification and carries on. It also reads how much memory is really free.
fit = (free memory − model − 6) ÷ 1.5, and the batch budget is the smallest of the default (8 GB), the cap it has learned from earlier failures, and fit, never below 1.4 GB (one large tensor). Section 7 lets you try it.
The batch runs as its own process. From outside, every 10 seconds, the controller records GPU memory and counts finished chunks. If no new chunk appears for 10 minutes while the sensitivity tables are still being built, it treats the process as stalled and stops it.
On success it saves peak GPU memory and seconds per chunk. After two clean batches in a row a lowered cap is raised by 50%, back up to the default: it recovers instead of staying cautious forever.
The failed attempt’s log is read for a GPU hang message, an out-of-memory sign or an interrupt, and the exit code is checked. The log is kept in logs/failed/.
At most two retries, following the rule for that class (the diagram above). A stall check applies only while sensitivity tables are being built, so the three-hour correction phase is never interrupted by it.
The lockup is not repeatable: identical work had already passed. Shrinking on the first failure would be acting on a guess. So the first retry is the same batch after the GPU is idle again; only a repeated lockup halves the size. Out-of-memory failures shrink at once, because there the size is the obvious suspect.
Both use the controller’s real formulas. The second reproduces the retry rules the automated tests check.
Model weights 56 GB and the 6 GB margin are as in the real controller. The 201 tensors are the ones not yet finished when the current run started; batches are packed in the same order as the real code. The extra-time figure assumes about 4 minutes of fixed overhead per batch (model load plus sensitivity measurement), an estimate from the logs; the correction time itself does not depend on how many batches there are.
Pick a failure and watch the decisions. The batch starts at the size the first tool chose.
These sequences are the automated test cases T2 to T7 of test_adaptive_controller.sh. The tool applies the same rules in the page; the test script is what proves the real script does the same, so run it after any edit to run_full_yaqa.sh.
Ten scenarios, run against the real script with a fake batch program that fails on command.
| Test | What is injected | What must happen | Result |
|---|---|---|---|
| T1 | Nothing fails | Three batches at 8 GB, no notification | PASS |
| T2 | One GPU lockup | Retry at the same size: 8, 8 | PASS |
| T3 | Two lockups | 8, 8, then half: 4; cap 4 remembered | PASS |
| T4 | Three lockups | Stop after 2 retries; report and notification | PASS |
| T5 | A genuine bug (NaN) | No retry, stop at once, notify | PASS |
| T6 | A silent stall | Killed after the stall limit, retried | PASS |
| T7 | Out of memory twice | 8, 4, 2 | PASS |
| T8 | Only 69 GB free | Batch sized down to 4.67 GB before starting | PASS |
| T9 | Another program keeps the GPU 60% busy | Waits, warns, notifies, proceeds | PASS |
| T10 | Ctrl-C mid-batch | Exit code 130, quiet, no leftover process | PASS |
| Question | Status | What would settle it |
|---|---|---|
| What triggers the firmware lockup? | Unknown | The next event, with the recorded peak GPU memory, chunk seconds and failed-attempt logs to compare against |
| Does the controller prevent the next lockup? | Not shown | Time. It has been built to recover, not proven to prevent |
| Are 20% idle, 10-minute stall limit, the 1.5 factor, 6 GB margin and halving the right numbers? | Engineering choices | Fitting them to the recorded telemetry after a few events |
| Which GPU-memory figure tracks danger? The controller records “in use” (first real batch, 25 Sep: peak 76.4 GB during Hessian collection, 17.9 GB mid-correction); the driver’s “allocated” figure read 105.9 GB, 92% of the 115.4 GB ceiling, at the same moment | Unknown | Record both at the next failure and compare; memory pressure itself stayed green |
| Does the stall watchdog work against a real hung GPU process? | Tested on a fake only | A real silent hang |
| Does the live dashboard show the controller? | Not yet | A dashboard panel reading the decision log (not built) |
| What if the terminal closes or the Mac loses power? | Not handled | The run stops; re-run the same command. There is no automatic restart after a reboot |
The command is unchanged. The controller is automatic.
Start the run (this is the exact command used for the probe build):
PROJECT_ROOT=/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara
"$PROJECT_ROOT/scripts/yaqa_port/run_full_yaqa.sh" \
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE \
--skip-lm-head-gptq \
--plan "$PROJECT_ROOT/04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.json" \
--n-calibration 4 \
--incoherence none \
--correct-mtp \
--godmode-multi-bit-checkpoint \
--godmode-candidate-bits 3,4,5,6,8
If it stopped, run the same command again: finished tensors are skipped and what it learned is kept. Watch its decisions live:
tail -f /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume/adaptive_events.log
Prove the controller still behaves after any edit to the script (about 2 minutes, no GPU, no real data):
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port && ./test_adaptive_controller.sh
The shell reads a running script piece by piece, so changing the file mid-run can corrupt the run. Edit between runs, then run the test above.
Nothing on this page is estimated without saying so.
| Number | Source |
|---|---|
| Firmware lockup, signature 579, guilty_dm 3, restart_reason 4, on 20 and 25 Sep | /Library/Logs/DiagnosticReports/gpuEvent-python3.11-2026-09-20-121024.ips and ...-2026-09-25-051653.ips |
| Chunk losses 2.1989, 1.2555, 0.9013, 1.3176, 3.7357, 2.1532 | <resume>/logs/batch_0.log and logs/batch_1.log (batch_1 is kept as the first failed log) |
| Crash at 05:16:53, batch 1 started 05:15:59, batch 0 finished 05:15:56 | log file timestamps and the run’s own printed times |
| 326 GB, 136 GB remaining, 18 batches, batch sizes 7.4 to 7.8 GB | plan plan_hessian_probe.json + tensor shapes in the model’s *.safetensors headers; size = (in² + out²) × 4 bytes |
| 115.4 GB GPU limit, 499000 handle limit | mx.device_info() (MLX 0.32.2) |
| 56 GB model | sum of the model’s weight files on disk (55.6 GB) |
| Batch durations 3 h 00 to 3 h 45 | differences between log-file modification times |
| “About 2 minutes” to build the tables, 4 minutes per extra batch | Estimates. Batch 1 reached chunk 6 in 54 s including loading; the current run finished table-building about 2 minutes after starting |
| Sensor check figures | a live test on this machine (2 GB allocation), read with ioreg |
One large table of numbers inside the model. This model has 497; 487 are corrected by this run.
Storing each number with fewer bits (here 4 to 8 instead of 16) so the model is smaller, at some cost in accuracy.
The compression method used here. It corrects the rounding error using measured sensitivities instead of rounding blindly.
A table, measured from sample text, of how much the model’s output changes when a tensor’s numbers are nudged. Big: tens of thousands squared entries.
A short piece of sample text passed through the model to measure sensitivity. 24 are used, across six kinds of text.
A group of tensors measured together in one pass, then corrected one by one.
One running copy of a program. Each batch is its own, so the GPU starts clean each time.
The list of finished tensors saved on disk, so a restart skips them.
The chip doing the heavy arithmetic, and the memory the CPU and GPU share on this Mac.
The GPU’s built-in watchdog decided some work stopped progressing and reset the GPU.
Memory that is really available right now for the run.
A ceiling on batch size remembered from earlier failures; it rises again after two clean batches.
A trillion arithmetic operations. A rough measure of how much GPU work a batch does per chunk.
One that recurs identically every time, like a bug. Retrying cannot help.
macOS ending programs to reclaim memory when it runs short.
The measurements the controller records so a future failure can be diagnosed from data.
Each batch reloads the model and re-runs the measurement pass. At 2 GB the run needs 105 batches instead of 18, about 10% more time, and the evidence does not show size caused the crash.
A lockup that repeats after a clean retry and a halved batch is not something more of the same will fix; the run stops and tells you rather than burning hours.
No. Each tensor is saved as it finishes. A failure costs at most the tensor in progress plus a model reload.