← Back to index
Engineering record · written 2026-09-25 · describes the controller as shipped that day

How the YAQA build now protects itself: adaptive batching

A compression run on this Mac takes days. On the night of 25 September the GPU reset itself three hours into a run, and the run sat stopped until someone noticed. This page explains what the run does, what went wrong, what the evidence says caused it (and what it does not), and how the new controller senses the machine, sizes each step to fit, and recovers on its own.

How to read this page. No prior knowledge is assumed: every term is explained where it first appears, and a glossary sits at the end. Every number comes from a real file on this machine, listed in section 10. Two interactive tools in section 7 let you try the controller’s rules yourself. Where the evidence stops, the page says so in an amber box.
487
large tables of numbers (tensors) that must each be measured and corrected
326 GB
of measurement tables if built all at once, against a 115 GB GPU limit
≈ 3.3 h
per batch, measured from real log timestamps
2 in 5 days
identical GPU firmware resets, on 20 and 25 September

Start here: the whole story in one minute

Read this section and you have the point. The rest is the detail and the evidence.

  1. The job. A language model is billions of numbers stored in large tables (tensors). To make it smaller, each table is compressed to fewer bits. The compression method here (YAQA) corrects the rounding error, and to do that it first needs a table of sensitivities for every tensor (a Hessian): how much the model’s output changes when that tensor’s numbers are nudged.
  2. The constraint. All those sensitivity tables together are far too big for the machine’s memory, so the work is cut into batches: a group of tensors is measured together, corrected, saved, and the next group starts. The run is about 18 more batches at roughly 3.3 hours each.
  3. What broke. During batch 1, at 05:16 on 25 September, the GPU (the chip doing the heavy arithmetic) reset itself. The script that runs the batches printed an error and exited. Nobody was told. The night was lost.
  4. What the evidence says. macOS recorded a firmware-detected lockup. The same fingerprint appeared five days earlier. The batch that crashed was the same size, on the same data, as batches that had run fine for hours. So it was not the batch size and not memory. What triggered it is unknown.
  5. The fix. The script now checks the machine before each batch, sizes the batch to fit what is really free, measures as it goes, works out why anything failed, retries at most twice with a rule that fits that kind of failure, and if it still fails it stops, writes a report and sends a notification. It never shrinks a batch on a guess, and never retries a genuine bug.
What is and is not proven
  • Verified Every rule behaves as designed under ten injected failures; the sensors respond to real load; the diagnosis of the crash (a firmware lockup, not a size or memory effect).
  • Unknown What actually triggers the lockup. The crash reports do not say which command stalled.
  • Not yet shown That the controller prevents the next lockup. It has not yet met a real one. It is built to recover in seconds and to record enough to diagnose the next event properly.

1. Why the work is cut into batches

The arithmetic that forces it.

The machine is a Mac with 128 GB of memory that the CPU and the GPU share (unified memory). The operating system lets the GPU use at most about 115 GB of it. The model’s own weights, kept loaded the whole time, take 56 GB. That leaves roughly 60 GB, for everything else.

For this plan, the sensitivity tables for all 487 tensors add up to 326 GB: more than five times the room available. So they cannot be built together. The run measures a group of tensors at a time, using a budget of 8 GB of tables per group by default, corrects them, saves them, and moves on.

All Hessian tables at once 326 GB Most the GPU may use 115 GB Model weights (always loaded) 56 GB Room left for Hessians 60 GB One batch (default budget) 8 GB
Memory arithmetic for this run. Sources: per-tensor shapes read from the model’s weight files (table size = (inputs² + outputs²) × 4 bytes per tensor), the GPU limit from mx.device_info(), the weight size from disk.
Term: batch

A group of tensors whose sensitivity tables are built together in one pass of the model over sample text. After the pass, each tensor is corrected one at a time and its result is saved to disk straight away.

2. One batch, step by step

What happens inside a batch, and how long each part takes.

1 · Fresh processA new program run perbatch, so the GPU startsfrom a clean state 2 · Measure24 sample-text chunks passthrough the whole model;tables fill up. About 2 min 3 · Correct, tensor by tensorEach tensor is compressed at3, 4, 5, 6 and 8 bits witherror correction. About 3 h 4 · SaveEach result iswritten to diskimmediately then a new process starts with the next group of tensors, skipping everything already saved The crash on 25 September happened in step 2, six chunks in.

Two design choices matter for everything that follows:

0h 1h 2h 3h 4h 3 h 10 batch A (Sep 19) 3 h 40 batch B 3 h 14 batch C 3 h 45 batch D 3 h 22 batch E (Sep 20) 3 h 00 today's batch 0 Almost all of each bar is per-tensor correction. The Hessian collection that crashed takes about2 minutes (see section 2), a sliver at this scale.
Real batch durations: the gap between the final write of one batch log and the next (five batches on 19–20 September) and today’s batch 0 (02:15 to 05:15). Measured from log-file timestamps, so each includes model loading. At this pace the 201 remaining tensors need about 18 batches, roughly 59 hours.

3. The night it stopped

What happened, in order, with the real timestamps.

THE NIGHT (02:00 to 11:30) 02:15 batch 0 startsbatch 0 runs 3 h 05:16:53 GPU reset run stopped; nobody was told until someone looked 11:02:54 new run, controller active 02:0005:1611:03 ZOOM: 05:15:50 to 05:17:00 (70 seconds) 05:15:56batch 0 finishes 05:15:59batch 1 starts, 3 s later loads the model, runs chunks 1 to 6 (54 s) 05:16:53 firmware-detected lockup GPU reset; script exits

The old script did exactly one thing on any failure: print the last lines of the log and exit. It had no way to tell a one-off GPU reset from a real bug, no retry, and no way of telling you. That, not the reset itself, is what turned a few minutes into a lost night.

What macOS wrote down

Whenever the GPU resets, macOS saves a small report. There are two, five days apart, and they match:

Report file (in /Library/Logs/DiagnosticReports/)TimeReason it recordsSignature
gpuEvent-python3.11-2026-09-20-121024.ips20 Sep 12:10:24firmware-detected lockup579 · guilty_dm 3 · restart_reason 4
gpuEvent-python3.11-2026-09-25-051653.ips25 Sep 05:16:53firmware-detected lockup579 · guilty_dm 3 · restart_reason 4
Term: firmware-detected lockup

The GPU has its own small built-in program (firmware). It watches the work it is given, and if a piece of work stops making progress it declares a lockup and resets the GPU. That is different from running out of memory, which produces a different report and a different error.

4. What it was not: not the size, not the data, not memory

The natural reflex is “make the batches smaller.” The evidence says that would not have prevented this.

Same data, same computation

The crashed batch and the batch that finished use the same sample text in the same order. The loss (how wrong the model is on each chunk) is the same number to four decimal places for the six chunks the crashed batch completed. Nothing about the input differed.

0 1 2 3 4 1 2 3 4 5 6 7 8 calibration chunk number (of 24) model loss on that chunk GPU reset here: batch 1 stops 3.7357 batch 0 (finished all 24 chunks) batch 1 (the one that crashed)
Loss per calibration chunk, from logs/batch_0.log and logs/batch_1.log. The rings sit exactly on the dots for chunks 1 to 6.

Same size as batches that worked

The batches are filled to the 8 GB budget, so nearly all of them are the same size. The batch that crashed was 7.76 GB and about 22 trillion arithmetic operations (22 TFLOP) per chunk. Batches of that size ran for hours without incident on 19 and 20 September.

0 2 4 6 8 1 7.8 2 7.8 3 7.8 4 7.8 5 7.7 6 7.8 7 7.7 8 7.8 9 7.8 10 7.7 11 7.8 12 7.8 13 7.5 14 7.8 15 7.5 16 7.8 17 7.8 18 4.4 the 18 batches still to run, in order (1 = the batch that crashed) Hessian memory (GB) crashed batch: 7.76 GB, 22.3 TFLOP
The 18 batches still to run, reconstructed from the plan and the model’s tensor shapes (this reconstruction reproduces the log’s 14 tensors and “18 batches remaining” exactly). The crashed batch (red) is not larger than the rest.

Not memory

After the crash, memory showed 95% free, and there was no out-of-memory event. There was a memory event the night before (a different Python program was by far the largest memory user and macOS ended work), so memory contention is real. It is a separate failure with a separate rule below.

Failure type A

GPU firmware lockup

Reason on the report: firmware-detected lockup. Seen twice. Cause unknown. Not related to size in the evidence.

Failure type B

Memory pressure

Seen on 24 September: a Python 3.14 process was by far the largest memory user and macOS ended work. Cause known; the fix is to size batches to what is really free.

Failure type C

A genuine bug

For example a number that is not a number (NaN) or a failed check in the code. Retrying gives the same failure, so the right response is to stop and say so.

A correction on the record

The first suggestion made during this investigation was to lower the batch budget to 4 GB. That was a guess and was withdrawn once the crash report and the identical-data evidence were read. A smaller batch does not address a firmware lockup that struck a batch the same size as ones that worked.

What is still unknown

The crash reports say a lockup happened; they do not say which piece of GPU work stalled. Candidates: a very long single calculation, another program using the GPU at that moment, or power and heat conditions. This is why the controller is built to defend and to record, not to pretend it has removed the cause.

5. What was wrong with the old script

Not the batching. The lack of any measurement, judgement or alert around it.

Before

  • Batch size: a fixed number (8 GB), never checked against the machine
  • Measured: nothing (the program never read memory or GPU load)
  • On any failure: print the log tail and exit
  • A one-off GPU reset and a real bug looked the same
  • Nobody was told; the run sat stopped
  • Failed-attempt logs could be overwritten

After

  • Batch size: chosen from real free memory and what was learned from failures
  • Measured: GPU busy %, free memory, peak GPU memory, seconds per chunk
  • On failure: classify it, apply that class’s rule, at most 2 retries
  • A real bug stops at once; a lockup is retried cleanly first
  • Stops loudly: a report file and a macOS notification
  • Every failed attempt’s log is kept

6. The controller: sense, size, run, classify, retry

It lives inside the same script (run_full_yaqa.sh) and changes nothing about the maths.

1 · SENSEGPU idle (20% or less, 3 checks)Real free memory (vm_stat) 2 · SIZE THE BATCHthe smallest of: the default,what it learned, what fits 3 · RUN + MEASUREits own fresh process; watchesGPU memory and chunk progress Finished OK? yes 4 · RECORD, NEXT BATCHsave peak memory, seconds/chunk2 clean in a row: cap grows 1.5x no 5 · CLASSIFY THE FAILUREfrom the log text and exit code GPU LOCKUP or SILENT STALLRetry 1: same size, once theGPU is idle againRetry 2: half the sizethen back to step 3 OUT OF MEMORYRetry 1 and 2: halve the size,never more than freshlysensed free memory allowsthen back to step 3 GENUINE BUGNo retry. Stop immediately:the same input fails thesame way, so retrying onlywastes hours After the 2nd failed retry (or a bug): write FAILED.txt, send a macOS notificationYou find out in minutes, not after a night. Your own Ctrl-C is never retried.
What actually steers decisions, and how well supported each signal is
  • Steers a decision, strongest support: the exit code and log text, classified. The real 05:16 crash log matches the lockup pattern, and ten injected-failure tests pass.
  • Steers a decision, hypothesis: the GPU-idle gate (the sensor works; that waiting prevents a lockup is unproven) and the free-memory sizing (an estimate, not binding so far: the batch stayed at 8 GB).
  • Steers a decision, tested on a fake only: the silent-stall kill.
  • Recorded only, decides nothing: peak GPU memory and seconds per chunk. Which memory figure matters is unknown.

So today’s proven value is fast recovery and reporting, which works whatever the lockup trigger turns out to be. A census on 25 September showed 20 or more programs sharing the GPU (WindowServer, Chrome, Safari, Spotlight and more) that the busy-percentage gate cannot tell apart, which is why a black-box recorder of the GPU client list and memory figures now runs alongside the build.

Step by step, in plain words

Step 1

Sense

Before starting (and before every retry) it checks that the GPU is quiet: the busy figure macOS reports, from every program, must be 20% or less on three checks in a row, ten seconds apart. It waits up to ten minutes; if the GPU is still busy it says so, sends a notification and carries on. It also reads how much memory is really free.

Step 2

Size

fit = (free memory − model − 6) ÷ 1.5, and the batch budget is the smallest of the default (8 GB), the cap it has learned from earlier failures, and fit, never below 1.4 GB (one large tensor). Section 7 lets you try it.

Step 3

Run and measure

The batch runs as its own process. From outside, every 10 seconds, the controller records GPU memory and counts finished chunks. If no new chunk appears for 10 minutes while the sensitivity tables are still being built, it treats the process as stalled and stops it.

Step 4

Record

On success it saves peak GPU memory and seconds per chunk. After two clean batches in a row a lowered cap is raised by 50%, back up to the default: it recovers instead of staying cautious forever.

Step 5

Classify

The failed attempt’s log is read for a GPU hang message, an out-of-memory sign or an interrupt, and the exit code is checked. The log is kept in logs/failed/.

Then

Retry or stop

At most two retries, following the rule for that class (the diagram above). A stall check applies only while sensitivity tables are being built, so the three-hour correction phase is never interrupted by it.

Why the first retry of a lockup is the same size

The lockup is not repeatable: identical work had already passed. Shrinking on the first failure would be acting on a guess. So the first retry is the same batch after the GPU is idle again; only a repeated lockup halves the size. Out-of-memory failures shrink at once, because there the size is the obvious suspect.

GPU busy (%) idle10 under a 2 GB test load100 2 s after31 GPU memory in use (GB) idle1.85 under a 2 GB test load4.93 2 s after1.88
The sensors respond to real load. Check performed on this machine: a 2 GB test allocation and a chain of matrix multiplications moved the GPU busy figure from 10% to 100% and GPU memory in use from 1.85 GB to 4.93 GB, and both fell back within about two seconds of the load ending.

7. Try it: two interactive tools

Both use the controller’s real formulas. The second reproduces the retry rules the automated tests check.

Tool 1: how big will the batches be?

Model weights 56 GB and the 6 GB margin are as in the real controller. The 201 tensors are the ones not yet finished when the current run started; batches are packed in the same order as the real code. The extra-time figure assumes about 4 minutes of fixed overhead per batch (model load plus sensitivity measurement), an estimate from the logs; the correction time itself does not depend on how many batches there are.

8 GB (default) 18 batches (baseline) 4.67 GB 34 batches, about +1.1 h 4 GB 36 batches, about +1.2 h 2 GB 105 batches, about +5.8 h 1.4 GB (one tensor) 106 batches, about +5.9 h
Number of batches the 201 remaining tensors need at each budget, computed with the real packing rule. Smaller batches are cheap down to about 4 GB and cost roughly 10% more time at 2 GB. This is why the controller shrinks only when a signal says to.

Tool 2: what does it do when something fails?

Pick a failure and watch the decisions. The batch starts at the size the first tool chose.

These sequences are the automated test cases T2 to T7 of test_adaptive_controller.sh. The tool applies the same rules in the page; the test script is what proves the real script does the same, so run it after any edit to run_full_yaqa.sh.

8. How we know it works, and what we do not know

Ten scenarios, run against the real script with a fake batch program that fails on command.

TestWhat is injectedWhat must happenResult
T1Nothing failsThree batches at 8 GB, no notificationPASS
T2One GPU lockupRetry at the same size: 8, 8PASS
T3Two lockups8, 8, then half: 4; cap 4 rememberedPASS
T4Three lockupsStop after 2 retries; report and notificationPASS
T5A genuine bug (NaN)No retry, stop at once, notifyPASS
T6A silent stallKilled after the stall limit, retriedPASS
T7Out of memory twice8, 4, 2PASS
T8Only 69 GB freeBatch sized down to 4.67 GB before startingPASS
T9Another program keeps the GPU 60% busyWaits, warns, notifies, proceedsPASS
T10Ctrl-C mid-batchExit code 130, quiet, no leftover processPASS

Two real defects the tests found

What the evidence does not cover

QuestionStatusWhat would settle it
What triggers the firmware lockup?UnknownThe next event, with the recorded peak GPU memory, chunk seconds and failed-attempt logs to compare against
Does the controller prevent the next lockup?Not shownTime. It has been built to recover, not proven to prevent
Are 20% idle, 10-minute stall limit, the 1.5 factor, 6 GB margin and halving the right numbers?Engineering choicesFitting them to the recorded telemetry after a few events
Which GPU-memory figure tracks danger? The controller records “in use” (first real batch, 25 Sep: peak 76.4 GB during Hessian collection, 17.9 GB mid-correction); the driver’s “allocated” figure read 105.9 GB, 92% of the 115.4 GB ceiling, at the same momentUnknownRecord both at the next failure and compare; memory pressure itself stayed green
Does the stall watchdog work against a real hung GPU process?Tested on a fake onlyA real silent hang
Does the live dashboard show the controller?Not yetA dashboard panel reading the decision log (not built)
What if the terminal closes or the Mac loses power?Not handledThe run stops; re-run the same command. There is no automatic restart after a reboot

9. Running it and reading what it leaves

The command is unchanged. The controller is automatic.

Start the run (this is the exact command used for the probe build):

PROJECT_ROOT=/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara

"$PROJECT_ROOT/scripts/yaqa_port/run_full_yaqa.sh" \
    /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE \
    --skip-lm-head-gptq \
    --plan "$PROJECT_ROOT/04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.json" \
    --n-calibration 4 \
    --incoherence none \
    --correct-mtp \
    --godmode-multi-bit-checkpoint \
    --godmode-candidate-bits 3,4,5,6,8

If it stopped, run the same command again: finished tensors are skipped and what it learned is kept. Watch its decisions live:

tail -f /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume/adaptive_events.log

Prove the controller still behaves after any edit to the script (about 2 minutes, no GPU, no real data):

cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port && ./test_adaptive_controller.sh
Do not edit the script while a run is in progress

The shell reads a running script piece by piece, so changing the file mid-run can corrupt the run. Edit between runs, then run the test above.

What it leaves behind

<output>.yaqa_resume/ manifest.json every finished tensor (the resume record) adaptive_events.log every decision, with its reason, timestamped adaptive_state.json what it has learned: cap, clean streak, last peak memory, last failure FAILED.txt only if it gave up: class, detail, log tail, newest GPU report logs/ batch_N.log the batch that is running or finished failed/ batch_N_attemptK_<class>_<time>.log every failed attempt, never overwritten

10. Where every number comes from

Nothing on this page is estimated without saying so.

NumberSource
Firmware lockup, signature 579, guilty_dm 3, restart_reason 4, on 20 and 25 Sep/Library/Logs/DiagnosticReports/gpuEvent-python3.11-2026-09-20-121024.ips and ...-2026-09-25-051653.ips
Chunk losses 2.1989, 1.2555, 0.9013, 1.3176, 3.7357, 2.1532<resume>/logs/batch_0.log and logs/batch_1.log (batch_1 is kept as the first failed log)
Crash at 05:16:53, batch 1 started 05:15:59, batch 0 finished 05:15:56log file timestamps and the run’s own printed times
326 GB, 136 GB remaining, 18 batches, batch sizes 7.4 to 7.8 GBplan plan_hessian_probe.json + tensor shapes in the model’s *.safetensors headers; size = (in² + out²) × 4 bytes
115.4 GB GPU limit, 499000 handle limitmx.device_info() (MLX 0.32.2)
56 GB modelsum of the model’s weight files on disk (55.6 GB)
Batch durations 3 h 00 to 3 h 45differences between log-file modification times
“About 2 minutes” to build the tables, 4 minutes per extra batchEstimates. Batch 1 reached chunk 6 in 54 s including loading; the current run finished table-building about 2 minutes after starting
Sensor check figuresa live test on this machine (2 GB allocation), read with ioreg

11. Glossary

Tensor

One large table of numbers inside the model. This model has 497; 487 are corrected by this run.

Quantization

Storing each number with fewer bits (here 4 to 8 instead of 16) so the model is smaller, at some cost in accuracy.

YAQA

The compression method used here. It corrects the rounding error using measured sensitivities instead of rounding blindly.

Hessian

A table, measured from sample text, of how much the model’s output changes when a tensor’s numbers are nudged. Big: tens of thousands squared entries.

Calibration chunk

A short piece of sample text passed through the model to measure sensitivity. 24 are used, across six kinds of text.

Batch

A group of tensors measured together in one pass, then corrected one by one.

Process

One running copy of a program. Each batch is its own, so the GPU starts clean each time.

Resume record (manifest)

The list of finished tensors saved on disk, so a restart skips them.

GPU / unified memory

The chip doing the heavy arithmetic, and the memory the CPU and GPU share on this Mac.

Firmware lockup

The GPU’s built-in watchdog decided some work stopped progressing and reset the GPU.

Headroom

Memory that is really available right now for the run.

Learned cap

A ceiling on batch size remembered from earlier failures; it rises again after two clean batches.

TFLOP

A trillion arithmetic operations. A rough measure of how much GPU work a batch does per chunk.

Deterministic failure

One that recurs identically every time, like a bug. Retrying cannot help.

Jetsam

macOS ending programs to reclaim memory when it runs short.

Telemetry

The measurements the controller records so a future failure can be diagnosed from data.

Questions a first-time reader asks

Q

Why not just always use small batches?

Each batch reloads the model and re-runs the measurement pass. At 2 GB the run needs 105 batches instead of 18, about 10% more time, and the evidence does not show size caused the crash.

Q

Why at most two retries?

A lockup that repeats after a clean retry and a halved batch is not something more of the same will fix; the run stops and tells you rather than burning hours.

Q

Will it lose finished work?

No. Each tensor is saved as it finishes. A failure costs at most the tensor in progress plus a model reload.

Written 2026-09-25 by Hakim Ghelab, VegaLaboratories LTD. A dated record of the controller as shipped that day; numbers are as of that day and are not recomputed live.