The main milestones of the project, read left to right. Scores are real n=200 or n=100 benchmark results.
The story at a glance
Milestones
Start. MLX OptiQ's allocator is the baseline. We found it ignores inversions and Pareto dominance, so we built our own MILP solver with boundary and regional floors.
Stratified calibration. OptiQ's calibration sampled flat. We fixed it with 24 windows spread across six writing domains. This is the calibration every later build uses.
YAQA on MLX. The rounding correction from the YAQA paper is ported to MLX and extended to the MTP speculative-decoding head. YAQA only handles rounding; the allocation work is ours.
Cascaded KL. Sensitivity measured with upstream layers already quantized. Round 1 gives plan P1, which later rounds barely change. Model B, built from P1, scored 92.14 at n=100.
Hessian and GODMODE. Real forward and backward passes give curvature-weighted error for every tensor at every bit width. This data feeds the solver directly.
HESSIAN-PROBE. A probe build that filled in the missing Hessian scores. Its n=200 mean is 93.04, the highest n=200 mean so far.
GODMODE-v2. The solver is fed pure Hessian data, with fixed rules for lm_head and the first and last layers. It reaches the top speed tier at 55.9 tok/s.
GODMODE-v2 at n=200. Mean 92.58 against 93.04 for HESSIAN-PROBE. The gap is inside the confidence intervals, so this is parity, not a win.
Still open. Stratified at n=200, the V3 build, and the BF16 reference with KL against it.
Reading guide
Read
For
This timeline
The arc of the project
HESSIAN_ZERO_TO_EXPERT.html
How the Hessian method works, from first principles