YAQA‑UMA names four original systems built as one: cascaded KL‑sensitivity measurement, Pareto/MILP bit‑width allocation, stratified calibration, and an independent MLX implementation of Cornell's YAQA rounding method — every line of the rounding step derived from the paper's own equations, not a translated port.
No MLX implementation of YAQA existed anywhere — checked directly against real GitHub search, not assumed.
The first real layer of a real 27B model, YAQA‑UMA vs. naive rounding — the result that justified building the rest.
The reference design needs 557GB held at once. Rebuilt as a memory‑bounded, resumable batching architecture that runs on this machine's real 128GB.
11 real tensors in, 99.40–99.93% error reduction each — the full 363‑tensor model, reprocessing live.
Before a single weight gets rounded, every tensor in the model is already scored by a cascaded KL‑sensitivity measurement system — a real, original pipeline for finding out which parts of a 27B‑parameter model actually matter, layer by layer. That score feeds a Pareto‑pruning and MILP allocation solver, built from scratch, that turns "which tensors matter" into an actual, optimal per‑tensor bit‑width plan — not a uniform guess. The calibration data that both of those systems and the final rounding step all depend on comes from a purpose‑built stratified calibration system, replacing naive flat sampling with real domain balance across the calibration corpus.
YAQA‑UMA is the name for that whole platform, not just its last stage. The rounding step itself was the hardest part to build: the reference algorithm had to be understood well enough to reject what didn't transfer — its trellis‑coded deployment format assumes a custom CUDA codebook that doesn't exist on Apple hardware, so it was deliberately left out rather than forced in. What replaced it was derived directly from the paper's own equations, verified independently at every stage, and hardened through real failures found the hard way: a missing incoherence step, a batching architecture rebuilt three times to survive Apple Silicon's actual memory limits, and — most recently — a silent 16‑bit precision bug that made one tensor 300× worse than doing nothing at all.
Albert Tseng, Zhaofeng Sun, and Christopher De Sa didn't just publish a paper — they solved a genuinely hard open problem in model compression and proved it, rigorously, with real numbers: a method that's provably better than the prior state of the art (GPTQ/LDLQ) and empirically backs that up with roughly a 30% real error reduction. That's the kind of research contribution that makes everything downstream of it possible. Full credit, always: arXiv:2505.22988, Cornell RelaxML, May 2025.
Precision here isn't modesty — a claim that's wrong in either direction is a claim that doesn't survive someone checking it.
Every stage below the rounding step is this project's own independent design — YAQA‑UMA plugs into it as the final, whole‑model‑aware correction.
Both correct rounding with a Hessian. The difference is what that Hessian is actually measuring — a local reconstruction error, or the real effect on the model's own output.
Both an input‑side and output‑side curvature estimate — not just local reconstruction error, real propagation to the model's own output.
Production curvature was silently accumulating in 16‑bit. Diagnosed, fixed, and verified against a 64‑bit reference — a real defect the original CUDA implementation's own environment never surfaced.
Adaptive damping with a hard, automatic fallback to naive rounding — built because unified‑memory constraints make some tensors genuinely harder to correct than the reference's own hardware ever has to face.
Whole‑model‑aware correction is what makes aggressive compression usable in production instead of just publishable. That has real consequences past this one model.
Apple Silicon inference is memory‑bandwidth‑bound, not compute‑bound — a 5‑bit model moves roughly a third of the bytes per token that a 16‑bit model does, which is where real throughput gains on this hardware actually come from.
A model that used to require multi‑GPU cloud infrastructure to hold in memory now runs on a single machine's unified memory — the difference between an hourly cloud bill and hardware you already own.
Naive quantization degrades unevenly — some tensors break badly. Whole‑model‑aware correction is what lets a model actually reach these bit‑widths without silently losing quality where it matters most.
A model small enough to run locally is a model whose prompts and outputs never have to leave the machine — relevant for anything cost‑, latency‑, or data‑residency‑sensitive.