← Back to index
Live on a real 27B model, right now · 2026-09-07

A complete quantization platform,
built end to end for Apple Silicon.

YAQA‑UMA names four original systems built as one: cascaded KL‑sensitivity measurement, Pareto/MILP bit‑width allocation, stratified calibration, and an independent MLX implementation of Cornell's YAQA rounding method — every line of the rounding step derived from the paper's own equations, not a translated port.

99.84%
mlp.up_proj
99.87%
mlp.down_proj
99.85%
mlp.gate_proj
99.57%
attn.out_proj
Real error reduction vs. naive rounding · captured live during the production run
The whole arc, not just tonight

From "does an MLX port even exist" to a live 27B model.

FIRST

Zero prior art

No MLX implementation of YAQA existed anywhere — checked directly against real GitHub search, not assumed.

PROVEN

75% error reduction

The first real layer of a real 27B model, YAQA‑UMA vs. naive rounding — the result that justified building the rest.

SOLVED

A 557GB memory wall

The reference design needs 557GB held at once. Rebuilt as a memory‑bounded, resumable batching architecture that runs on this machine's real 128GB.

LIVE

99.81% average, right now

11 real tensors in, 99.40–99.93% error reduction each — the full 363‑tensor model, reprocessing live.

HG
Hakim Ghelab
Founder · VegaLaboratories LTD
Zero to live4 days, Sep 3–6 2026
Original systems built4, integrated as one
Reference code reusedNone — derived from equations
Root cause foundConfirmed, not guessed
Tensors protected363 / 363

YAQA‑UMA isn't one port. It's a platform — four original systems, built end to end by one engineer, that only work because they were built to fit together.

Before a single weight gets rounded, every tensor in the model is already scored by a cascaded KL‑sensitivity measurement system — a real, original pipeline for finding out which parts of a 27B‑parameter model actually matter, layer by layer. That score feeds a Pareto‑pruning and MILP allocation solver, built from scratch, that turns "which tensors matter" into an actual, optimal per‑tensor bit‑width plan — not a uniform guess. The calibration data that both of those systems and the final rounding step all depend on comes from a purpose‑built stratified calibration system, replacing naive flat sampling with real domain balance across the calibration corpus.

YAQA‑UMA is the name for that whole platform, not just its last stage. The rounding step itself was the hardest part to build: the reference algorithm had to be understood well enough to reject what didn't transfer — its trellis‑coded deployment format assumes a custom CUDA codebook that doesn't exist on Apple hardware, so it was deliberately left out rather than forced in. What replaced it was derived directly from the paper's own equations, verified independently at every stage, and hardened through real failures found the hard way: a missing incoherence step, a batching architecture rebuilt three times to survive Apple Silicon's actual memory limits, and — most recently — a silent 16‑bit precision bug that made one tensor 300× worse than doing nothing at all.

SYSTEM 1
Built the cascaded KL‑sensitivity pipeline — real per‑tensor scoring of what actually matters in the model.
SYSTEM 2
Built the Pareto + MILP allocator that turns those scores into an optimal bit‑width plan, tensor by tensor.
SYSTEM 3
Built stratified calibration, replacing flat sampling with real domain balance across every stage that touches calibration data.
SYSTEM 4
Ported YAQA's rounding math to MLX from its own equations — rejecting the parts that couldn't transfer, rebuilding the memory architecture for Apple Silicon, and finding and fixing a real 16‑bit precision bug the original environment never surfaced.
UNIFIED
All four, combined, is what YAQA‑UMA actually names — live, right now, reprocessing the full model.
Standing on real shoulders

None of this exists without Cornell RelaxML's real work.

Albert Tseng, Zhaofeng Sun, and Christopher De Sa didn't just publish a paper — they solved a genuinely hard open problem in model compression and proved it, rigorously, with real numbers: a method that's provably better than the prior state of the art (GPTQ/LDLQ) and empirically backs that up with roughly a 30% real error reduction. That's the kind of research contribution that makes everything downstream of it possible. Full credit, always: arXiv:2505.22988, Cornell RelaxML, May 2025.

Correct attribution, stated plainly

One real algorithm. One real port. Neither one pretends to be the other.

Precision here isn't modesty — a claim that's wrong in either direction is a claim that doesn't survive someone checking it.

The published method

YAQA — Cornell RelaxML

  • Albert Tseng, Zhaofeng Sun, Christopher De Sa · arXiv:2505.22988
  • The Sketch B Hessian formula and the two‑sided LDLQ rounding update
  • Original reference stack: PyTorch + CUDA
  • Deployed via a trellis‑coded (QTip) codebook, not affine quantization
  • Targets multi‑GPU systems with CPU‑offload for large Hessians
“Provably better than GPTQ/LDLQ, ≈30% error reduction.”
— abstract, verbatim
This implementation

YAQA‑UMA — Hakim Ghelab, VegaLaboratories LTD

  • Independently derived from the paper's own equations — no PyTorch/CUDA source reused
  • Cross‑validated against independent NumPy references at every stage
  • Native MLX affine packing — no trellis codebook, no custom loader required
  • A real, found‑and‑fixed precision bug the original paper never had to confront
  • Adaptive‑damping safety gate: never ships a correction worse than naive
  • Memory‑bounded batching architecture built for unified memory, not multi‑GPU
Unified Memory Architecture — the constraint this port is built around, distinct from the reference's own design.
The broader system

YAQA‑UMA is one stage in a larger, original pipeline.

Every stage below the rounding step is this project's own independent design — YAQA‑UMA plugs into it as the final, whole‑model‑aware correction.

STAGE 1
Cascaded sensitivity measurement
Per‑tensor KL‑sensitivity checkpointing across the model.
Original system
STAGE 2
Pareto pruning + MILP allocation
Mixed‑integer solve for the real per‑tensor bit‑width plan.
Original system
STAGE 3
Stratified calibration
Domain‑balanced calibration text, replacing naive flat sampling.
Original system
STAGE 4
YAQA‑UMA rounding
Whole‑model‑aware correction — the independent MLX port documented here.
This page
Why whole‑model awareness matters

GPTQ sees one tensor. YAQA‑UMA sees the whole model.

Both correct rounding with a Hessian. The difference is what that Hessian is actually measuring — a local reconstruction error, or the real effect on the model's own output.

01

Two‑sided correction

Both an input‑side and output‑side curvature estimate — not just local reconstruction error, real propagation to the model's own output.

02

Found a real precision bug

Production curvature was silently accumulating in 16‑bit. Diagnosed, fixed, and verified against a 64‑bit reference — a real defect the original CUDA implementation's own environment never surfaced.

03

A safety net the paper doesn't need

Adaptive damping with a hard, automatic fallback to naive rounding — built because unified‑memory constraints make some tensors genuinely harder to correct than the reference's own hardware ever has to face.

Why this matters beyond one model

Quantization isn't a research trick — it's the difference between renting a GPU cluster and running a 27B‑class model on the machine already on your desk.

Whole‑model‑aware correction is what makes aggressive compression usable in production instead of just publishable. That has real consequences past this one model.

SPEED

Acceleration on unified memory

Apple Silicon inference is memory‑bandwidth‑bound, not compute‑bound — a 5‑bit model moves roughly a third of the bytes per token that a 16‑bit model does, which is where real throughput gains on this hardware actually come from.

COST

Cloud GPU spend, replaced

A model that used to require multi‑GPU cloud infrastructure to hold in memory now runs on a single machine's unified memory — the difference between an hourly cloud bill and hardware you already own.

QUALITY

Compression without the naive‑rounding tax

Naive quantization degrades unevenly — some tensors break badly. Whole‑model‑aware correction is what lets a model actually reach these bit‑widths without silently losing quality where it matters most.

PRIVACY

On‑device, not on someone else's server

A model small enough to run locally is a model whose prompts and outputs never have to leave the machine — relevant for anything cost‑, latency‑, or data‑residency‑sensitive.

60–70
tok/s · Qwen3.8‑27B, YAQA‑UMA 5bpw
0
generation errors observed
1
consumer Apple Silicon Mac — no cluster
5bpw
vs. 16‑bit original weights
Reported by the developer during interactive local testing on this machine · not yet captured by an automated, logged benchmark harness