← Back to index YAQA Pipeline Comparison
MLX_OptiQ · yaqa_port

YAQA correction vs. real deployment format

Real reference repo (Cornell-RelaxML/yaqa) vs. this MLX port — where they diverge, and why.

Real reference pipeline (PyTorch / CUDA)

Calibration dataforward + backward passCollect H_I, H_Otwo-sided HessianHadamard rotate H_I, H_Ovia SU, SVLDLQ_2hessblock-sequential correctionRound each block:TRELLIS CODEBOOK (VQ)Save: trellis idx+ SU + SV + Hadamard buffersInference: custom CUDA kernelapplies SU/SV/Hadamardevery forward pass

This MLX port (current)

A: carry rotationB: drop rotation(matches this project's GPTQ)Calibration dataforward + backward passCollect H_I, H_Osame formula (MLX)Hadamard rotate H_I, H_Ovia SU, SV (verified)LDLQ_2hesssame algorithm (yaqa_core.py)Round each block:PLAIN AFFINE (mx.quantize)Deploy rotatedresult?New rotation-aware module+ custom loader -- NOT BUILTStandard QuantizedLinearreuses GPTQ save path AS-IS
Why rotation mattered less than it looked

In the real repo, SU/SV/Hadamard rotation exist to condition weights for a fixed trellis codebook — a VQ scheme with no per-group scale adaptation, so uncontrolled outliers are catastrophic for it. mx.quantize is a per-group affine quantizer — it already fits its own scale/zero-point to each group's actual range, absorbing most of what rotation buys. This project's own GPTQ baseline (08_gptq_apply_plan.py) is also affine, also unrotated, and already ships in production. Rotation stays real and beneficial for the correction step (Step 4's 75% error reduction used it) — the open question was only whether it's worth carrying into the deployed representation, which needs new, unverified machinery to do. Decision: v1 deploys unrotated (Option B), matching how this project's own GPTQ ships.

Not portable to MLX / not yet built Deliberate substitution, now made explicit Real, standard, reuses proven code