Real reference repo (Cornell-RelaxML/yaqa) vs. this MLX port — where they diverge, and why.
In the real repo, SU/SV/Hadamard rotation exist to condition weights for a fixed trellis codebook — a VQ scheme with no per-group scale adaptation, so uncontrolled outliers are catastrophic for it. mx.quantize is a per-group affine quantizer — it already fits its own scale/zero-point to each group's actual range, absorbing most of what rotation buys. This project's own GPTQ baseline (08_gptq_apply_plan.py) is also affine, also unrotated, and already ships in production. Rotation stays real and beneficial for the correction step (Step 4's 75% error reduction used it) — the open question was only whether it's worth carrying into the deployed representation, which needs new, unverified machinery to do. Decision: v1 deploys unrotated (Option B), matching how this project's own GPTQ ships.