VegaLaboratories · MLX_OptiQ · Qwen3.8-27B-heretic-ara

YAQA-UMA Documentation

The first known port of YAQA (Cornell-RelaxML) to Apple's MLX, and the real architecture built to make its two-sided Hessian correction fit on a single unified-memory Mac. Everything below is a real file on this machine — verified to exist, not a placeholder.

scripts/yaqa_port/exports/
→ YAQA‑UMA — the brand page
The flagship page for the whole project: live animated hero, founder story, real attribution to Cornell RelaxML's original YAQA work, architecture overview, and live results. Start here if you're new to this project.
brand/YAQA_UMA_Brand.html
→ Timeline Audit — the real chronology, and every stale conclusion found & fixed
Cross-checks every investigation doc in this project against what actually happened afterward. Found 4 docs whose real conclusion was later superseded and never updated, plus 2 stale project ledgers — all fixed, this page shows exactly where and why.
exports/TIMELINE_AUDIT.html
→ Live Mission Control — current Hessian coverage-probe run
Real-time, auto-refreshing status of the real 487-tensor godmode multi-bit probe closing the Hessian coverage gap (2026-09-15). Local only, tied to this specific run — will go stale once it completes or the resume folder is cleaned up.
LIVE · .mtplx/models/...HESSIAN-PROBE.yaqa_resume/dashboard.html

Start here

→ README
What YAQA is, why it's different from GPTQ, current real status of every step (1-5), and install/requirements.
→ Complete Running Guide
Every real command, every flag, exactly what each one does. Start/resume/monitor a full run, including the live mission-control dashboard.
→ Command Reference
Quick flag lookup — the condensed version of the running guide for when you already know the pipeline.

The engineering record

→ Port Ledger
The full, honest log: every real bug found and fixed, every command actually run, the batching architecture, and the live RCA of the one tensor YAQA made worse than naive — root cause, options considered, decision made.
→ Changelog
The short, chronological version of every real change — for when the full ledger is more detail than you need.
→ Batching + Resume Architecture
Spatial diagram of how 363 tensors get processed on 128GB of unified memory without ever holding all their Hessians at once, plus the real lm_head finding: this port compresses it, the original reference doesn't. Updated 2026-09-25 with a section on the adaptive controller.
00_DOCS/HTML/
→ Adaptive Batching: How the Run Protects Itself
What the run does and why it is cut into batches (326 GB of tables against a 115 GB GPU limit), the night the GPU reset itself, what macOS's own crash reports show (a firmware lockup, not size or memory) and what is still unknown, then the controller that now senses the machine, sizes each batch, classifies failures, retries at most twice and notifies. Real charts, a labelled flowchart and two interactive tools.
research_hadamard_blowup/exports/ · 2026-09-25
→ Timeline Audit
A meta-audit of this project's own documentation integrity, not its engineering — checks the real site's docs against each other for stale claims and contradictions, same spirit as this changelog's own 2026-09-16 catch-up.
exports/ · 2026-09-07

Brand & outreach

→ Brand Page
The public-facing identity: name, positioning, timeline, and impact — every factual claim sourced from this project's own verified work.
brand/ · 2026-09-07
→ Launch Kit
Local-only draft kit for X, LinkedIn, Reddit, Hacker News, a Hugging Face model card, and an Apple outreach email. Nothing has been posted anywhere — review before using any of it.
outreach/ · 2026-09-06

2026-09-06: the BF16 curvature bug

→ The Hessian, Fully Explained — Zero to Expert
Reference companion to everything below: what a Hessian actually is, how H_I/H_O get built from real forward/backward passes, why eigenvalues must stay positive, and how this project caught real corruption by checking that.
research_hadamard_blowup/
→ Hadamard Blowup — Root-Cause Research
The full Define/Hypothesize/Research/Test/Measure/Validate investigation into the one catastrophic tensor — rotation ruled out, calibration size ruled out, before the real cause was found.
research_hadamard_blowup/
→ The BF16 Curvature Bug — Plain English
Production was computing YAQA's curvature in 16-bit, not 32-bit as assumed. Every term defined, a before/after flow diagram, the exact one-line fix, and the exact reprocessing command.
research_hadamard_blowup/
→ Hessian Precision — the Math, Root Cause, and Fix
The from-first-principles technical companion to the plain-English writeup above: what a Hessian is, why it must be positive-semidefinite, why BF16 broke that, and the exact corrected production architecture.
research_hadamard_blowup/
→ GPTQ lm_head — Safety Gate + Precision Fix
lm_head's separate GPTQ correction had no safety net and the same precision issue. Exact before/after code, verified against the original GPTQ reference implementation, real dry-run results.
research_hadamard_blowup/
→ The Power of Adapting a Technology
Porting YAQA's adaptive-damping concept into GPTQ's own algorithm — real results, and a mechanical proof (not just an assurance) that the improvement is genuine, not a metric being gamed.
research_hadamard_blowup/

2026-09-07: the slow-inference investigation

→ The "Missing Affine Mode" Bug — Full Code Review
First of two real speed regressions found this day: config.json never declared mode: "affine", even though every real quantize call always used it. A metadata-only bug — exhaustive evidence that the packed weight bytes were never wrong, just undescribed.
research_hadamard_blowup/
→ Scales/Biases in F32, and the embed_tokens Mismatch — Full Root-Cause Analysis
Second, larger regression: every quantized layer's scales/biases were saved at float32 instead of bfloat16 (double the per-group metadata MLX reads on every forward pass), plus embed_tokens was quantized differently than the reference. Both fixed, rebuilt, and independently re-verified against the live model — full tensor forensics, weight-provenance chain-of-custody check, and real post-fix mtplx tune numbers.
research_hadamard_blowup/ · not the same bug as the 2026-09-06 curvature section above
→ The MTP Sidecar Mismatch — Plan to Fix It
Third real gap found the same day: the MTP speculative-decode draft head was quantized naively in both the YAQA build and the reference baseline (same shared code path) — but the reference's own trunk is also naive (confirmed from its real build manifest: 35.8s total build time, no correction code), while YAQA's trunk is genuinely corrected. Naive-trunk/naive-sidecar vs. corrected-trunk/naive-sidecar was never an apples-to-apples comparison. Plan: extend real YAQA correction to the MTP sidecar too, integrated into the main build script, not a new one-off tool.
research_hadamard_blowup/ · implemented and validated with a real smoke test against the actual model
→ MTP Sidecar — Every Real Tensor, What It Is, What Changed
Real data read directly from each model's actual safetensors header and config.json — reference, the naive build, and the YAQA-corrected build, tensor by tensor. No estimates.
research_hadamard_blowup/

2026-09-08/10: does the free Hessian signal predict real risk? (Hessian investigation, part 1 of 3)

→ The Free Signal Hiding in Your Hessians — Live
Auto-refreshed every 30s from the real, currently-running cascaded measurement. Every round this run has completed or is still working through, shown separately — round 1's real, complete picture is never silently replaced by a later partial round. Includes a real cinematic scrubber replaying each round's tensors in their true measurement order.
research_hadamard_blowup/ · LIVE · will go stale once the run finishes
→ The Full Investigation, Reasoned From Zero
Every term defined before use, every derivation shown step by step, every real number traced to a real file — including two real mistakes made and corrected mid-analysis, left visible rather than edited away. Answers: is there a real link between the free Hessian score and real risk, does the early-QKV floor over-protect anything, and is round 1 vs. round 2 a yo-yo or a real, one-directional shift.
research_hadamard_blowup/exports/
→ What "Effective Rank 1.6 out of 5120" Actually Looks Like
The danger-score formula derived by hand from first principles, then rendered as real, rotatable 3D shapes — five real tensors from this model's own Hessian log, each shape's axis ratios exactly matching a hand-verified calculation.
research_hadamard_blowup/exports/ · real WebGL 3D, drag to rotate

2026-09-10/15: the improvement ledger, 00–10 — every real investigation, start to finish (Hessian investigation, part 2 of 3)

→ Absolute Zero, Start Here
One picture held the whole way through: a sound-mixing board covered in volume knobs. Every idea — vector, matrix, Hessian, effective rank — is that same picture seen a little more precisely each time. No formula appears before the plain idea behind it has already landed.
IMPROVEMENT_LEDGER/
→ Zero to Expert, Redone
Every concept worked on one real tensor, start to finish, with this project's own real numbers — instead of defining terms up front and jumping straight to results.
IMPROVEMENT_LEDGER/
→ MTP-YAQA Log Walkthrough
The exact real terminal output from a real run_full_yaqa.sh --correct-mtp run, broken into its real sections, each one explained in plain language directly underneath.
IMPROVEMENT_LEDGER/
→ The Brainstorm — Multi-Bit YAQA Sensitivity
Every number from a real file on disk or the real, then-running code. Read top to bottom — each answer only uses facts established in an earlier one.
IMPROVEMENT_LEDGER/
→ The lm_head Q4 Incident
A step-by-step, code-verified trace of the mechanism, the exact real blast radius (all 497 tensors checked, not estimated), and a precise verdict on what could be salvaged for free versus what genuinely needed re-measurement.
IMPROVEMENT_LEDGER/
→ lm_head Fix — Exact Next Steps
The real, default-on hard floor that came out of the incident above, and the exact sequence to actually use it.
IMPROVEMENT_LEDGER/
→ Trunk, Hidden State, Sidecar
Makes a real mechanism actually visible instead of just described, so "the trunk's hidden state changed" stops being a phrase and becomes something you can see happen.
IMPROVEMENT_LEDGER/
→ The Hessian Signal — Four Real Questions in Order
Why isolated KL can't be fully trusted, why cascading it is a real fix with a real cost, why the flat 100× floors were a defensible patch, and the free per-tensor signal buried inside YAQA's own Hessian that doesn't share the cascade's snowball problem. Includes a real, rotatable 3D chart of measured danger score by tensor role.
IMPROVEMENT_LEDGER/ · updated 2026-09-15
→ Hessian Hybrid Solver — Full Technical Debrief
Every flag, every command, every real result, in order — from the first soft-weight experiments through the real feasibility ceiling, the per-tensor hard floor, and the lexicographic Hessian-primary solve. The canonical reference for exactly what's shipped, what's opt-in, and what's still open.
IMPROVEMENT_LEDGER/ · updated 2026-09-15 · canonical solver-strategy reference
→ The First Real Benchmark Results
Real, complete MMLU/GSM8K/IFEval/BFCL/HumanEval scores for the stratified-calibration, cascade-corrected, YAQA-rounding-protected build — deterministic (temperature=0), against five other real models including a naive flat-bit control. Highest 5-test mean of the group, never worst on any single metric.
IMPROVEMENT_LEDGER/ · new 2026-09-15
→ The Pure Sensitivity Hypothesis
States precisely, with real solver line citations, what happens once the godmode multi-bit sweep reaches 100% coverage and the hybrid checkpoint from §8 has nothing left to fall back to — and names three questions that are genuinely still open, not yet answered by any page on this site.
IMPROVEMENT_LEDGER/ · new 2026-09-17 · open hypothesis, not a confirmed result

2026-09-14/16: MILP-aware Hessian — from coverage gap to hybrid checkpoint (Hessian investigation, part 3 of 3)

→ Hessian vs. KL — Deep-Dive Companion
Self-contained deep dive into the whole investigation: what the Hessian danger score is, how it was verified against isolated and cascaded KL, where the correlation holds and where it flips sign, and how it led to the hybrid-checkpoint proposal below. For the current, continuing story, start at Improvement Ledger §07 instead.
research_hadamard_blowup/exports/ · 2026-09-16
→ Hessian Score, Origin Trace
Where the Hessian danger score actually comes from, traced line-by-line from the real Hessian construction through trace → Frobenius norm → effective rank in/out → final score, with a hand-verified 2×2 matrix worked example and the exact print statement in safety_gate() that fires on every real correction pass.
research_hadamard_blowup/exports/ · 2026-09-16
→ How the Solver Turns a Hessian Score into a Bit-Width
Mechanism reference for --hessian-primary versus the floor: raw score → percentile rank → danger weight, the phase-1 objective and budget worked step by step, why tiny tensors win 16-bit (danger per million parameters), what happens to unscored tensors, and why a second phase exists. Every number re-checked by script; an independent re-solve reproduces plan C2 exactly.
research_hadamard_blowup/exports/ · 2026-09-19
→ Hessian vs. Isolated KL, Tensor Explorer
Interactive, live-refreshed explorer over every real profiled tensor — Hessian danger score plotted against isolated KL divergence per tensor, so the real correlation (and its exceptions) can be inspected tensor by tensor rather than taken on faith from a single summary number. Now flags a real, confirmed pattern: every sharp disagreement tensor is a late layer (48–63).
research_hadamard_blowup/exports/ · 2026-09-16
→ YAQA Paper vs. Code, Verified Line by Line
Cornell-RelaxML's real YAQA paper (arXiv:2505.22988) checked directly against this project's actual rounding code — confirms the codebase implements Sketch B exactly, with real 3D diagrams of the Kronecker-factored Hessian, and a correction of the difference between "YAQA's rounding beats LDLQ by ≈30%" (paper-confirmed) and "using Hessian to pick bit allocation" (never tested by the paper).
research_hadamard_blowup/exports/ · 2026-09-16
→ Hessian Hybrid Checkpoint, Design
The V4 hybrid design: merge godmode's real per-candidate-bit rounding error into the same sensitivities[bit] field the MILP solver already reads, with cascaded KL as fallback for uncovered tensors — so the existing, completely unmodified solver consumes richer, real per-candidate cost data. Built via the consolidated 08_extract_hessian_scores.py hybrid subcommand.
research_hadamard_blowup/exports/ · 2026-09-16
→ Closing the Hessian Coverage Gap
Plain-language, every-command guide to re-measuring the 136 real tensors this project's own protection rules kept at full precision (and so never got a real Hessian score) — the exact real godmode command sequence, step by step, nothing assumed known.
research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/
→ Making the Solver Listen to Curvature, Not Just KL
Real, tested proof that a soft Hessian weight saturates — even 500× stronger changes nothing for the most dangerous tensors — and the two new opt-in solver mechanisms that fix it: a real per-tensor hard floor, and a lexicographic Hessian-first solve. Full real comparison against production, plus an independent code review. Updated 2026-09-18: real, confirmed finding — --pareto none is required whenever a Hessian file drives allocation, or it can silently produce genuine MILP infeasibility.
research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/ · solver flags documented in RUNNING_GUIDE.md §11.8
→ Plan Comparison, Every Tensor Across 7 Plans
All 497 real tensors and all 7 real solver plans side by side: production, plans A/B/C (original settings) and A2/B2/C2 (corrected: --pareto none, --hessian-weight-strength 0). Sortable, searchable, filterable to the tensors where plans disagree, with each plan’s floors, the Hessian score and the isolated KL kept as reference.
research_hadamard_blowup/exports/ · 2026-09-18
→ Hessian-Weight Strength Sweep, Raw Data
The real 2026-09-14 sweep of the soft Hessian weight: correlation and stability tables, plus the logs behind them. Kept as raw supporting data; the current recommendation lives in ledger §08 and the MILP-aware guide.
research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/exports/ · 2026-09-14

Why each method matters

→ YAQA for MLX — Why It Matters
How Sketch B's two-sided Hessian sketch actually works, verified equation-by-equation against the real paper. Written before the port started — the original feasibility case.
scripts/yaqa_port/exports/
→ GPTQ for MLX — Why It Matters
Why compensated rounding beats naive quantization, and how this project's per-tensor, plan-aware integration goes further than Apple's own shipped GPTQ.
scripts/yaqa_port/exports/
→ GPTQ MLX Integration
The technical reference: exact commands, every flag, and the real fix that extended Apple's GPTQ packer from bits {2,4,8} to {2,3,4,5,6,8}.
research/quantization_methods_2026/exports/
→ YAQA Pipeline Comparison
Side-by-side: the real reference's trellis-VQ codebook pipeline vs. this port's affine-quantized, unrotated-by-default pipeline.
scripts/yaqa_port/exports/

Real measured results

→ Benchmark Master Recap
The living source of truth for real quality/speed numbers across every quantized variant of this model — pulled from real log files, with honest statistical-significance checks, not eyeballed deltas.
scripts/yaqa_port/exports/
→ Benchmark Run Guide
Companion to the recap above: how to re-run any existing repro_check_*.sh script, the exact copy-paste template to design a new one for a new model, real tool gotchas, and the mtplx tune speed commands.
scripts/yaqa_port/exports/