← Back to index
Paper vs. code, verified line by line · 2026-09-16

The real YAQA paper, next to our real code, with no unverified claims

Every equation here is quoted from the actual paper PDF. Every code excerpt is the actual file, with its real line context. Where something wasn't checked, that's stated plainly instead of assumed.

Start here if you're new to any of this — the foundation everything else is built on (click to collapse)

A neural network, physically, is just a huge pile of numbers called weights. To process text, the model repeatedly does one basic operation: take a list of numbers representing the text so far, and combine them with a group of weights to produce a new list of numbers — its next guess. That combining step has a real name: matrix multiplication (written @ in the real code below). It sounds abstract, but the actual rule is simple and mechanical — here it is on real, tiny numbers:

input:   [1, 2]
weights: [[3],
          [4]]

result = (1 × 3) + (2 × 4) = 11

That's the entire idea, just done many times over on much bigger lists (thousands of numbers instead of 2) and repeated across many weights at once, which is why it gets written as a grid ("matrix") instead of one line. Every @ anywhere on this page — in the pseudocode, in the real code excerpts, in the paper's own formula — is exactly this same rule, just at real project scale.

Quantizing a weight means writing it down with fewer decimal digits than the model originally used — storing 3.14159265 as just 3.14, or even coarser. This saves real memory and makes the model faster, but each rounded number is now slightly wrong. The real question this whole project deals with: some weights can be rounded aggressively with almost no effect on the model's actual behavior, and others can't — round the wrong one too hard and the model's real output measurably gets worse. Working out which weights are safe to round harder, and by how much, is what everything else on this page exists to do.

The more specialized terms below build on that (click to collapse)
Gradient (the ∇ symbol)

A number that answers: "if I nudge this one weight slightly, how much does the model's error change, and in which direction?" Every weight has its own gradient. ∇Wℓ means "the gradient of the loss ℓ with respect to the weight W." This is the real signal the code below is capturing — not training the model further, just measuring how sensitive each weight already is.

Backward pass / backpropagation

The standard technique for computing every weight's gradient at once: start from the model's final error and work backward through each layer using calculus's chain rule. It's the real mechanism this project's code hooks into to capture grad_output below.

Loss (ℓ) / cross-entropy

A single number meaning "how wrong was the model's prediction" — lower is better. Cross-entropy is the specific, standard way of scoring that for next-token prediction: how confident and correct the model's real guess was for the actual next word.

Transpose (the ᵀ symbol, or .T in code)

Flip a grid of numbers so its rows become columns and its columns become rows. Needed here purely so the real matrix multiplications line up — two grids can only be multiplied together when their inner dimensions match.

Expectation, E[...]

Just "the average" — here, averaged over every real calibration sequence (below) that was run through the model.

Trace, trace(H)

Sum of the numbers running diagonally through a grid, top-left to bottom-right. One single number that summarizes part of what the grid "contains" — used below to build the real danger score.

Eigenvalue

A real number describing how much a matrix stretches space along one particular direction. A Hessian built from real squared gradients (below) can mathematically never have a negative eigenvalue — a sum of squares can't point "backward." Finding one in practice (section 3) is proof of a real numerical bug, not a real property of the data.

Kronecker product (the ⊗ symbol)

A way of combining two smaller grids into one giant one, structured so the giant version never actually has to be built or stored. This is the paper's real efficiency trick: work with H_I and H_O separately (each only as big as one dimension of the real weight) instead of their full, enormous combined product.

KL divergence

A standard, real way of measuring how different two probability distributions are — here, how different the quantized model's real output predictions are from the original, unquantized model's. Lower means the quantized model behaves more like the original.

Calibration (data / sequences)

A small, real, representative sample of text run through the model beforehand, specifically to measure how each weight actually behaves in practice. Used to inform the real rounding decision below — it never changes or trains the model itself.

MILP (Mixed-Integer Linear Program)

A standard optimization technique for making many whole-number decisions at once (here: how many bits should each tensor get) so a real total cost is minimized under real constraints (like a total size budget). This project's own solver on top of YAQA — not something the YAQA paper itself does.

1 · The verdict: Sketch B, one confirmed round

Directly verified, not inferred

Our code implements Sketch B's real formula — specifically, its exact first-round reduction from an identity start.

Confirmed by matching our accumulation code line-by-line against the paper's own stated algebra, not by pattern-matching or vibes.

Why it's Sketch B and not Sketch A

The paper's two sketches differ in one structural way: Sketch A treats each token in a sequence as an independent sample (cheaper, biased). Sketch B treats each whole sequence as one sample (more expensive, matches the true Hessian more closely, because attention mixes tokens together so they aren't really independent).

The paper — Section 3.2.2, real quote

"If (H_I)₀ and (H_O)₀ are both I, then (H_I)₁ = E[(∇Wℓ)ᵀ(∇Wℓ)]/m and (H_O)₁ = E[(∇Wℓ)(∇Wℓ)ᵀ]/n, where the expectation… is taken over entire sequences."

Our code — 05_full_model_quantize.py, real excerpt
for b in range(x.shape[0]):   # b = one whole sequence
    xb = x[b]; gb = grad_output[b]
    Gb = gb.T @ xb             # = ∇_Wℓ for this sequence
    st["hin"]  += Gb.T @ Gb
    st["hout"] += Gb @ Gb.T

x.shape[0] is the batch dimension — each index b is one full calibration sequence, and Gb = gb.T @ xb is exactly ∇_Wℓ for that sequence (gradient-of-loss w.r.t. weight, standard linear-layer backprop formula). Accumulating Gb.T@Gb and Gb@Gb.T across sequences is exactly the paper's own stated Sketch-B reduction — not similar to it, the same formula.

The paper's formula, translated into real pseudocode, symbol by symbol

The dense equation quoted above says the same thing as this — nothing extra, nothing skipped. Two things worth naming before the pseudocode: in_dim/out_dim just mean "how many numbers go into, and come out of, this one weight grid" (e.g. a grid that turns a 5120-number input into a 5120-number output has in_dim = out_dim = 5120); and getting each real grad below takes two real steps, not one — run the sequence through the model to get its guess (the forward pass), then compare that guess to the real next word to get a loss (how wrong it was), then work backward through the model using calculus's chain rule to figure out exactly how much this specific weight contributed to that error (the backward pass / backpropagation) — that final per-weight number is the gradient.

# what "(H_I)₁ = E[(∇_Wℓ)ᵀ(∇_Wℓ)] / m" literally means

H_I = zeros(in_dim, in_dim)     # starts empty, see note on "identity" below
H_O = zeros(out_dim, out_dim)

for each of the m real calibration sequences:
    guess = run_sequence_through_model(sequence)         # forward pass
    loss  = how_wrong_was(guess, real_next_word)          # = ℓ
    grad  = backward_pass(loss, this_weight)               # = ∇_Wℓ, one real grid of numbers
    H_I += transpose(grad) @ grad     # the ᵀ(∇ℓ)(∇ℓ) part of the formula
    H_O += grad @ transpose(grad)     # the (∇ℓ)(∇ℓ)ᵀ part

H_I = H_I / m     # E[...] just means "average" — divide by how many
H_O = H_O / m     # sequences you actually summed over

# real code, same page, one line above: "Gb" IS this loop's "grad";
# "hin"/"hout" ARE this loop's H_I/H_O, accumulating the exact same way.

On "identity start": the paper's own notation calls the theoretical starting point (H_I)₀ = I — the mathematical identity matrix (a grid of all zeros except a diagonal line of 1s, the "do nothing" grid, the matrix equivalent of the number 1). That's a statement about what neutral, uninformative belief you start from before any real data arrives, not a term literally added into the sum above. The quoted formula's own right-hand side — E[(∇_Wℓ)ᵀ(∇_Wℓ)]/m — has no identity term in it at all, which is why the pseudocode above just starts from empty and accumulates: for this first round specifically, the paper's own math already works out to exactly that plain average.

The same formula, as a real shape — many sequences, averaged into two grids

H_I = E[GbᵀGb]  ·  H_O = E[Gb Gbᵀ]

one real Gb, one real calibration sequence H_I, after averaging over all m sequences H_O, after averaging over all m sequences

Every thin purple slice is one real calibration sequence's Gb. Stacking and averaging them — exactly what E[...]/m means — is what turns many individual, noisy single-sequence measurements into the two stable, confident grids (H_I, H_O) that everything else on this page is built from.

Corrected 2026-09-17, checked against the real reference code, not just the paper

Single-pass is the correct, reference-matching behavior for Sketch B — this was never a gap.

An earlier version of this page raised "further rounds of power iteration" as an open, unresolved question, based only on the paper's prose and a search of this project's own two files. Checked directly against Cornell-RelaxML's actual reference implementation (hessian_llama/get_hess_llama.py) instead: it has a real --power_iters argument, default 1. For Sketch A, the code hard-fails if you pass fewer than 2 — that sketch structurally needs multiple refinement passes. For Sketch B — what our port implements — there is no such requirement, and the reference's own official example command runs it at --power_iters 1: one pass, exactly matching our port. The loop (for pit in range(args.power_iters)) just re-runs one full calibration sweep per iteration; at 1, that's the same single accumulation pass shown above. Our port isn't missing a second round — a second round was never part of Sketch B to begin with.

2 · What the paper is actually claiming, precisely

Real quotes, and exactly what they do and don't cover.

Real claimSourceWhat it actually covers
"YAQA… reduces the KL divergence by ≈30% over… [LDLQ/GPTQ]"Abstract, ConclusionRounding quality at a fixed, already-decided bit-width. Table 1 compares YAQA vs. LDLQ row-by-row at the same bits (2/3/4-bit).
H̃ = H_O ⊗ H_I (Kronecker-factored sketch)Section 3.1The real structural approximation our own H_I/H_O naming comes directly from.
Sketch A vs. Sketch BSection 3.2.1/3.2.2Two real, named ways to estimate that sketch — per-token vs. per-sequence.
What the paper does not cover

Bit-width allocation across tensors — which tensor gets how many bits under a fixed total budget. No MILP, no Pareto solver, no budget question anywhere in this paper. That's entirely this project's own extension on top of YAQA. The real 30% number is about rounding quality at bits someone else already chose, not about choosing those bits in the first place — a distinction worth keeping precise, since it's easy to accidentally apply that number to the wrong question.

3 · The real code, walked through in plain language

What each part of the real accumulation is actually doing, step by step.

1

Run real calibration text through the model

Real sequences of tokens are fed through the whole model, forward and backward, computing the real training-style loss (cross-entropy on next-token prediction).

2

Intercept one tensor's own input and gradient

For each target tensor, a real wrapper (YaqaCatcher + mx.custom_function) captures exactly two real things as the backward pass runs through it: x (what fed into this tensor) and grad_output (how much the final loss cares about this tensor's output) — both real, not approximated.

3

Form the real per-sequence gradient, Gb

Gb = gb.T @ xb — the real weight-gradient for one sequence, cast to float32 first (see the real bug story below).

4

Accumulate two real matrices

hin += Gb.T @ Gb, hout += Gb @ Gb.T — summed across every real calibration sequence. These become H_I and H_O for that tensor.

5

Turn H_I/H_O into the danger score

effective_rank(H) = trace(H)²/sum(H²), averaged for H_I and H_O — the same real number every Hessian-vs-KL comparison tonight has used. Full trace: the origin page.

A real bug this exact code already found and fixed — a concrete illustration of why the float32 cast matters

Documented directly in the code's own comment: x and grad_output are bfloat16 (this model's native compute type). Forming Gb and accumulating in bfloat16 produced real, impossible negative eigenvalues (H_I min eigenvalue −5.618, H_O −7.534 — a true sum of squared matrices can never be negative). Casting xb/gb to float32 before forming Gb (not after) reduced that to float32's own noise floor (−3e-4 to −8e-4, six to seven orders of magnitude smaller) and matched a gold-reference float64 CPU implementation to 4 decimal places.

3b · The same three steps, in 3D

Three real diagrams, one per real math step from section 3. Each one rotates slowly so the 3D shape is visible — not decoration, each shape directly represents the actual matrices in the actual code.

Step 3 — forming Gb, one sequence's real weight-gradient

Gb = gbᵀ @ xb

xb — this sequence's real input, [seq_len × in_dim] gb — this sequence's real gradient, [seq_len × out_dim] Gb — the real result, [out_dim × in_dim]

Two real sheets of numbers, one per sequence, multiplied together along the sequence-length dimension — every token in that sequence gets folded into one combined sheet. That folding is exactly why this is a per-sequence sample, not a per-token one, and exactly why it matches Sketch B's real formula rather than Sketch A's.

Step 4 — accumulating Gb across every real calibration sequence

hin += Gbᵀ@Gb  ·  hout += Gb@Gbᵀ

one real Gb, one real sequence H_in, growing H_out, growing

Every real calibration sequence contributes one thin slice. += is literal stacking — more real sequences means a taller, more confident real H_in/H_out, not a different formula. This is what "the real per-sequence sample" from Step 3 turns into once you sum over the whole calibration set.

Step 5 — what effective_rank actually measures, as a shape

effective_rank(H) = trace(H)² / sum(H²)

concentrated — low score, dangerous spread out — high score, safe

One tall spike among short bars (left) means the real Hessian's energy sits almost entirely in one direction — narrow, fragile, a low score. Many bars of similar height (right) means the energy is spread broadly — forgiving, a high score. The formula is just a precise way of measuring "spiky vs. flat."

Illustrative bar heights, chosen to show the two shapes clearly — not a specific real tensor's measured spectrum.

4 · Real scale, this project vs. the paper's own experiments

Paper's real YAQA-B setupThis project
Sequence length2K tokensproject-configured, not re-verified on this page
Sequence count64K sequences (≈128M tokens)real calibration count set via --n-calibration, real project scale, much smaller
Real GPU cost (paper's own 10B-model reference)≈30 GPU-hoursnot independently re-measured on this page

This row is intentionally incomplete rather than filled in with a guess — the exact real calibration scale used for any specific build should be read from that build's own real launch command, not assumed from the paper's reference numbers.

Author: Hakim Ghelab, VegaLaboratories LTD · Paper equations quoted directly from YAQA_2505.22988.pdf. Code excerpts read directly from 05_full_model_quantize.py and yaqa_core.py tonight. Every claim on this page is either a direct quote, a direct code excerpt, or explicitly marked as unverified.