Every equation here is quoted from the actual paper PDF. Every code excerpt is the actual file, with its real line context. Where something wasn't checked, that's stated plainly instead of assumed.
A neural network, physically, is just a huge pile of numbers called weights. To process text, the model repeatedly does one basic operation: take a list of numbers representing the text so far, and combine them with a group of weights to produce a new list of numbers — its next guess. That combining step has a real name: matrix multiplication (written @ in the real code below). It sounds abstract, but the actual rule is simple and mechanical — here it is on real, tiny numbers:
input: [1, 2]
weights: [[3],
[4]]
result = (1 × 3) + (2 × 4) = 11
That's the entire idea, just done many times over on much bigger lists (thousands of numbers instead of 2) and repeated across many weights at once, which is why it gets written as a grid ("matrix") instead of one line. Every @ anywhere on this page — in the pseudocode, in the real code excerpts, in the paper's own formula — is exactly this same rule, just at real project scale.
Quantizing a weight means writing it down with fewer decimal digits than the model originally used — storing 3.14159265 as just 3.14, or even coarser. This saves real memory and makes the model faster, but each rounded number is now slightly wrong. The real question this whole project deals with: some weights can be rounded aggressively with almost no effect on the model's actual behavior, and others can't — round the wrong one too hard and the model's real output measurably gets worse. Working out which weights are safe to round harder, and by how much, is what everything else on this page exists to do.
A number that answers: "if I nudge this one weight slightly, how much does the model's error change, and in which direction?" Every weight has its own gradient. ∇Wℓ means "the gradient of the loss ℓ with respect to the weight W." This is the real signal the code below is capturing — not training the model further, just measuring how sensitive each weight already is.
The standard technique for computing every weight's gradient at once: start from the model's final error and work backward through each layer using calculus's chain rule. It's the real mechanism this project's code hooks into to capture grad_output below.
A single number meaning "how wrong was the model's prediction" — lower is better. Cross-entropy is the specific, standard way of scoring that for next-token prediction: how confident and correct the model's real guess was for the actual next word.
.T in code)Flip a grid of numbers so its rows become columns and its columns become rows. Needed here purely so the real matrix multiplications line up — two grids can only be multiplied together when their inner dimensions match.
E[...]Just "the average" — here, averaged over every real calibration sequence (below) that was run through the model.
trace(H)Sum of the numbers running diagonally through a grid, top-left to bottom-right. One single number that summarizes part of what the grid "contains" — used below to build the real danger score.
A real number describing how much a matrix stretches space along one particular direction. A Hessian built from real squared gradients (below) can mathematically never have a negative eigenvalue — a sum of squares can't point "backward." Finding one in practice (section 3) is proof of a real numerical bug, not a real property of the data.
A way of combining two smaller grids into one giant one, structured so the giant version never actually has to be built or stored. This is the paper's real efficiency trick: work with H_I and H_O separately (each only as big as one dimension of the real weight) instead of their full, enormous combined product.
A standard, real way of measuring how different two probability distributions are — here, how different the quantized model's real output predictions are from the original, unquantized model's. Lower means the quantized model behaves more like the original.
A small, real, representative sample of text run through the model beforehand, specifically to measure how each weight actually behaves in practice. Used to inform the real rounding decision below — it never changes or trains the model itself.
A standard optimization technique for making many whole-number decisions at once (here: how many bits should each tensor get) so a real total cost is minimized under real constraints (like a total size budget). This project's own solver on top of YAQA — not something the YAQA paper itself does.
Confirmed by matching our accumulation code line-by-line against the paper's own stated algebra, not by pattern-matching or vibes.
The paper's two sketches differ in one structural way: Sketch A treats each token in a sequence as an independent sample (cheaper, biased). Sketch B treats each whole sequence as one sample (more expensive, matches the true Hessian more closely, because attention mixes tokens together so they aren't really independent).
"If (H_I)₀ and (H_O)₀ are both I, then (H_I)₁ = E[(∇Wℓ)ᵀ(∇Wℓ)]/m and (H_O)₁ = E[(∇Wℓ)(∇Wℓ)ᵀ]/n, where the expectation… is taken over entire sequences."
for b in range(x.shape[0]): # b = one whole sequence
xb = x[b]; gb = grad_output[b]
Gb = gb.T @ xb # = ∇_Wℓ for this sequence
st["hin"] += Gb.T @ Gb
st["hout"] += Gb @ Gb.T
x.shape[0] is the batch dimension — each index b is one full calibration sequence, and Gb = gb.T @ xb is exactly ∇_Wℓ for that sequence (gradient-of-loss w.r.t. weight, standard linear-layer backprop formula). Accumulating Gb.T@Gb and Gb@Gb.T across sequences is exactly the paper's own stated Sketch-B reduction — not similar to it, the same formula.
The dense equation quoted above says the same thing as this — nothing extra, nothing skipped. Two things worth naming before the pseudocode: in_dim/out_dim just mean "how many numbers go into, and come out of, this one weight grid" (e.g. a grid that turns a 5120-number input into a 5120-number output has in_dim = out_dim = 5120); and getting each real grad below takes two real steps, not one — run the sequence through the model to get its guess (the forward pass), then compare that guess to the real next word to get a loss (how wrong it was), then work backward through the model using calculus's chain rule to figure out exactly how much this specific weight contributed to that error (the backward pass / backpropagation) — that final per-weight number is the gradient.
# what "(H_I)₁ = E[(∇_Wℓ)ᵀ(∇_Wℓ)] / m" literally means
H_I = zeros(in_dim, in_dim) # starts empty, see note on "identity" below
H_O = zeros(out_dim, out_dim)
for each of the m real calibration sequences:
guess = run_sequence_through_model(sequence) # forward pass
loss = how_wrong_was(guess, real_next_word) # = ℓ
grad = backward_pass(loss, this_weight) # = ∇_Wℓ, one real grid of numbers
H_I += transpose(grad) @ grad # the ᵀ(∇ℓ)(∇ℓ) part of the formula
H_O += grad @ transpose(grad) # the (∇ℓ)(∇ℓ)ᵀ part
H_I = H_I / m # E[...] just means "average" — divide by how many
H_O = H_O / m # sequences you actually summed over
# real code, same page, one line above: "Gb" IS this loop's "grad";
# "hin"/"hout" ARE this loop's H_I/H_O, accumulating the exact same way.
On "identity start": the paper's own notation calls the theoretical starting point (H_I)₀ = I — the mathematical identity matrix (a grid of all zeros except a diagonal line of 1s, the "do nothing" grid, the matrix equivalent of the number 1). That's a statement about what neutral, uninformative belief you start from before any real data arrives, not a term literally added into the sum above. The quoted formula's own right-hand side — E[(∇_Wℓ)ᵀ(∇_Wℓ)]/m — has no identity term in it at all, which is why the pseudocode above just starts from empty and accumulates: for this first round specifically, the paper's own math already works out to exactly that plain average.
H_I = E[GbᵀGb] · H_O = E[Gb Gbᵀ]
Every thin purple slice is one real calibration sequence's Gb. Stacking and averaging them — exactly what E[...]/m means — is what turns many individual, noisy single-sequence measurements into the two stable, confident grids (H_I, H_O) that everything else on this page is built from.
An earlier version of this page raised "further rounds of power iteration" as an open, unresolved question, based only on the paper's prose and a search of this project's own two files. Checked directly against Cornell-RelaxML's actual reference implementation
(hessian_llama/get_hess_llama.py) instead: it has a real --power_iters argument, default 1.
For Sketch A, the code hard-fails if you pass fewer than 2 — that sketch structurally needs multiple refinement passes.
For Sketch B — what our port implements — there is no such requirement, and the reference's own official example command runs it at --power_iters 1: one pass, exactly matching our port. The loop
(for pit in range(args.power_iters)) just re-runs one full calibration sweep per iteration; at 1, that's the same single accumulation pass shown above. Our port isn't missing a second round — a second round was never part of Sketch B to begin with.
Real quotes, and exactly what they do and don't cover.
| Real claim | Source | What it actually covers |
|---|---|---|
| "YAQA… reduces the KL divergence by ≈30% over… [LDLQ/GPTQ]" | Abstract, Conclusion | Rounding quality at a fixed, already-decided bit-width. Table 1 compares YAQA vs. LDLQ row-by-row at the same bits (2/3/4-bit). |
| H̃ = H_O ⊗ H_I (Kronecker-factored sketch) | Section 3.1 | The real structural approximation our own H_I/H_O naming comes directly from. |
| Sketch A vs. Sketch B | Section 3.2.1/3.2.2 | Two real, named ways to estimate that sketch — per-token vs. per-sequence. |
Bit-width allocation across tensors — which tensor gets how many bits under a fixed total budget. No MILP, no Pareto solver, no budget question anywhere in this paper. That's entirely this project's own extension on top of YAQA. The real 30% number is about rounding quality at bits someone else already chose, not about choosing those bits in the first place — a distinction worth keeping precise, since it's easy to accidentally apply that number to the wrong question.
What each part of the real accumulation is actually doing, step by step.
Real sequences of tokens are fed through the whole model, forward and backward, computing the real training-style loss (cross-entropy on next-token prediction).
For each target tensor, a real wrapper (YaqaCatcher + mx.custom_function) captures exactly two real things as the backward pass runs through it: x (what fed into this tensor) and grad_output (how much the final loss cares about this tensor's output) — both real, not approximated.
Gb = gb.T @ xb — the real weight-gradient for one sequence, cast to float32 first (see the real bug story below).
hin += Gb.T @ Gb, hout += Gb @ Gb.T — summed across every real calibration sequence. These become H_I and H_O for that tensor.
effective_rank(H) = trace(H)²/sum(H²), averaged for H_I and H_O — the same real number every Hessian-vs-KL comparison tonight has used. Full trace: the origin page.
Documented directly in the code's own comment: x and grad_output are bfloat16 (this model's native compute type). Forming Gb and accumulating in bfloat16 produced real, impossible negative eigenvalues (H_I min eigenvalue −5.618, H_O −7.534 — a true sum of squared matrices can never be negative). Casting xb/gb to float32 before forming Gb (not after) reduced that to float32's own noise floor (−3e-4 to −8e-4, six to seven orders of magnitude smaller) and matched a gold-reference float64 CPU implementation to 4 decimal places.
Three real diagrams, one per real math step from section 3. Each one rotates slowly so the 3D shape is visible — not decoration, each shape directly represents the actual matrices in the actual code.
Gb = gbᵀ @ xb
Two real sheets of numbers, one per sequence, multiplied together along the sequence-length dimension — every token in that sequence gets folded into one combined sheet. That folding is exactly why this is a per-sequence sample, not a per-token one, and exactly why it matches Sketch B's real formula rather than Sketch A's.
hin += Gbᵀ@Gb · hout += Gb@Gbᵀ
Every real calibration sequence contributes one thin slice. += is literal stacking — more real sequences means a taller, more confident real H_in/H_out, not a different formula. This is what "the real per-sequence sample" from Step 3 turns into once you sum over the whole calibration set.
effective_rank(H) = trace(H)² / sum(H²)
One tall spike among short bars (left) means the real Hessian's energy sits almost entirely in one direction — narrow, fragile, a low score. Many bars of similar height (right) means the energy is spread broadly — forgiving, a high score. The formula is just a precise way of measuring "spiky vs. flat."
Illustrative bar heights, chosen to show the two shapes clearly — not a specific real tensor's measured spectrum.
| Paper's real YAQA-B setup | This project | |
|---|---|---|
| Sequence length | 2K tokens | project-configured, not re-verified on this page |
| Sequence count | 64K sequences (≈128M tokens) | real calibration count set via --n-calibration, real project scale, much smaller |
| Real GPU cost (paper's own 10B-model reference) | ≈30 GPU-hours | not independently re-measured on this page |
This row is intentionally incomplete rather than filled in with a guess — the exact real calibration scale used for any specific build should be read from that build's own real launch command, not assumed from the paper's reference numbers.
YAQA_2505.22988.pdf. Code excerpts read directly from 05_full_model_quantize.py and yaqa_core.py tonight. Every claim on this page is either a direct quote, a direct code excerpt, or explicitly marked as unverified.