One picture, held the whole way through this page: a giant sound-mixing board covered in volume knobs. Every idea below — vector, matrix, Hessian, effective rank — is that same picture, just seen a little more precisely each time. No formula appears before the plain idea behind it has already landed.
Strip away every buzzword: an AI language model is one enormous machine for guessing the next word. You feed it some text, it guesses what word probably comes next, over and over, and that's how it writes.
Inside that machine are billions of tiny adjustable numbers. Think of each one as a physical knob you could turn — turn it one way, the machine's guesses change slightly; turn it the other way, they change differently. During training, all these knobs got tuned, very slowly and carefully, until the machine's guesses became good. This project's model has around 27 billion of these knobs.
Each knob's exact position is normally stored using 16 digits of computer precision (16 "bits"). With 27 billion knobs, that's a lot of storage, and the machine has to read all of it every time it guesses a word — that costs time and memory.
The fix this whole project is built around: store each knob's position using fewer digits — say 4 instead of 16. That's 4x less storage and a faster machine. This is called quantization. The catch: with only 4 digits of precision, you can no longer record a knob's exact original position — you have to round it to the nearest position you can record with 4 digits. That rounding is a small error, on every single knob you compress.
Go back to the mixing board. You want to know: if I nudge this one knob slightly, does the final sound change a lot or a little? The only real way to find out is to actually play something through the board, nudge the knob, and listen to how much the output changed.
For the AI model, "play something through it" means: run real text through the model, nudge one knob (round it to fewer digits), and measure how different the model's output became. Do that separately for every candidate knob, and you've measured, one at a time, which knobs are risky and which are safe.
The mixing board's 27 billion knobs aren't just one giant pile — they're organized into groups. One group might be "everything that decides how much attention this word pays to the word three positions back." A group like that could have, say, 5,120 knobs in it.
Write those 5,120 knob values down as one long ordered list: [0.31, -0.04, 1.2, ..., 0.08], 5,120 numbers long. That ordered list is called a vector. It's not a new idea — it's just the word for "a group of knob values, written down in order." How many numbers are in the list is called its dimension — this list has dimension 5,120.
Now stack many of those lists together, one under another, like rows in a spreadsheet. That stack of lists — rows and columns of numbers — is called a matrix. Same idea as the vector, just many vectors stacked into a table instead of one list on its own.
Step 3's one-at-a-time test has a real gap: it never checks what happens when two knobs get nudged together. Two knobs might amplify each other's damage when both are off — or they might partly cancel each other out. Testing each knob completely alone can never see that.
So instead of just recording "how sensitive is this one knob," we want to record, for every pair of knobs in a group: "how sensitive is knob A, how sensitive is knob B, and how much do they affect each other when nudged together." Written down for every pair in a group of, say, 5,120 knobs, that's a lot of numbers — exactly enough to fill a 5,120×5,120 grid. In other words: a matrix, from Step 5.
For one group of knobs (one real "tensor" in this project), YAQA actually builds two of these Hessian matrices, not one:
Both are real matrices, built the way Step 6 described, from real text actually run through the real model — not estimated, not guessed.
H_I shape=(5120, 5120), H_O shape=(12288, 12288)
Read this exactly the way Step 5 taught: H_I is a grid, 5,120 rows by 5,120 columns — one number for every pair among this group's 5,120 input knobs. H_O is a separate grid, 12,288×12,288, same idea for the output side.
A 5,120×5,120 grid is over 26 million numbers — far too many to look at directly. What we actually want is one simple summary: is this group's risk concentrated in just a few knobs, or spread out evenly across all of them?
A few-knobs-concentrated risk is dangerous — those few knobs are load-bearing, and rounding them badly breaks things. A spread-out risk is safer — no single knob matters enough on its own to cause real damage.
Back to the 3×3 toy grid from Step 5. The trace is simply: add up the diagonal (purple) numbers, ignore everything else.
Next, the Frobenius norm squared: square every number in the whole grid (not just the diagonal), then add all 9 of those squares up.
This project's real formula, combining the two numbers from Step 8:
Why this specific division produces "how spread out is the risk": if every knob in the group carried exactly equal risk, this number comes out equal to the group's real dimension (a 3×3 grid would give exactly 3). If instead almost all the risk sits in one single knob, this number comes out close to 1, no matter how big the grid technically is. It's a sliding scale from "1 knob really matters" to "every knob matters equally" — and it can land anywhere in between, including fractional values like 2.45, because real risk is rarely split among a whole, round number of knobs.
real effective rank -- H_I=1.1/5120, H_O=1.8/12288
1.1, out of a possible 5,120. Read exactly as just explained: this group's real input-side risk is concentrated in barely more than one effective knob, out of the 5,120 it technically has. That's about as concentrated — as dangerous — as this measurement can show.
Two more small, plain steps, both real code, both arithmetic you can redo by hand: