← Back to index ← Back to research index
Draft — not live
Improvement Ledger · Start here · not yet promoted to the live site · 2026-09-10

Start here — no matrix, no vector, no LLM knowledge assumed

One picture, held the whole way through this page: a giant sound-mixing board covered in volume knobs. Every idea below — vector, matrix, Hessian, effective rank — is that same picture, just seen a little more precisely each time. No formula appears before the plain idea behind it has already landed.

Read this one straight through, in order. Nothing later needs anything except what's already been said. If you get lost, the break is at the exact sentence where a new word showed up before its idea did — that's a real bug in this page, not something to push past.
Step 1

What an AI language model actually is

Strip away every buzzword: an AI language model is one enormous machine for guessing the next word. You feed it some text, it guesses what word probably comes next, over and over, and that's how it writes.

Inside that machine are billions of tiny adjustable numbers. Think of each one as a physical knob you could turn — turn it one way, the machine's guesses change slightly; turn it the other way, they change differently. During training, all these knobs got tuned, very slowly and carefully, until the machine's guesses became good. This project's model has around 27 billion of these knobs.

The picture, part 1
k1
k2
k3
k4
k5
Imagine a mixing board, except instead of a few dozen knobs for volume and tone, it has 27 billion of them. Each knob's exact position was tuned during training. From here on, every time this page says "knob," it means one of these — the technical word for it is a weight, but "knob" is the same thing.
Step 2

Why we'd want to store the knobs less precisely (and why that's risky)

Each knob's exact position is normally stored using 16 digits of computer precision (16 "bits"). With 27 billion knobs, that's a lot of storage, and the machine has to read all of it every time it guesses a word — that costs time and memory.

The fix this whole project is built around: store each knob's position using fewer digits — say 4 instead of 16. That's 4x less storage and a faster machine. This is called quantization. The catch: with only 4 digits of precision, you can no longer record a knob's exact original position — you have to round it to the nearest position you can record with 4 digits. That rounding is a small error, on every single knob you compress.

Check this landed
For some knobs, that small rounding error changes the machine's guesses so little that nobody would ever notice. For other knobs, the exact same size of rounding error breaks something real — a wrong word, a garbled sentence. Same-sized nudge, wildly different consequences, depending on which knob you nudged. That difference is the entire subject of this project.
Step 3

How you'd actually test which knobs are risky

Go back to the mixing board. You want to know: if I nudge this one knob slightly, does the final sound change a lot or a little? The only real way to find out is to actually play something through the board, nudge the knob, and listen to how much the output changed.

For the AI model, "play something through it" means: run real text through the model, nudge one knob (round it to fewer digits), and measure how different the model's output became. Do that separately for every candidate knob, and you've measured, one at a time, which knobs are risky and which are safe.

Check this landed
This is real, and this project actually does it — it's called an "isolated" measurement, because every other knob is left untouched (full precision) while you test just the one. You'll see this again later as "isolated KL." Nothing new needed yet — just: nudge one knob, measure the change, record it.
Step 4

A list of knobs — what "vector" and "dimension" mean

The mixing board's 27 billion knobs aren't just one giant pile — they're organized into groups. One group might be "everything that decides how much attention this word pays to the word three positions back." A group like that could have, say, 5,120 knobs in it.

Write those 5,120 knob values down as one long ordered list: [0.31, -0.04, 1.2, ..., 0.08], 5,120 numbers long. That ordered list is called a vector. It's not a new idea — it's just the word for "a group of knob values, written down in order." How many numbers are in the list is called its dimension — this list has dimension 5,120.

Step 5

A grid of lists — what "matrix" means

Now stack many of those lists together, one under another, like rows in a spreadsheet. That stack of lists — rows and columns of numbers — is called a matrix. Same idea as the vector, just many vectors stacked into a table instead of one list on its own.

A tiny, made-up 3×3 matrix — just 9 knob values in a grid, for practice
4
1
0
1
3
1
0
1
2
The purple cells running from top-left to bottom-right are called the diagonal. We'll use exactly this 9-number grid again in a few steps — remember it.
Check this landed
Vector = one list of numbers. Matrix = a grid (many lists stacked into rows and columns). Dimension = how long one list is. That's genuinely everything those three words mean — nothing hidden, nothing more advanced coming for these three specifically.
Step 6

Knobs don't act alone — why we need more than "test one at a time"

Step 3's one-at-a-time test has a real gap: it never checks what happens when two knobs get nudged together. Two knobs might amplify each other's damage when both are off — or they might partly cancel each other out. Testing each knob completely alone can never see that.

So instead of just recording "how sensitive is this one knob," we want to record, for every pair of knobs in a group: "how sensitive is knob A, how sensitive is knob B, and how much do they affect each other when nudged together." Written down for every pair in a group of, say, 5,120 knobs, that's a lot of numbers — exactly enough to fill a 5,120×5,120 grid. In other words: a matrix, from Step 5.

Check this landed
That specific matrix — one number for every pair of knobs, capturing "how sensitive, and how do they interact" — is called the Hessian. Not a new kind of object. It's a matrix, exactly as defined in Step 5, just filled in with a specific kind of number: sensitivity-and-interaction, not arbitrary values.
Step 7

H_I and H_O — the two real Hessians this project actually computes

For one group of knobs (one real "tensor" in this project), YAQA actually builds two of these Hessian matrices, not one:

Both are real matrices, built the way Step 6 described, from real text actually run through the real model — not estimated, not guessed.

Real, from disk — this project's own log, for one real tensor
H_I shape=(5120, 5120), H_O shape=(12288, 12288)

Read this exactly the way Step 5 taught: H_I is a grid, 5,120 rows by 5,120 columns — one number for every pair among this group's 5,120 input knobs. H_O is a separate grid, 12,288×12,288, same idea for the output side.

Step 8

Turning a huge grid into one simple number — trace and "spread"

A 5,120×5,120 grid is over 26 million numbers — far too many to look at directly. What we actually want is one simple summary: is this group's risk concentrated in just a few knobs, or spread out evenly across all of them?

A few-knobs-concentrated risk is dangerous — those few knobs are load-bearing, and rounding them badly breaks things. A spread-out risk is safer — no single knob matters enough on its own to cause real damage.

Back to the 3×3 toy grid from Step 5. The trace is simply: add up the diagonal (purple) numbers, ignore everything else.

By hand
trace = 4 + 3 + 2 = 9

Next, the Frobenius norm squared: square every number in the whole grid (not just the diagonal), then add all 9 of those squares up.

By hand
4²+1²+0² + 1²+3²+1² + 0²+1²+2² = 16+1+0+1+9+1+0+1+4 = 33
Step 9

Effective rank — the one number this project actually uses

This project's real formula, combining the two numbers from Step 8:

The formula, only now that both halves are already known
effective_rank = trace² ÷ Frobenius-norm-squared = 9² ÷ 33 = 81 ÷ 33 = 2.45

Why this specific division produces "how spread out is the risk": if every knob in the group carried exactly equal risk, this number comes out equal to the group's real dimension (a 3×3 grid would give exactly 3). If instead almost all the risk sits in one single knob, this number comes out close to 1, no matter how big the grid technically is. It's a sliding scale from "1 knob really matters" to "every knob matters equally" — and it can land anywhere in between, including fractional values like 2.45, because real risk is rarely split among a whole, round number of knobs.

Real, from disk — same tensor as Step 7
real effective rank -- H_I=1.1/5120, H_O=1.8/12288

1.1, out of a possible 5,120. Read exactly as just explained: this group's real input-side risk is concentrated in barely more than one effective knob, out of the 5,120 it technically has. That's about as concentrated — as dangerous — as this measurement can show.

Step 10

From "1.1 out of 5120" to the one number the whole site calls the "danger score"

Two more small, plain steps, both real code, both arithmetic you can redo by hand:

Step A — turn each into a fraction of its own size
hi_frac = 1.1 ÷ 5120 = 0.000215
ho_frac = 1.8 ÷ 12288 = 0.0001465
(Dividing by the group's own size makes a 5,120-wide group and a 12,288-wide group comparable on the same 0-to-1 scale — otherwise a bigger group would look automatically "safer" just for being bigger, which isn't real information.)
Step B — average the two
hess_score = (0.000215 + 0.0001465) ÷ 2 = 0.00018
Check this landed — this is the whole chain, no step skipped
27 billion knobs → grouped into vectors → stacked into matrices → two of those matrices (H_I, H_O) measure real sensitivity-and-interaction for one group → trace and Frobenius norm compress each huge matrix into one "how spread out is the risk" number (effective rank) → dividing by group size and averaging the input/output versions gives one final number per group: hess_score. Low hess_score = this group's real risk is concentrated in almost no knobs at all = dangerous to round carelessly. That's the entire definition — nothing about it was ever more mysterious than the ten steps above.
Hakim Ghelab, VegaLaboratories LTD · improvement ledger, start-here draft · every real number traced to this project's own files, listed inline · not yet promoted to the live site · read 01_ZERO_TO_EXPERT_REDO.html next for isolated vs. cascaded KL and Pearson vs. Spearman, then 02_MTP_YAQA_LOG_WALKTHROUGH.html for a real build log decoded line by line