← Back to index ← Back to research index
Hessian Danger Score · Geometric Derivation · 2026-09-10

What "effective rank 1.6 out of 5120" actually looks like

The danger score plotted throughout this project's charts is a single number derived from a real matrix. This page derives, by hand, exactly what that number means geometrically, then renders it — in real, rotatable 3D — using five real tensors from this model's own Hessian log.

Auditability statement, read first: the exact full eigenvalue spectrum of any tensor's Hessian was never saved to disk — only two derived numbers were logged (trace(H) and ||H||_F, combined into the "effective rank" you see below). The 3D ellipsoids on this page are therefore not a picture of the literal unknown eigenvectors. They are a mathematically exact reconstruction: a geometric-decay spectrum, solved so its own effective rank reproduces the real measured number precisely, at the tensor's real dimension. Every step of that solve is shown below, in full, so it can be checked by hand.

Added 2026-09-19: this page is about what the raw score measures. How the solver turns it into a rank, a danger weight and a bit-width is explained in How the Solver Turns a Hessian Score into a Bit-Width.

1. What "effective rank" is measuring

Direct from this project's own source, yaqa_core.py:328 — not paraphrased.

def effective_rank(H):
    return trace(H)**2 / frobenius_norm(H)**2
1

trace(H) — sum the diagonal

Add up the numbers running down the matrix's diagonal. Toy example, a 3×3 case with diagonal values 5, 3, 1: trace = 5+3+1 = 9. For the kind of matrix used here, these diagonal values are (or closely track) the matrix's real eigenvalues — the actual "how much curvature exists in each independent direction" numbers.

2

||H||_F² — square every entry, then sum

Same toy example: 5² + 3² + 1² = 25+9+1 = 35.

3

The ratio, and why it detects concentration

Two matrices can have the same trace but a completely different shape:

Casevaluestracetrace²Σvalue²effective rank
Spread evenly (safe)3, 3, 39812781/27 = 3.0
Concentrated (dangerous)9, 0, 09818181/81 = 1.0

Same total risk (trace=9 both times), opposite shape. The spread-out case scores the full 3 (matches the true dimension — nothing hidden). The concentrated case collapses to 1, correctly flagging that only one direction actually carries the real risk, even though the matrix is technically 3-dimensional.

4

Dividing by real dimension gives the plotted "danger score"

hi_frac = effective_rank / hi_dim

Example, read straight from a real log line: H_I=15.6/6144 means effective_rank=15.6, hi_dim=6144, so hi_frac = 15.6/6144 = 0.00254. This fraction — near 0 when risk is concentrated, near 1 when it's spread out — is exactly what hessian_story_lib.py plots as "danger score."

2. Reconstructing the real spectrum — full derivation

To render a 3D shape, we need more than one summary number — we need an actual spectrum. This section derives, in closed form, the simplest real spectrum whose own effective rank exactly equals the measured one.

1

The model

A geometric-decay spectrum: λi = ri for i = 0, 1, ..., N-1, with 0 < r < 1. N is the tensor's real dimension, read straight from its log line (5120, 10240, or 17408 below).

2

Sum the spectrum (finite geometric series)

Σλi = (1 − rN)/(1 − r)

Since every real N here is ≥5120 and r<1, rN is astronomically close to 0 (even at r=0.88, 0.885120 has thousands of leading zeros). So Σλi ≈ 1/(1−r).

3

Sum the squares (same series, ratio is now r²)

Σλi² = Σ(r²)i = (1−r2N)/(1−r²) ≈ 1/(1−r²)

Same large-N justification — r² is also <1, so (r²)N is also ≈0.

4

Form the participation ratio and simplify

PR = (Σλi)² / Σλi² ≈ [1/(1−r)]² × (1−r²) = (1−r²)/(1−r)²

Factor (1−r²) = (1−r)(1+r) (difference of squares) and cancel one (1−r):

PR = (1+r)/(1−r)
5

Solve for r

PR(1−r) = 1+r  →  PR − PR·r = 1+r  →  PR−1 = r(1+PR)
r = (PR − 1) / (PR + 1)
Hand-check — verify this yourself
Real safest tensor in this dataset, L51.self_attn.k_proj, real measured PR = 15.9: r = 14.9/16.9 = 0.8817. Plug back into step 4's formula: (1+0.8817)/(1−0.8817) = 1.8817/0.1183 = 15.90 — matches the input exactly.
6

The three visible axes

Once r is known, the top three eigenvalues of this spectrum are λ0=1, λ1=r, λ2=r². An ellipsoid's radius along a given eigen-direction is proportional to 1/√λ (standard error-ellipse convention: high curvature = short, tightly-constrained axis; low curvature = long, forgiving axis). So the three real axis lengths, normalized to the longest:

axis0 = 1    axis1 = 1/√r    axis2 = 1/r

3. The five real cases used below — full audit trail

Every row traceable to an exact real log line. Source: ~/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume/logs/batch_*.log.

TensorRaw log linePRNraxes (norm.)
L51.self_attn.k_proj
safest tensor in this dataset
H_I=15.9/5120 15.951200.8817 1.00, 1.06, 1.13
L14.linear_attn.in_proj_qkv
H_I side
H_I=1.6/5120 1.651200.2308 1.00, 2.08, 4.33
L14.linear_attn.in_proj_qkv
H_O side — the documented, confirmed real cause of a real 302× quantization failure
H_O=1.9/10240 1.9102400.3103 1.00, 1.80, 3.22
L63.mlp.up_proj
H_I side — most dangerous tensor in round 1's full dataset
H_I=1.3/5120 1.351200.1304 1.00, 2.77, 7.67
L63.mlp.up_proj
H_O side — the single most concentrated tensor found
H_O=1.2/17408 1.2174080.0909 1.00, 3.32, 11.00

4. The real 3D shapes

Drag to rotate. Each ellipsoid's three axis lengths are exactly the numbers derived above — nothing decorative. A near-sphere means risk is genuinely spread out; a needle means risk is concentrated in one real direction.

drag to rotate · scroll to zoom
L51.self_attn.k_proj (safest, PR=15.9) L14.in_proj_qkv H_I (PR=1.6) L14.in_proj_qkv H_O (PR=1.9) L63.mlp.up_proj H_I (PR=1.3) L63.mlp.up_proj H_O (PR=1.2, most concentrated)
Axis ratios shown at real, non-uniform scale relative to each other within one ellipsoid (a sphere would be 1:1:1); ellipsoids are placed side by side for comparison, not to a shared absolute size scale, since the point being shown is shape (concentration), not the tensor's raw magnitude.