Every file, function, and real command in this chain, confirmed by reading the actual source tonight — not recalled from memory. Godmode did not invent this number; it was always a free byproduct of ordinary correction.
Deep-dive companion, not the story's spine — for the current, continuing account, start at Improvement Ledger §07 and read forward through §10. Everything on this page is still accurate as of 2026-09-16.
This page ends at the raw score. How the solver turns that score into a percentile rank, a danger weight and a bit-width is in How the Solver Turns a Hessian Score into a Bit-Width. Note that “danger score” there means three different numbers, and that page keeps them apart.
Four real steps, none of them built for godmode, all of them older than it.
The formula itself: trace(H)² / sum(H²). A cheap spectral participation-ratio diagnostic, chosen specifically to avoid a full eigendecomposition of a 10240-dim matrix on this machine.
Called during every single real correction, godmode or not — not a special path. Computes r_in = effective_rank(Hin_raw), r_out = effective_rank(Hout_raw), and prints them. This is the exact print every project script this session has been reading logs for.
The regex that scans any real logs/batch_*.log file and parses that print line back into {hi_frac, ho_frac} per tensor. Shared, reused logic — not reimplemented anywhere else.
106 real lines. Calls hessian_story_lib.load_hessian() (step 3, reused unchanged), averages hi_frac/ho_frac per tensor, writes one flat {tensor_name: score} JSON file. This is the file every Hessian-vs-KL comparison tonight has used.
Six real steps, in the exact order the code actually computes them. Each one defined in plain words first, then in real notation — the notation on its own (specifically H²) is genuinely ambiguous unless someone tells you which meaning it is, so that gets said explicitly, not assumed.
Before any of the formula below runs, H already exists — it's built from real gradients during the backward pass (full real trace here). Every step from here on takes that already-built matrix H as its input. This step is upstream of the score formula, not part of it.
Add up only the numbers on H's own diagonal (top-left to bottom-right). One real number out.
H² disambiguated hereThis is the step where H² is genuinely ambiguous notation. It does not mean matrix multiplication (H @ H). It means: take every individual number in the matrix, square that one number, then add up all of those squares into one real total. A student would write it as: square each cell, then sum the whole grid.
Here trace(H)² is ordinary squaring — trace(H) is already just one real number by this point (from step 1), so squaring it means multiplying that number by itself, nothing more exotic. Divide that by step 2's total. Real range: from 1 (this tensor's whole curvature sits in one direction — narrow, dangerous) up to H's own dimension (spread evenly across every direction — forgiving).
Dividing by the matrix's own real dimension turns step 3's raw number into a fraction, roughly 0 to 1 — which is what actually makes tensors of different sizes comparable on the same scale. This is the one step most easily missed: hi_frac is not effective_rank itself, it's effective_rank already divided by dimension.
A plain mean of the two fractions from step 4. This is the one real number every comparison in this project's whole Hessian-vs-KL investigation actually uses.
Two tiny, made-up 2×2 matrices — not from a real tensor, just small enough that every step below can be verified on paper. Real numbers, computed and checked directly before writing them here.
H_I = [ 4 0 ] dim_in = 2
[ 0 1 ]
Step 1 trace(H_I) = 4 + 1 = 5
Step 2 sum of squares = 4² + 0² + 0² + 1² = 16 + 0 + 0 + 1 = 17
Step 3 effective_rank = trace² ÷ sum of squares = 5² ÷ 17 = 25 ÷ 17 ≈ 1.4706
Step 4 hi_frac = 1.4706 ÷ 2 ≈ 0.7353
H_O = [ 2.5 0 ] dim_out = 2
[ 0 2.5 ]
Step 1 trace(H_O) = 2.5 + 2.5 = 5
Step 2 sum of squares = 2.5² + 0² + 0² + 2.5² = 6.25 + 0 + 0 + 6.25 = 12.5
Step 3 effective_rank = 5² ÷ 12.5 = 25 ÷ 12.5 = 2.0000
Step 4 ho_frac = 2.0000 ÷ 2 = 1.0000
Step 5 hess_score = (0.7353 + 1.0000) ÷ 2 = 0.8677
Notice H_O's effective_rank lands exactly at 2.0 — its own full dimension. That's not a coincidence: its energy is spread perfectly evenly (2.5 and 2.5, identical), which is exactly what "high score = forgiving" means concretely. H_I's energy is concentrated far more in one direction (4 vs. 1), so its effective_rank sits closer to 1 than to its own dimension — concretely what "low score = dangerous" looks like as actual numbers, not just a description.
def effective_rank(H: mx.array) -> float:
tr = float(mx.trace(H))
fro2 = float(mx.sum(H ** 2))
return tr * tr / fro2 if fro2 > 0 else 0.0
r_in = effective_rank(Hin_raw)
r_out = effective_rank(Hout_raw)
if verbose:
print(f" {label}: real effective rank -- H_I={r_in:.1f}/{Hin_raw.shape[0]}, "
f"H_O={r_out:.1f}/{Hout_raw.shape[0]}")
Real, literal output: language_model.model.layers.1.mlp.gate_proj: real effective rank -- H_I=14.2/5120, H_O=14.4/17408
safety_gate() runs as part of the normal correction math every real tensor goes through — computing H_in/H_out is required to build the LDL correction itself, not extra work done to produce a score. The score is a side effect of work that was always happening, on every real build this project has ever run.
def extract_hessian_scores(resume_dir: Path) -> dict[str, float]:
"""Real per-tensor hess_score, scanned directly from resume_dir/logs/batch_*.log."""
raw = hsl.load_hessian(str(resume_dir))
return {
name: (entry["hi_frac"] + entry["ho_frac"]) / 2.0
for name, entry in raw.items()
if "hi_frac" in entry and "ho_frac" in entry
}
Note it imports and reuses hessian_story_lib directly — the same shared library the tensor explorer and every research page in this folder already trust. Nothing here is a second, parallel implementation.
Run once against a completed, ordinary production build — not a godmode run.
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume \
--output /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json
Every path absolute — runs correctly no matter what directory your terminal is currently in.
| Moment | Correct? | Why |
|---|---|---|
| Day zero — no master file exists yet | Yes | Creates the file for the first time, from one finished build's real logs. |
Any time after — the file already has real data from elsewhere (e.g. a later watch run) | No | extract overwrites; it only ever sees the one resume-dir you point it at, so anything watch added from a different build is gone from what it writes. Use watch instead. |
This project's own file already had its real day zero — the correct command for it today is always watch, never this one.
Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected is a real, ordinary, already-finished production build — the same one used as this project's own KL-truth reference throughout. No sweep, no candidate bits, no godmode flag anywhere in that command.
Since 2026-09-14 this extraction also runs automatically, with zero manual step, at the end of any real run_full_yaqa.sh build — the script above is what you'd run by hand only against an older build that predates that hook.