← Back to research index
Solver Deep Dive · Full Code Review · 2026-10-02

Hessian Score vs. GODMODE: the code, the pipeline, and what they actually disagree on

Two signals this project calls "the Hessian signal" are, on inspection, genuinely different measurements that happen to be loosely related. This page walks the real solver code line by line, then renders a real, rotatable 3D comparison of nine real tensors' actual curves — nothing illustrative, every number traced to a real file.

Every number below is real. Formulas are quoted directly from 08_extract_hessian_scores.py and 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py with real line numbers. The 3D scene plots nine real tensors' real GODMODE curves from hessian_hybrid_checkpoint.json, color-coded by their real hess_score percentile.

1. Two signals, one vocabulary

This project's own scripts and conversation both use "Hessian score" loosely. There are really two distinct files, computed completely differently, feeding completely different parts of the solver.

The aggregate score — one number per tensor, no bit-width axis

08_extract_hessian_scores.py, line 20-22
hess_score(tensor) = mean(effective_rank(H_I)/dim_in, effective_rank(H_O)/dim_out) where effective_rank(H) = trace(H)^2 / sum(H^2)

A spectral participation-ratio — how concentrated the tensor's curvature is across its own internal directions. Low score = energy concentrated in very few directions (narrow, dangerous). High = spread broadly (forgiving). It needs nothing about any candidate bit-width — it's a property of the curvature matrix alone.

Where it actually comes from

extract_hessian_scores(), line 90-97, doesn't run a dedicated measurement at all — it scans YAQA's own real correction logs for "effective rank" lines the correction pass already prints as a byproduct of computing H_I/H_O for its own two-sided correction math. A repurposed side-effect, not a purpose-built predictor.

GODMODE — a real curve, one point per real candidate bit-width

08_extract_hessian_scores.py, _merge_hybrid_checkpoint(), line 241-260
for bit_str, cand in godmode_sweep[tensor]["candidates"].items(): sensitivities[bit_str] = cand["weighted_err"] # the real, measured, Hessian-weighted # quantization error AT that exact bit-width

A direct, outcome-based measurement: "if I actually quantize this tensor to exactly 4-bit (or 5, 6, 8…), what does the real resulting error turn out to be?" — not a geometric proxy for an answer, the measured answer itself, at every real candidate.

2. How the solver actually uses each one

solve_once(), line 587-891. Real code, real line numbers for every claim below.

Every checkpoint's real per-bit value becomes o["kl"] — generically, regardless of which file it came from

line 1132, 1141
sens = {int(k): float(v) for k, v in item["sensitivities"].items()} ... options.append({ ..., "kl": float(sens[b]), ... }) # whatever file this is

The field is named "kl" for historical reasons — the checkpoint used to always be real KL-divergence data. The code never checks what kind of number it actually received. Point it at the real KL cascade or at GODMODE's real curve and the solver behaves identically: it optimizes whatever is in that file.

That becomes the real, final allocation — always

line 601, 824
weighted_kl = [o["kl"] * structural_weight[o["name"]] for o in options] # Q_global_weighted_KL ALWAYS runs -- this decides the real, final bits for every tensor

The aggregate score only ever drives an optional, separate, earlier phase

line 752-818
if hessian_primary and hessian_danger_weight: hess_deficit = [danger_rank[name] * (16 - bits) for each option] # LINEAR -- no curve shape at all hess_opt = solve(minimize hess_deficit) pin: all future solves must match hess_opt within tiny tolerance

Only when --hessian-primary and a real --hessian-score-file are both given. Its objective is a straight line in bits — it cannot see that a tensor's real benefit from extra bits might flatten out early, because it was never given the real curve to see that with.

Pseudocode, the whole pipeline

load checkpoint.json # per-tensor, per-bit-width real sensitivities for each tensor, each candidate bit-width: options[tensor, bits].kl = checkpoint[tensor].sensitivities[bits] if --hessian-score-file given: danger_rank[tensor] = percentile_rank(hess_scores_file[tensor]) # ONE number, no bit axis PHASE 1 (only if --hessian-primary AND --hessian-score-file): minimize sum(danger_rank[tensor] * (max_bits - bits)) pin best_deficit = achieved objective PHASE 2 (ALWAYS runs -- this is the real, final decision): minimize sum(kl[tensor, bits] * structural_weight[tensor]) subject to budget, floors, (if Phase 1 ran) deficit <= best_deficit + tiny_tolerance

3. How different these signals really are — measured, not assumed

Spearman correlation between hess_score and GODMODE's real error, across all 496 real tensors with both, every candidate bit-width:

Bit-width34568
Spearman ρ-0.528-0.529-0.530-0.529-0.529

Negative is expected (the two danger scales point opposite directions) — the magnitude is the real story. ρ≈0.53 squared is roughly 28% shared variance. A true "GODMODE aggregated to one number" relationship would read 0.9+. These are related, not interchangeable.

A real, named example of where they disagree sharply

layer.33.linear_attn.in_proj_qkv ranks only 25th-percentile dangerous by hess_score — comfortably mid-pack. Its real GODMODE error at 3-bit is 0.303 — the worst of the nine tensors sampled below, including the one hess_score calls the single most dangerous in the model (layer 63 mlp.up_proj, 0.053). A plan built from the aggregate score alone would under-protect this tensor; GODMODE catches it directly.

4. Nine real tensors, in 3D

Each ribbon is one real tensor's real GODMODE error curve across bit-widths 3→8. Drag to rotate, scroll to zoom, hover any sphere for its exact real numbers, click to fly the camera to that tensor's row.

How to read this scene — every encoding, explained

What a ribbon is: one real tensor, traced through five real, measured data points — its actual GODMODE-sweep error at 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit. The curve between points is a smooth interpolation for legibility; the five spheres on it are the only real measurements.

Height axis (vertical): real quantization error at that bit-width, square-root scaled (not linear) — the real values span roughly 0.0 to 0.3, a range wide enough that a linear scale would flatten every small-error tensor invisibly against the one or two large ones. Taller = more real error at that bit-width. A ribbon that drops fast and flattens early is a tensor that gets most of its benefit cheaply; a ribbon that stays tall even at 8-bit is one that's still paying a real cost that far in.

Depth axis (into the scene, Z): the nine tensors, laid out in order of their real hess_score percentile — most dangerous by the old aggregate signal at the back, safest at the front. This ordering is itself a claim the scene lets you test: if that ordering were a good predictor of real error, ribbon height should fall smoothly as you move from back to front. It doesn't.

Color: the same real hess_score percentile, as a gradient — deep red at 0.0 (the aggregate signal's "most dangerous"), through amber, to green at 1.0 ("safest"). Color and depth encode the identical real number; they're redundant on purpose, so the eye can read danger-rank from either axis while judging height against it.

The one thing to actually look for: whether tall ribbons are reliably red and short ribbons are reliably green. If hess_score and GODMODE agreed closely, they would be — color would predict height. They don't: the tallest ribbon in the whole scene (L33.in_proj_qkv, real 3-bit error 0.303) sits at a muted amber-red, 25th-percentile color, not the deep red you'd expect for the worst performer here. That single visual mismatch is the real ρ≈-0.53 correlation finding (section 3) made visible, not asserted.

drag · zoom · hover · click
HEIGHT = real GODMODE error at that bit-width
taller = more real error (worse)
COLOR = real hess_score percentile (a DIFFERENT signal)
red = "dangerous" · green = "safe" (old signal)
Tallest ribbon = most real error.
If color predicted height, tallest would always be reddest. Watch for when it isn't.
Real error ranking
Drag the slider to rank all nine tensors at a specific bit-width.
bit-width: full curve (drag to reveal)
Baseline rails mark y=0 for each tensor's row. Floating labels along the front edge name each bit-width; labels on the left name each tensor, with its own worst-case real error printed in. Hovering a sphere shows its exact real tensor name, bit-width, measured error, and hess_score percentile. Drag the slider below to sweep a highlight plane through the real data and re-rank every tensor by its actual error at that exact bit-width — nothing in any of this is interpolated or rounded for effect.

5. One real tensor, start to finish: granularity changing a real MILP decision

language_model.model.layers.1.linear_attn.in_proj_z — the Gated DeltaNet output gate for layer 1. Real curve, real param count, real bits the actual MILP solver chose when this exact tensor was part of a real 497-tensor budget-constrained solve.

Bit-width3456816
Real GODMODE error0.045930.010240.002400.000600.000040.0

Real param count: 31,457,280. Real bits the MILP actually chose for this tensor, in the live build's own plan file: 4-bit.

Why granularity matters here, concretely

The curve shows real, steep, early returns (3→4-bit alone removes 78% of the error) and then rapid diminishing returns after that (4→5 removes another 0.008 of error; 5→6 only another 0.0018; 6→8 only another 0.0006). A solver with only a single aggregate danger number for this tensor cannot see that shape — it can only ask "is this tensor dangerous or not," never "how much of this tensor's real risk is already captured by 4-bit." GODMODE's granularity is what lets the real MILP make an informed trade: give this tensor enough bits to capture the steep early improvement, then spend the rest of the budget on tensors whose curves are still steep at that point — instead of either starving it entirely (no data) or over-paying for it (a single flat "dangerous" flag with no sense of where the curve actually flattens).

6. The complete explanation, in full

Full detail, not an abbreviated summary — kept current as mistakes are found and corrected, rather than preserved as a historical transcript once they're known wrong.

On the real code mechanism (solve_once, Phase 1 vs. Phase 2)

Found it, and it fully reverses my earlier verdict. Layer 1's in_proj_z:

  • Aggregate hess_scores rank: 32nd of 496 most dangerous — genuinely near the top, which is why --hessian-primary pushed it to 16-bit.
  • Real GODMODE curve: 0.046 (3-bit) → 0.010 (4-bit) → 0.0024 (5-bit) → 0.0006 (6-bit) → 0.00004 (8-bit) → 0 (16-bit). Steep early, then flattens hard.

Here's the mechanism: --hessian-primary's Phase 1 objective is Σ weight × (16 − bits) — a linear deficit function. It treats every bit of shortfall as equally valuable for a tensor, scaled only by its rank. It has no idea this tensor's real curve already captures ~94% of its achievable benefit by 5-bit. It just sees "ranked 32nd most dangerous" and keeps demanding more bits all the way to 16, because the linear model can't see the diminishing returns.

The GODMODE-only run uses the real, measured curve directly — so it correctly recognizes this tensor's marginal benefit beyond 5-bit is tiny, and spends the saved budget on tensors whose real curves are still steep at that point. That's not a gap in protection — that's the clean data doing exactly what you'd want: real curve shape beats a crude linear proxy of the same underlying signal.

This reverses my earlier verdict. The aggregate hess_scores + --hessian-primary mechanism isn't a refinement on top of GODMODE — it's a lossier proxy of the same data, and layering it back on top of the real curve throws away exactly the information that makes GODMODE valuable in the first place. 293 of 497 tensors (59%) differ between the two plans — this isn't a minor tweak, it's a fundamentally different, more information-respecting allocation.

Can you derive the Hessian score from a GODMODE sweep, and why the plans still diverge

Yes — this project's own tooling is built around exactly that. 10_build_hessian_probe_plan.py exists specifically because 136 of 497 real target tensors had no hess_score coverage; it builds a "probe plan" so those tensors clear the pipeline's selection gate, then a real run_full_yaqa.sh build runs with --godmode-multi-bit-checkpoint so the GODMODE sweep itself recovers their missing real effective-rank data, in one real pass (its own docstring: "every probed tensor's real effective rank gets captured... in one real pass"). Verified directly: running 08_extract_hessian_scores.py extract against Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume (that probe build's own resume directory, the same one the real godmode_sensitivity_checkpoint.jsonl sweep lives in) pulls 128 real tensors' worth of hess_score straight out of that sweep's own logs — a bit-for-bit exact match (max abs diff 0.0) against the separately-maintained master hess_scores_qwen38_v2_fp32_mtpcorrected.json. Same real H_I/H_O Hessian, same correction run, same logs.

So why do the plans still diverge on 293 of 497 tensors (59%)? Because "derivable from the same sweep" and "the same ranking once derived" are different claims. hess_score is effective_rank(H)/dim — a bit-width-independent spectral participation ratio (how spread out the Hessian's eigenvalues are). GODMODE's weighted_err is the actual measured quantization loss at a specific bit-width. Both are real functions of the exact same real H_I/H_O data — they just extract different information from it. Real Spearman correlation between the two resulting rankings: ρ ≈ -0.53, consistent across every bit-width 3 through 8 (negative sign expected, their danger scales point opposite directions) — a moderate relationship (~28% shared ranking variance), not the 0.9+ you'd expect from two views of the same single number.

Why the plans diverge, then — it's two compounding reasons, not one:

  1. Derivable from the identical real sweep data, yet the two derived rankings only moderately agree (ρ≈-0.53) — roughly half the real tensor-by-tensor ranking information in one isn't explained by the other.
  2. Even the half they do agree on gets used completely differently — hess_scores only ever drives a crude linear deficit in an optional Phase 1 that can't see curve shape at all (danger_rank × (16 − bits)); GODMODE drives the real, curve-aware Phase 2 objective directly.

On the real formulas themselves

hess_scores (the old "Hessian danger score") — formula, from the script's own docstring:

hess_score(tensor) = mean(effective_rank(H_I)/dim_in, effective_rank(H_O)/dim_out) where effective_rank(H) = trace(H)^2 / sum(H^2)

This is a spectral participation-ratio — a pure, geometric property of the curvature matrix's eigenvalue spread. Low score = the Hessian's energy is concentrated in very few directions (narrow, dangerous); high score = spread broadly (forgiving). Crucially, confirmed directly in extract_hessian_scores(): this is computed by scanning YAQA's own correction logs for "effective rank" lines the correction pass already prints as a byproduct of computing its H_I/H_O curvature matrices for the real correction math — it was never a purpose-built measurement at all, it's a repurposed side-effect of something else. And critically: it needs nothing about any specific bit-width. It's asking "how narrow/concentrated is this tensor's curvature, in the abstract" — never "what happens if I actually quantize this to 4-bit."

GODMODE, confirmed directly in _merge_hybrid_checkpoint(): each tensor's real sweep data has a "candidates" dict, one real "weighted_err" value per actual candidate bit-width, copied directly into the per-bit sensitivities. This is asking a completely different, much more direct question: "if I actually quantize this tensor to exactly this bit-width, what is the real resulting error turn out to be?" — an outcome-based measurement, not a geometric proxy.

Why GODMODE is more powerful, precisely: hess_score answers "is this tensor's curvature shape narrow or wide" — a reasonable but indirect proxy for danger. GODMODE answers "what actually happens at each real bit-width" — the literal thing you want to know. A tensor can have a narrow, concentrated curvature spectrum (hess_score flags it dangerous) without that necessarily producing large real error at a given bit-width — or the reverse. That's exactly why the correlation between them is only moderate (ρ≈-0.53, not near ±1): they're asking related but genuinely different questions, and GODMODE is the one that measures the real outcome directly instead of inferring it from the curvature's shape.

The honest limit of what's been shown

Not yet a validated breakthrough — a well-grounded, real finding

Shown, real, verified: GODMODE vs. the aggregate hess_score — moderate correlation (ρ≈-0.53), and GODMODE is structurally more direct (real measured outcome vs. a spectral proxy). Shown, real, verified: in a real solver run, GODMODE directly replacing KL as the Phase 2 objective produced a very different plan. Not shown: that a GODMODE-built model actually scores higher on real benchmarks than a KL-built or hess_score-built one. That is the one claim that would make this a breakthrough, and it remains unmeasured until a real build finishes and gets benchmarked.

7. The real paper's own KL methodology — a third, separate meaning of "KL"

Read directly from the real reference repo (arXiv 2505.22988), eval/eval_kl.py — not the paper's prose, the actual evaluation code.

loss_fct = torch.nn.KLDivLoss(reduction='batchmean') ... output = model(input, ...) # the QUANTIZED model's forward pass orig_output = orig_model(input, ...) # the ORIGINAL full-precision model's forward pass loss = loss_fct(output.log_softmax(dim=-1), orig_output.softmax(dim=-1))

Both the full model and the quantized model run forward over real wikitext2 text (8192-token sequences), and the KL divergence between their final next-token probability distributions is measured — averaged over every token, every sequence. This is a single, global, whole-model, end-to-end number, computed only after the entire model is already quantized. The paper's real ≈30% figure is this number, compared against GPTQ/LDLQ's version of this same global number, at the same fixed bit-width.

Three real things, one shared name

1. The paper's real 30% claim — global, whole-model, post-quantization KL, their actual benchmark metric.
2. This project's "isolated KL" — local, per-tensor, pre-decision, the signal cascaded-KL allocation is built on.
3. hess_score — not KL at all, a spectral curvature proxy (effective rank).
None of these three are the same measurement. Citing the paper's 30% number as evidence for any claim about per-tensor allocation quality applies it to a question the paper never measured.

8. The real benchmark: flat vs. allocated, full table

Every number read directly from a real, completed eval log (temperature=0, deterministic decoding), Improvement Ledger §09.

ModelMMLUGSM8KIFEval strict/looseBFCLHumanEval5-test mean
YAQA (P1, cascaded-KL, this build)87.7%98.0%88.0 / 88.094.0%93.0%92.14%
v689.5%97.0%90.0 / 91.090.0%94.0%92.10%
v33floor89.5%100.0%86.0 / 87.093.0%92.0%92.10%
stock588.6%99.0%86.0 / 86.094.0%92.0%91.92%
v6matched86.8%99.0%89.0 / 90.091.0%93.0%91.76%
flat5-naive (no allocation)90.4%99.0%82.0 / 84.093.0%94.0%91.68%
What this real table actually proves

Allocation wins the aggregate (92.14% vs. 91.68%, a real 0.46-point margin) and wins IFEval decisively (88.0% vs. 82.0% strict — flat is "the lowest of any model by a wide margin"). But flat wins real MMLU outright (90.4%, highest of six) — raw factual recall barely notices naive quantization. The mechanism isn't a strict free lunch across every axis; it's a real, net-positive trade that helps most where structured, rule-following behavior is at stake, and costs least where it doesn't.

9. Why reallocation works, mechanistically — two real tensors, proving the trade

The MILP objective (section 2 above) minimizes total weighted error under a fixed average-bits budget. That forces exactly one trade: take bits from tensors with little marginal benefit, give them to tensors with a lot. Two real tensors from this model show the real mechanism making exactly that call.

A

The "resistant, low-value" case — real budget correctly withheld

L1.linear_attn.in_proj_z: real curve flattens hard after 5-bit (6→8-bit buys only another 0.0006 of error reduction). The real MILP gave it just 4-bit — declining to spend more budget where the real data shows it wouldn't help much.

B

The "looks safe, isn't" case — real budget correctly spent

L33.linear_attn.in_proj_qkv: only 25th-percentile dangerous by the old aggregate score — yet its real GODMODE error is the single worst of the nine-tensor sample, worse even than the tensor the aggregate score calls most dangerous. A signal that can't see this tensor's real curve would under-protect it; GODMODE catches it directly.

The mechanism is proven; the open question is narrower than it looks

Both real examples show the reallocation logic working exactly as intended. What remains genuinely open isn't whether "take from dumb tensors, give to critical ones" helps — the MMLU-vs-mean split in section 8 already shows it's a real, net-positive trade, not a universal one. What's open is whether GODMODE's specific picture of which tensors are dumb and which are critical is more accurate than cascaded KL's picture — since the two signals disagree on roughly half their real ranking information (ρ≈-0.53, section 3). That's the one number only the finished, benchmarked build can supply.