← Back to index
Deep-dive companion · self-contained · 2026-09-16

Hessian vs. KL: the complete investigation, from zero

Everything from that night's investigation in one place — every term defined, every number verified against real data, every script and command traced to its real source. Written so it stands alone: no prior context required.

If you're picking this up cold (a new session, a compacted conversation, someone else entirely): read section 1 first for the vocabulary, then read top to bottom — each section only depends on the ones before it.
Where this sits in the real story (added 2026-09-17)

This page is a deep-dive companion, not the story's spine. The chronological, continuing account of this investigation lives in IMPROVEMENT_LEDGER — start at §07 and read forward through §10 for where the investigation stands now. This page is still accurate for everything it covers as of 2026-09-16 — it just isn't the up-to-date entry point.

1 · The cast of characters

Every term this page uses, defined once, here, before it's used anywhere else.

Hessian (H)

A matrix describing how a tensor's own output error curves as its weights change. This project builds a separate H_I (input-side) and H_O (output-side) per tensor from real calibration data — the same Hessians YAQA's own correction math already needs.

effective_rank(H)

trace(H)² / sum(H²) — a cheap diagnostic of how "spread out" a Hessian's energy is. Low = concentrated in a few directions (narrow, fragile). High = spread broadly (forgiving).

Hessian score / danger score

mean(effective_rank(H_I)/dim_in, effective_rank(H_O)/dim_out). One fixed number per tensor. Low = dangerous. High = safe. Computed once, never changes with bit-width.

Isolated KL

The real, measured output-distribution shift from quantizing one tensor to a given bit-width, with every other tensor still at full precision (bf16). A clean, "vacuum" measurement — but unrealistic, since a real deployed model never has everything else at full precision.

Cascaded KL (P0, P1, P2, P3)

The fix for isolated KL's unrealism: measure each tensor's real KL cost against the model's actual, already-partially-quantized state. P0 = the isolated baseline (no cascade yet). P1/P2/P3 = successive real rounds, each measured against the previous round's own real, applied plan.

YAQA correction

This project's real, curvature-aware rounding-error correction — uses a tensor's own H_I/H_O to correct quantization error smarter than naive rounding. Requires building the real Hessian anyway, which is where the danger score comes from for free.

MILP solver

02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py — the real optimizer that decides each tensor's final bit-width under a fixed total-BPW budget, by minimizing a real, weighted cost across all tensors at once.

BPW

Bits per weight — the real, average number of bits used per parameter across the whole model. The solver's fixed budget constraint.

Godmode sweep

A real, additional measurement pass (--godmode-multi-bit-checkpoint) that tests every candidate bit-width (3,4,5,6,8) for a tensor and records the real YAQA-corrected rounding error at each — not just one number, a whole curve.

Hybrid checkpoint

A real, merged input file for the (unmodified) MILP solver: godmode's real rounding-error data for tensors it has covered, real KL data as fallback for everything else.

2 · Why this investigation started

The real, expensive problem this was trying to solve.

Building a real, production-quality bit-allocation plan the way this project always has requires cascaded KL — which means real GPU time: an isolated KL pass (~30 hours), then a real MILP solve, then re-measuring the whole model against that applied plan (round 1), then re-solving, then re-measuring again (round 2), and again (round 3) before the plan stabilizes. Real, substantial GPU cost, every single time a new plan is needed.

The Hessian danger score is nearly free — a byproduct of computation YAQA's correction already has to do. The real question: does the free signal agree closely enough with the expensive one that it could replace some of that cost?

3 · What we actually found — the honest version

This investigation went through a real self-correction tonight. Both the overclaim and the correction are recorded here, because the correction is the actually-useful part.

The claim that didn't hold up

An earlier project document claimed a Hessian-driven solve reproduced the real, benchmarked KL plan almost exactly (3 of 497 tensors differing). Re-running the exact real comparison directly, controlled, tonight: 157 of 497 tensors (31.6%) actually differ. The original claim measured a different, narrower question (how often a tie-break mechanism engages) and got read as "the plans match" — they don't, not closely.

The real correlation, tensor by tensor

Directly computed, real data, all 360 tensors with both a Hessian score and an isolated-KL measurement:

0.00 +0.50 -0.50 P0 isolated P1 round 1 P2 round 2 P3 round 3

● Pearson   ● Spearman — real values: P0 −0.34/−0.40, P1 +0.28/+0.29, P2 +0.30/+0.31, P3 +0.31/+0.31. Both cross from negative to positive between P0 and P1 and stay there.

What "negative" and "positive" mean here, concretely

Hessian's danger direction is low = dangerous. KL's danger direction is high = dangerous. So real agreement between them produces a negative number (opposite directions, moving together toward "this tensor is risky"). A flip to positive means the two signals start naming different tensors as risky — not "no risk detected," an active disagreement about where the risk is.

A concrete example of that disagreement

TensorHessian rank (of 360)Isolated-KL rank (of 360)Reading
layers.3.self_attn.q_proj22agree
layers.63.mlp.up_proj1 (most dangerous)300 (near-safest)sharp disagreement

The layer-63 case lines up with a real, independent reason this project already protects layer 0 and layer 63 regardless of any signal — the "boundary floor." Full per-tensor list (356 tensors, excluding the boundary layers): hessian_scores/isolated_hessian_vs_kl_L0L63_excluded_FULL.csv, or browse it live in the tensor explorer.

Where Hessian turned out to be strong, not weak

Correlating the danger score against something else — not KL, but the real, measured benefit of YAQA's curvature-aware correction over naive rounding (pct_reduction in the manifest) — produced the strongest real result all night:

Hessian score vs.Real Pearson rReading
final error after correction (frob_yaqa)+0.02near zero — doesn't predict what's left over
improvement YAQA provides over naive (pct_reduction)−0.74strong — predicts how much correction was needed
The real, current interpretation

Hessian danger looks like a strong signal for "how much does curvature-aware correction help this tensor" — not a strong signal for "how much does this tensor's quantization hurt the model's output." Two different questions. Tonight's data answers the first one well and the second one only loosely.

4 · Where the Hessian score itself comes from

Real chain, verified directly against source tonight. Full depth on the dedicated origin page; summarized here so this page stands alone.

1

yaqa_core.py:343 — effective_rank(H)

trace(H)² / sum(H²) — the formula.

2

yaqa_core.py:434-437 — inside safety_gate()

Printed during every real correction pass, godmode or not — a free byproduct of math YAQA already needs.

3

hessian_story_lib.py:44 — RANK_RE

The regex that parses that print line back out of any real logs/batch_*.log.

4

08_extract_hessian_scores.py

Consolidates into one real {tensor_name: score} JSON. Built from an ordinary production run — Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected, no godmode involved. Since 2026-09-14 this also runs automatically at the end of any real run_full_yaqa.sh build.

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  extract \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume \
  --output /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json
The real, full lifecycle of this exact command — same command, two different moments, two different answers
MomentIs this command correct?Why
Day zero — no master file exists yetYES — this exact commandThis is precisely what creates the file for the first time, from one finished build's real logs. Correct, necessary, exactly this command.
Any time after — the file already has real data, e.g. from a later watch run against a different buildNO — do not re-run thisextract reads only the one resume-dir you point it at and overwrites the output. Real tensors that came from any other build (only watch adds those) are not in what it just read, so they're gone from what it just wrote. Use watch instead — section 5 below — which reads the existing file first and only ever adds.

This file's own real day zero already happened, before this investigation started — so today, right now, the correct command for this file is always watch, never this one.

5 · The godmode sweep — what it adds

The flat score is one number per tensor, forever fixed. The sweep is what turns that into a real curve.

Before (flat score only)After (godmode sweep)
What you know about one tensordanger = 0.02danger = 0.02, plus real error at 3/4/5/6/8-bit
Can you tell how many bits it actually needs?noyes — the real curve shows the payoff at each level

Real, complete command that produces it — part of a normal run, not a separate step. This is the exact command actually run tonight for the HESSIAN-PROBE job, every flag included:

PROJECT_ROOT=/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara

"$PROJECT_ROOT/scripts/yaqa_port/run_full_yaqa.sh" \
    /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE \
    --skip-lm-head-gptq \
    --plan "$PROJECT_ROOT/04_LANGUAGE_PLANS/HESSIAN_PROBE/plan_hessian_probe.json" \
    --n-calibration 4 \
    --incoherence none \
    --correct-mtp \
    --godmode-multi-bit-checkpoint \
    --godmode-candidate-bits 3,4,5,6,8

Real output, one line per tensor, appended to as the run progresses — for this exact command, that's /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume/godmode_sensitivity_checkpoint.jsonl:

{
  "layer_name": "language_model.model.layers.0.linear_attn.in_proj_a",
  "effective_rank_in": 1.726, "effective_rank_out": 1.695,
  "candidates": {
    "3": {"weighted_err": 0.000287, "safe": true},
    "4": {"weighted_err": 0.0000249, "safe": true},
    "5": {"weighted_err": 0.0000147, "safe": true},
    "6": {"weighted_err": 0.0000029, "safe": true},
    "8": {"weighted_err": 0.00000012, "safe": true}
  }
}

Note the score (effective_rank_in/out) is still right there at the top — free, same diagnostic, no extra work. The real new part is candidates.

To keep the master score file growing as a live run produces more real coverage — real, complete command:

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  watch \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume \
  /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --interval 30
This is the one to run today, for a job that's live right now

Same output file as section 4's command, different resume-dir, different subcommand — on purpose. This one only ever adds real tensors it finds, never removes or shrinks what's already there. Safe to run any time, against any live or finished job.

6 · The Hessian Hybrid checkpoint — real design, real result

Full depth on the dedicated design page; the essential mechanics here.

The one real insight that made this possible

The solver's existing per-tensor, per-candidate-bit cost already comes from one field: item["sensitivities"][bit]. Godmode's candidates[bit].weighted_err is the same real shape under a different name. So a small, separate converter — not a solver change — can produce a merged checkpoint the real, completely unmodified solver already knows how to read.

cascaded_checkpoint _round1.json (real KL) godmode_sensitivity _checkpoint.jsonl 08_extract_hessian_ scores.py hybrid per-tensor real merge merged checkpoint.json 02_optimize _v3_3.py unmodified

The real, disclosed merge rule

Per tensor, never blended: real godmode error at every bit it tested if the sweep has reached that tensor; otherwise the real, unmodified KL sensitivities. Every output entry tagged cost_source: "godmode" | "kl_fallback", so any finished plan can be audited afterward.

The exact real commands, in order, full paths, nothing to fill in

Step 1 — build the merged checkpoint:

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  hybrid \
  /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume/godmode_sensitivity_checkpoint.jsonl \
  --output /tmp/hessian_hybrid_checkpoint.json

Step 2 — solve against it, real solver, unmodified:

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  /tmp/hessian_hybrid_checkpoint.json \
  --target-bpw 5.028075122774708 --candidates 4,5,6,8,16 --group-size 64 \
  --late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 \
  --output /tmp/plan_hessian_hybrid.json

Run Step 1, then Step 2, in order — every path absolute, runs from any directory. Step 2 reads the exact file Step 1 just wrote.

The first real result (21/497 tensors covered)

MetricReal value
Tensors with real godmode coverage21 / 497
Of those, got a different bit than the benchmarked plan8 / 21
KL-fallback tensors that also changed (shared-budget ripple)11
Total differing from the real benchmarked plan19 / 497
Honest limitation

Only 21 of 497 tensors are real so far — this confirms the pipeline works end to end on real data, it is not yet a verdict on Hessian-driven allocation as a whole. Re-run the same two commands any time for a fresh, larger-coverage result.

7 · Every real command, in one flowing sequence

Same five commands as the sections above, gathered here in the order you'd actually run them — full absolute paths, nothing abbreviated, nothing to fill in yourself. Copy each block exactly as written.

1. Watch whichever live job is running

/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/watch.sh

Run with no arguments — it lists every real job it finds and asks which one to watch.

2. Extract Hessian scores from a finished build

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  extract \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume \
  --output /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json

3. Grow the master score file live, from any running job

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  watch \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume \
  /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/research_hadamard_blowup/hessian_scores/hess_scores_qwen38_v2_fp32_mtpcorrected.json \
  --interval 30

Swap the first path for whichever job's resume-dir you actually want to watch.

4. Build the hybrid checkpoint

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
  hybrid \
  /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
  /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume/godmode_sensitivity_checkpoint.jsonl \
  --output /tmp/hessian_hybrid_checkpoint.json

5. Solve against it — real solver, completely unmodified

python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
  /tmp/hessian_hybrid_checkpoint.json \
  --target-bpw 5.028075122774708 --candidates 4,5,6,8,16 --group-size 64 \
  --late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 \
  --output /tmp/plan_hessian_hybrid.json

8 · What's still open

9 · Related real pages

Author: Hakim Ghelab, VegaLaboratories LTD · Every number, file path, and code excerpt on this page is real — read directly from source or computed directly from real data tonight, none estimated, none illustrative.