Betting the same solver on a cleaner signal — not a new one
The real, unmodified MILP solver from §8 has never cared what produced its
sensitivities[bit] array — only that one real number exists per tensor, per candidate
bit-width. Real hypothesis, not yet confirmed: replacing cascaded KL in that array with real, measured
godmode rounding error, once coverage is complete, may give the solver a cleaner signal than it has ever
had. This page states the hypothesis precisely and says exactly what is still unproven.
Where §7–09 left things
This page picks up exactly where the solver debrief and the first benchmark left off — it does not restate their content.
§07 established that cascaded KL alone is an imperfect signal (it can be near-zero for genuinely dangerous tensors) and that the flat Hessian danger score is a second, KL-independent signal. §08 built the real hard-floor and lexicographic-priority mechanism on top of that flat score, and closed by validating the full P0→P1→P2→P3 cascade. §09 benchmarked the resulting build and found no regression — a real, positive result, on a build that predates everything below.
Real, confirmed trigger (2026-09-17): going to overlay real Hessian score data onto that clean, pre-Hessian §09 plan is exactly when the coverage gap was discovered — most tensors simply didn't have a real Hessian measurement yet. That discovery is what triggered building two things together, not sequentially: a sweep to close the coverage gap, and, in parallel, the better (lexicographic, hard-floor) solver mechanism in §08. It was specifically while building that better solver that the realization landed: godmode's richer per-bit data (the real error rate at every candidate bit-width, not just one flat number) could let a solver leverage 100% of the real Hessian weighting error, in theory — the real origin of this page's hypothesis.
The project's multi-bit godmode sweep (--godmode-multi-bit-checkpoint) has since been measuring
real, per-tensor, per-candidate-bit rounding error directly — not a flat score, a full sensitivity
profile, in the same shape the solver already consumes. The real, current merge mechanism
(08_extract_hessian_scores.py hybrid, see
its own design doc) blends this real measured
data in wherever the sweep has reached a tensor, and falls back to cascaded KL everywhere else.
What happens once there's nothing left to fall back to
Stated as precisely as the real code allows — this is a hypothesis being tested by real, ongoing coverage, not a conclusion.
The real solver, 02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py, reads a per-tensor
sensitivities dict at real line 1082: sens = {int(k): float(v) for k, v in
item["sensitivities"].items()}, and builds its real objective from it at real line 563:
weighted_kl = np.array([o["kl"] * structural_weight[o["name"]] for o in options]). Neither
line, nor anything else in the solver, checks where those numbers came from.
The real merge mechanism from §8 already writes into that exact same shape, one real entry per
tensor, tagged cost_source: "godmode" where the sweep has real measured data and
"kl_fallback" where it hasn't. The hypothesis is narrow: at 100% real coverage, every
entry is godmode and zero are kl_fallback — not because of any new code,
but because there is nothing left to fall back to. The array the solver reads becomes, for the first time,
entirely real measured rounding error instead of a mix of measured error and cascaded KL's own noise. Only
the input changes. The solver, the floors, the boundary protection all stay exactly as validated in §8.
§07 already showed cascaded KL can sit near zero for a tensor that real Hessian curvature says is dangerous — the entire reason the flat score and the floors exist. A pure, measured sensitivity checkpoint would remove that specific failure mode at the source, for every candidate bit, not just as a floor correction after the fact. Whether that actually produces a better real plan is exactly what is untested — see the closing section.
Four real, small steps, not one big idea
Each step exists because the one before it hit a real, measured limit.
| Step | What it did | Real limit that motivated the next step |
|---|---|---|
| Flat score as soft weight | Multiply KL by the flat danger score | KL near zero × anything is still near zero (§07) |
| Flat score as hard floor + lexicographic solver | Guarantee dangerous tensors a minimum bit-width directly, built as the real better solver (§08) | Built alongside the sweep below, both triggered by the same coverage-gap discovery |
| Multi-bit sweep built | Measure real YAQA rounding error at every candidate bit, per tensor | Confirmed trigger: overlaying Hessian data onto the clean, pre-Hessian §09 plan exposed the coverage gap — building the better solver in parallel is what surfaced the idea that richer per-bit data could let it use 100% of the real Hessian weighting error |
| Hybrid merge (today) | Real measured data where swept, cascaded KL elsewhere, tagged per tensor | Coverage is partial — 81 of 497 tensors, as of this page |
| Pure checkpoint (this hypothesis) | Same merge command, once coverage reaches 100% | Not yet reached — nothing to test against yet |
81 of 497 — not close enough to test the hypothesis yet
This number moves as the sweep runs — check it directly rather than trusting this page over time:
Three real questions this page does not answer
Named honestly so nothing here gets read as a settled result.
The early-QKV and late-attention floors exist to compensate for cascaded KL's own real noise. Whether they're still needed once the checkpoint is pure godmode data is untested — the real Hessian-measured data could carry its own, different reasons those regions need protection, unrelated to KL's noise specifically. This needs a real, measured test at full coverage, not an assumption either way.
This page's hypothesis assumes cascaded KL should be fully replaced once godmode coverage allows it. An open alternative: cascaded KL and real Hessian sensitivity may be measuring genuinely different real things — KL measures actual output divergence, the Hessian measures local curvature/rounding sensitivity — and the better real answer might be a measured way to combine both signals rather than discard one. No such combination formula has been derived or tested here; this is flagged as a real open question, not a plan.
A complete sweep means real, corrected weights get computed at every candidate bit for every tensor
during the run (even though only the scores are kept — see §08's godmode candidate-sweep
mechanism). The real disk cost is the resume directory's cached .safetensors files, not the
small .jsonl checkpoint itself. Not yet measured at full coverage; flag before assuming it's
free.
research_hadamard_blowup/exports/HESSIAN_HYBRID_CHECKPOINT_DESIGN.html
· Real solver line numbers verified directly against
02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py at write time · The three open questions
above are open because they are unverified, not because they are assumed false.