2026-09-14
Author: Hakim Ghelab, VegaLaboratories LTD
Superseded as the recommendation document — kept as the raw
supporting data. The current, authoritative strategy (updated
2026-09-15, including --boundary-min-bits and the
per-tensor hard floor / lexicographic mechanism this document predates)
lives in IMPROVEMENT_LEDGER/08_HESSIAN_HYBRID_SOLVER_DEBRIEF.html
§7 and §8. The correlation and stability tables below are still real and
accurate; the exact command at the bottom of this file is
stale — it points at V4_CLEAN_STRATIFIED_CASCADE
(the original checkpoint lineage) rather than the corrected
V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected used in every
real run since, and it predates --boundary-min-bits
entirely. Do not copy it directly — use the command in §8 of the debrief
above instead.
This is the final, corrected version of the underlying investigation. Earlier drafts wrongly called the cascade measurement “broken” based on a narrow 3-tensor spot check. That was wrong and is retracted here. The cascade is real and legitimate — see the direct evidence below.
This is a planning-time-only simulation. No model
has been built. No real eval (MMLU/IFEval) has been run against any plan
named in this document. Every plan JSON’s own warning field
says it plainly: “Planning proxy only. Validate the final physical
artifact separately.” The next real step, when ready, is building a
real model from the real recommended plan below and running a real eval
against it.
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-OptiQ-5bpw_CLEAN_STRATIFIED/sensitivity_checkpoint.json).
Measures each tensor alone, assuming every other tensor stays bf16 — a
real, known limitation, since that’s never true in the real deployed
(all-quantized) model. Confirmed correct: matches
03_LANGUAGE_SENSITIVITY/analysis_Stratified exactly, and
its own no-Hessian, floors-on baseline plan reproduces the real shipped
production plan’s raw KL almost exactly (1.9412 vs. the real production
plan’s own recorded 1.9411543848545987).04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE/cascaded_checkpoint_round{1,2,3}.json),
seeded from P0’s own resulting plan
(04_LANGUAGE_PLANS/V3_3_APPLES_TO_APPLES_0005/plan_Qwen38_V3_3_5.028075122774708_BPW.json
as --baseline-plan), built with
--calibration-mode stratified. Exists specifically to fix
P0’s blind spot: each round measures every tensor against the real,
already- quantized state the model actually has, not a pristine one.
Real, direct evidence this is legitimate, not broken:
checked the full 497-tensor spread, not just a few tensors — P1’s real
coefficient of variation is 0.58 (min 0.0108, max 0.1994), genuine
per-tensor differentiation, not flat/uniform data.Correlated each round’s own real plan (bits assigned, no Hessian) against the real, independent Hessian danger score, across all 360 scored tensors:
| Round | Real Pearson r | Real Spearman r | Reading |
|---|---|---|---|
| P0 (isolated) | −0.1173 | −0.3597 | negative — agrees with Hessian |
| P1 (cascade, round 1) | +0.0505 | −0.1596 | negative — agrees with Hessian |
| P2 (cascade, round 2) | −0.0841 | −0.1389 | negative — agrees with Hessian |
| P3 (cascade, round 3) | −0.0579 | −0.1211 | negative — agrees with Hessian |
Negative correlation (a real, dangerous tensor — low score — getting more real bits) holds at every single round. The independent Hessian signal agrees with the cascade’s own correction direction the whole way through — this is the real confirmation that the cascade’s correction is sound, not noise.
Scatter-plot view of the same P0-P3 rounds, real Hessian percentile rank vs. real assigned bits, all 360 scored tensors plotted individually: hessian_rank_vs_bits.html. Note: that chart’s own labeled Pearson/Spearman numbers are from an earlier computation pass and don’t exactly match the table above (Spearman is close; Pearson differs) — treat it as a supplementary visual of the same negative-correlation shape, not a verified duplicate of this table’s exact figures.
| Transition | Real tensors that change bits | % of 497 |
|---|---|---|
| P0 → P1 | 167 | 34% |
| P1 → P2 | 83 | 17% |
| P2 → P3 | 68 | 14% |
The real, large structural correction happens at P0→P1. What follows is real but shrinking — genuine stabilization around an already-corrected plan, not a second major restructuring. P1 is the real, correct stopping point. P2/P3 are not worth the real GPU cost of chasing further.
--late-attn-min-bits and
--early-qkv-min-bits exist because isolated KL (P0) has a
real, documented, measured unreliability in specific regions (34.4%
sensitivity-ranking inversion rate in layers 0-15 — see
RUNNING_GUIDE.md §11 for the real incident this was built
to prevent, a real MMLU regression). The floors are a hard-constraint
workaround for that specific, known measurement flaw — not an
independently necessary rule in their own right.
The cascade (P1) exists specifically to fix that exact flaw — measuring each tensor against the real, already-quantized state instead of a pristine, misleading one. The Hessian’s continued real agreement with P1 (negative correlation, same as at P0) is the real confirmation that this correction worked: the noise the floors were built to guard against is addressed at the source once P1 is used instead of P0. This removes the justification for keeping the floors on.
This section is intentionally short. The full,
current, maintained recommendation — including the real per-tensor hard
floor and lexicographic mechanisms this sweep predates — lives in IMPROVEMENT_LEDGER/08_HESSIAN_HYBRID_SOLVER_DEBRIEF.html
§7 (this sweep’s own P0→P3 validation, reproduced there) and §8 (the
real, current strategy and exact command). Duplicating it here would
just create a second copy to keep in sync; this file’s job is the raw
sweep data above, not the conclusion.
What this sweep did real, correctly establish, and which still holds:
strength≥3 triggers a real, sharp cliff in this configuration (43
tensors reshuffle at once) with no independent justification beyond a
raw-KL trend already shown to be potentially misleading —
--hessian-weight-strength 2.0 remains the sensible default
for the soft weight specifically, separate from whether the newer
hard-floor/lexicographic mechanisms are also active.
Standing caveat, unchanged: none of this is quality. The real next step is building a real model from the current recommended plan (§8 of the debrief) and running a real eval against it.