Full, step-by-step trace of the pipeline built tonight — two real data sources, one small converter script, and the exact same MILP solver this project has always used, completely unmodified. Every code line and number on this page is real, pulled directly from the actual files.
This design doc is a deep-dive companion to Improvement Ledger §08. What happens once the mechanism described here reaches 100% real coverage — a real, open hypothesis, not yet confirmed — is written up in §10, the pure sensitivity hypothesis.
Two real, independent files. Neither is new tonight — both already existed; what's new is combining them.
04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json — the same real cascade-round-1 measurement this whole project's production plans have always been built from. One entry per tensor:
{
"layer_name": "language_model.model.layers.0.mlp.up_proj",
"sensitivities": {"4": 0.0189, "5": 0.0163, "6": 0.0136, "8": 0.0143, "16": 0.0},
"param_count": 245760
}
sensitivities is the real, measured KL cost of quantizing this one tensor to each candidate bit-width — this is what the solver has always consumed.
Real example, this exact job: /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume/godmode_sensitivity_checkpoint.jsonl — produced live by 05_full_model_quantize.py --godmode-multi-bit-checkpoint, one line per tensor, appended to as the real run progresses:
{
"layer_name": "language_model.model.layers.0.linear_attn.in_proj_a",
"effective_rank_in": 1.726, "effective_rank_out": 1.695,
"candidates": {
"3": {"weighted_err": 0.000287, "frob": 2.02, "safe": true},
"4": {"weighted_err": 0.0000249, "frob": 0.96, "safe": true},
"5": {"weighted_err": 0.0000147, "frob": 0.47, "safe": true},
"6": {"weighted_err": 0.0000029, "frob": 0.23, "safe": true},
"8": {"weighted_err": 0.00000012, "frob": 0.06, "safe": true}
}
}
candidates[bit].weighted_err is the real, measured YAQA-corrected rounding error at that bit-width — a fundamentally different measurement than KL (structural/curvature-driven, not output-distribution-driven), but the exact same real shape: one number per tensor per candidate bit.
Both files describe the same thing — a real number per (tensor, candidate bit-width) — just under different key names (sensitivities vs. candidates[bit].weighted_err). The solver never cares which measurement produced its numbers, only that they exist in this shape. That's the whole reason no solver code needed to change.
One subcommand of the consolidated real toolkit: 08_extract_hessian_scores.py hybrid. It never touches the solver — its only job is producing a third file, in Source A's exact shape.
| Tensor has real godmode data? | What its sensitivities become | cost_source tag |
|---|---|---|
| Yes | Real weighted_err at every bit the sweep tested (3,4,5,6,8). Bit 16 always kept from Source A's own 0.0 convention — the sweep never tests 16-bit. | godmode |
| No (not reached yet) | Untouched, exactly as in Source A. As if this script never ran for that tensor. | kl_fallback |
Nothing is averaged, blended, or guessed. Every output tensor is 100% one real signal or the other, tagged so a finished plan can be audited afterward for exactly which tensors were Hessian-driven.
02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py. Not one line edited. Real proof, from the actual source:
options.append({
"tensor_i": i,
"name": item["layer_name"],
"bits": int(b),
"kl": float(sens[b]), # <- straight from sensitivities[bit], whatever produced it
"params": int(item["param_count"]),
})
Real line ~1091. sens comes directly from item["sensitivities"] — the solver has no branch, no flag, no special case for where that number came from.
weighted_kl = np.array([o["kl"] * structural_weight[o["name"]] for o in options])
Real line 563. Every existing mechanism — the boundary floor, the run guard, late-attn/early-qkv floors, --hessian-primary, everything — sits on top of this exact same array, untouched.
This isn't a new optimizer. It's the same real, already-trusted solver, fed a checkpoint file whose numbers happen to come from a different real measurement for some tensors. Every floor, every constraint, every piece of production logic still applies exactly as it always has.
Two commands. Safe to re-run any time as the live sweep covers more tensors — each run is a fresh, real snapshot, nothing cached or stale.
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/08_extract_hessian_scores.py \
hybrid \
/Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/cascaded_checkpoint_round1.json \
/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-HESSIAN-PROBE.yaqa_resume/godmode_sensitivity_checkpoint.jsonl \
--output /tmp/hessian_hybrid_checkpoint.json
Real, complete, copy-pasteable — every path absolute, runs from any directory. Swap the godmode-checkpoint path for whichever job's sweep you actually want to use.
python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py \
/tmp/hessian_hybrid_checkpoint.json \
--target-bpw 5.028075122774708 --candidates 4,5,6,8,16 --group-size 64 \
--late-attn-min-bits 5 --early-qkv-min-bits 6 --lm-head-min-bits 6 \
--output /tmp/plan_hessian_hybrid.json
Takes the exact file Step 1 just wrote. Every flag shown — nothing truncated, nothing left for you to guess.
The floor values above match the exact real command that produced the actual benchmarked, shipped plan — kept identical here on purpose, so the only variable between the two plans is where each tensor's cost number came from.
Run tonight, with the live sweep at 21 of 497 tensors covered. Small coverage, but the pipeline runs end to end on real data.
| Metric | Real value |
|---|---|
| Tensors with real godmode coverage | 21 / 497 |
| Of those, got a different bit than the benchmarked plan | 8 / 21 |
| KL-fallback tensors that also changed (shared-budget ripple) | 11 |
| Total tensors differing from the real benchmarked plan | 19 / 497 |
Layer 0's mlp.down_proj/gate_proj/up_proj all dropped hard under real rounding-error data — 16-bit (KL-driven) down to 6/4/4-bit. This test used the exact same floors as the real benchmarked build, which did not include a boundary floor at all — so nothing here was artificially protected or held back. Worth re-checking once the sweep covers more of layer 0's neighborhood, and before this ever becomes anyone's default.
This is not yet a verdict on Hessian-driven allocation — only 21 tensors are real so far. It confirms the mechanism works correctly end to end on real data, and gives a first, honest, very partial look at where the two signals actually disagree.
IMPROVEMENT_LEDGER/08 and the Hessian-vs-KL tensor explorer.