Both models use the identical solver (02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py, single-stage Q_global_weighted_KL) and the identical generic 100x structural-weight protection on layer 0, layer 63, and lm_head. What differs is which extra hard floors are stacked on top of that weight — and that difference, not the sensitivity data itself, is the leading explanation for the Stratified build's speed edge.
plan_godmode_ONLY.json, plan_Qwen38_V3_3_5.028075122774708_BPW.json), cross-checked byte-for-byte against each model's real deployed config.json quantization map. Both models' generation quality was independently verified with mtplx forge verify on the uncapped prompt suite — results in Section 1, not assumed.Before anything else: neither model's speed advantage comes at the cost of broken generation. Verified directly, not inferred.
Both ran against mtplx forge verify's default uncapped suite (long-code-uncapped, up to 10,000 tokens per case) — the same real test, same command shape, run independently against each model's own real resume directory.
Real mtplx tune results, both measured on the same Mac (Apple M3 Max, 128GB), same day, same profile.
| Mode | GODMODE tok/s | GODMODE acceptance | Stratified tok/s | Stratified acceptance |
|---|---|---|---|---|
| AR | 18.19 | — | 18.19 | — |
| D1 | 39.51 | 96.3% | 40.15 | 98.4% |
| D2 | 52.47 | 98.0% | 52.60 | 97.6% |
| D3 (default) | 54.83–55.20 | 84.5% | 56.04–56.07 | 93.5% |
The Stratified build wins D1 and D3 outright on both speed and acceptance — D2 is close either way. D3 is the mode MTPLX's own tuner selects by default for both models (verdict: "mtp_depth_wins", zero rejections, zero collapses on both), so the D3 gap is the one that actually matters in deployment.
lm_headEvery tensor the structural floors touch, both models, exact bits and exact isolated-KL cost at that bit — no ranges, no rounding. group_size=64, mode=affine, and effective_kl_weight=100.0 are identical in every row for both models, so they're omitted from the table rather than repeated 32 times.
| Tensor | GODMODE bits | GODMODE isolated_kl | Stratified bits | Stratified isolated_kl | Params |
|---|---|---|---|---|---|
| L0 in_proj_a | 8 | 1.1697e-07 | 16 | 0.0 | 245,760 |
| L0 in_proj_b | 8 | 7.3127e-07 | 16 | 0.0 | 245,760 |
| L0 in_proj_qkv | 8 | 1.0828e-04 | 16 | 0.0 | 52,428,800 |
| L0 in_proj_z | 8 | 1.4380e-04 | 16 | 0.0 | 31,457,280 |
| L0 out_proj | 16 | 0.0 | 16 | 0.0 | 31,457,280 |
| L0 mlp.down_proj | 6 | 4.2735e-04 | 16 | 0.0 | 89,128,960 |
| L0 mlp.gate_proj | 6 | 2.2037e-05 | 16 | 0.0 | 89,128,960 |
| L0 mlp.up_proj | 6 | 2.2891e-05 | 16 | 0.0 | 89,128,960 |
| L63 k_proj | 8 | 3.3825e-06 | 16 | 0.0 | 5,242,880 |
| L63 o_proj | 6 | 2.3504e-05 | 16 | 0.0 | 31,457,280 |
| L63 q_proj | 6 | 5.4534e-05 | 8 | 1.6342e-04 | 62,914,560 |
| L63 v_proj | 8 | 4.2745e-06 | 16 | 0.0 | 5,242,880 |
| L63 mlp.down_proj | 6 | 4.1892e-04 | 8 | 9.7445e-05 | 89,128,960 |
| L63 mlp.gate_proj | 6 | 1.0380e-04 | 8 | 8.8229e-05 | 89,128,960 |
| L63 mlp.up_proj | 6 | 6.6777e-04 | 8 | 1.0966e-04 | 89,128,960 |
| lm_head | 16 | 0.0 | 6 | 5.2681e-04 | 1,271,398,400 |
At 16-bit, isolated_kl is always exactly 0.0 by construction — it's the lossless reference point every other bit-width is measured against, not evidence of "no cost."
The generic 100x boundary/output weight is identical in both builds — it is not what differs. What differs is which extra hard-minimum floors are stacked on top of it.
| Floor | GODMODE setting | Stratified setting | Real effect |
|---|---|---|---|
| early_qkv_floor | off — 0 bits, 0 tensors | on — 6-bit min, 48 tensors | Forces Q/K/V input-projections in layers 0–15 above their natural KL-optimal bit. Only applies to 3 of layer 0's 8 tensors. |
| late_attention_floor | off — 0 bits, 0 tensors | on — 5-bit min, 76 tensors | Same idea for late attention. Covers 4 of layer 63's 7 tensors (k/o/q/v_proj). |
| boundary_floor | on — 6-bit min, 15 tensors | not present (older plan schema) | GODMODE-only mechanism; Stratified's plan predates it and relies on the generic 100x weight alone for the tensors this would cover. |
| lm_head_floor | on — 8-bit min | not present | The one floor that actually decides an outcome here. GODMODE's floor forces lm_head above 8-bit; the 100x weight then pushes it the rest of the way to 16. Stratified has only the weight, no floor — and under the identical real sensitivity numbers (both checkpoints carry the same KL-fallback value for lm_head), it settles at 6-bit on its own. |
| max_low_bit_run (run guard) | off (0) | on (3) | Never triggered in either build's real solve (run_guard_promotions: [] in both) — empirically had zero effect on either real outcome. |
"Not present" never means "no bit value" — every Stratified tensor still landed on a real, specific bit, decided by the generic 100x weight plus that build's own real sensitivity data, with no extra floor needed. In every "no floor" case except lm_head, the weight alone was already enough to land at 16-bit. lm_head is the one case where the extra floor is the actual deciding factor, not just insurance.
Your read on why the Stratified build wins, quoted directly:
The per-tensor table in Section 3 is consistent with this: every one of layer 0's 8 tensors sits at full 16-bit in the Stratified build (zero quantization noise entering the stack), versus GODMODE's 6–8 bit layer 0. Layer 63 mostly follows the same pattern before lm_head. The open question the table can't answer on its own is whether that upstream cleanliness is what lets lm_head compress safely to 6-bit without hurting MTP agreement — or whether lm_head's own compression would do just as well sitting downstream of a noisier stack, and the two effects are independent. Section 6 is the real test for that.
A new GODMODE plan, same real GODMODE sweep checkpoint, same target BPW — only the early/late floors changed to match the Stratified build's real values, and the lm_head floor lowered to let the (identical, real, shared) sensitivity data decide on its own instead of forcing it. This isolates the floor settings as the variable, keeping the GODMODE sweep data itself unchanged — a real test of your hypothesis, not a guess.
4 is the lowest real candidate bit-width in this plan's own candidate_bits list ([4, 5, 6, 8, 16] — no 3-bit candidate here). Setting the floor to 4 is the closest real equivalent to "no floor at all," since the solver can never choose below its lowest candidate anyway.
This changes only the floor settings — the real GODMODE sweep sensitivity data, target BPW, and candidate bits are untouched. If lm_head lands at 6-bit under this command (expected, since its real sensitivity number is identical in both checkpoints) and the resulting build's D3 speed closes the gap with the Stratified build's 56 tok/s, that's real evidence the floor settings — not the underlying sweep data — were the speed driver. It does not yet test the quality side of your hypothesis (whether a cleaner layer 0 is what protects MTP agreement specifically) — that question is answered in section 7 (GODMODE-v2 at n=200, completed 3 October 2026).
GODMODE-v2 (Hessian middle network, lm_head floor 6, boundary floor 16) was run through the full n=200 suite and compared with HESSIAN-PROBE (cascaded KL, YAQA-corrected), the same protocol and seeded question set.
| Task | GODMODE-v2 | HESSIAN-PROBE | Leader |
|---|---|---|---|
| MMLU (n=171) | 87.7% | 90.1% | HESSIAN-PROBE |
| GSM8K | 98.0% | 97.0% | GODMODE-v2 |
| IFEval strict, prompt (instruction) | 89.5% (92.6%) | 90.5% (93.5%) | HESSIAN-PROBE |
| BFCL | 92.0% | 92.5% | HESSIAN-PROBE |
| HumanEval (n=164) | 95.7% | 95.1% | GODMODE-v2 |
| Mean of 5 core | 92.58 | 93.04 | HESSIAN-PROBE by 0.46 |
| D3 decode speed (tuned) | 55.9 tok/s (3.11x) | 51.4 tok/s, D2 best (2.86x) | GODMODE-v2 |
Speed: confirmed. With the floor settings changed, the Hessian build reaches D3 and runs at 55.9 tok/s, within 0.2 tok/s of the Stratified build (56.0). The speed gap closed.
Quality: not decided. The mean gap (0.46 points) is inside the confidence intervals (about ±3–5 points per task). Each model wins three tasks. Quality parity is the honest reading, not a win for either side.
Still missing for a full three-way verdict: V3 (exact Stratified boundary) is not built; Stratified has no n=200 run; no BF16 reference or KL-against-BF16 measurement exists for any model.