Structural Floor Deep-Dive · Verified Against Real Plans · 2026-10-02

GODMODE vs. Stratified: tensor-by-tensor, why one is faster, and the exact command to test why

Both models use the identical solver (02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py, single-stage Q_global_weighted_KL) and the identical generic 100x structural-weight protection on layer 0, layer 63, and lm_head. What differs is which extra hard floors are stacked on top of that weight — and that difference, not the sensitivity data itself, is the leading explanation for the Stratified build's speed edge.

Every number on this page is real. Bits, isolated-KL costs, and floor settings are read directly from each model's own real solved plan (plan_godmode_ONLY.json, plan_Qwen38_V3_3_5.028075122774708_BPW.json), cross-checked byte-for-byte against each model's real deployed config.json quantization map. Both models' generation quality was independently verified with mtplx forge verify on the uncapped prompt suite — results in Section 1, not assumed.

1. Both models generate clean, complete output

Before anything else: neither model's speed advantage comes at the cost of broken generation. Verified directly, not inferred.

GODMODE-v1

ARquality_passed: true
D1quality_passed: true
D2quality_passed: true
D3quality_passed: true
finish_reasonsstop (every depth)
hit_token_budgetfalse (every depth)

HybridPareto5bpw‑V3.3_Stratified

ARquality_passed: true
D1quality_passed: true
D2quality_passed: true
D3quality_passed: true
finish_reasonsstop (every depth)
hit_token_budgetfalse (every depth)

Both ran against mtplx forge verify's default uncapped suite (long-code-uncapped, up to 10,000 tokens per case) — the same real test, same command shape, run independently against each model's own real resume directory.

2. Speed and acceptance, side by side

Real mtplx tune results, both measured on the same Mac (Apple M3 Max, 128GB), same day, same profile.

ModeGODMODE tok/sGODMODE acceptanceStratified tok/sStratified acceptance
AR18.19—18.19—
D139.5196.3%40.1598.4%
D252.4798.0%52.6097.6%
D3 (default)54.83–55.2084.5%56.04–56.0793.5%

The Stratified build wins D1 and D3 outright on both speed and acceptance — D2 is close either way. D3 is the mode MTPLX's own tuner selects by default for both models (verdict: "mtp_depth_wins", zero rejections, zero collapses on both), so the D3 gap is the one that actually matters in deployment.

3. The complete per-tensor table — layer 0, layer 63, lm_head

Every tensor the structural floors touch, both models, exact bits and exact isolated-KL cost at that bit — no ranges, no rounding. group_size=64, mode=affine, and effective_kl_weight=100.0 are identical in every row for both models, so they're omitted from the table rather than repeated 32 times.

TensorGODMODE bitsGODMODE isolated_klStratified bitsStratified isolated_klParams
L0 in_proj_a81.1697e-07160.0245,760
L0 in_proj_b87.3127e-07160.0245,760
L0 in_proj_qkv81.0828e-04160.052,428,800
L0 in_proj_z81.4380e-04160.031,457,280
L0 out_proj160.0160.031,457,280
L0 mlp.down_proj64.2735e-04160.089,128,960
L0 mlp.gate_proj62.2037e-05160.089,128,960
L0 mlp.up_proj62.2891e-05160.089,128,960
L63 k_proj83.3825e-06160.05,242,880
L63 o_proj62.3504e-05160.031,457,280
L63 q_proj65.4534e-0581.6342e-0462,914,560
L63 v_proj84.2745e-06160.05,242,880
L63 mlp.down_proj64.1892e-0489.7445e-0589,128,960
L63 mlp.gate_proj61.0380e-0488.8229e-0589,128,960
L63 mlp.up_proj66.6777e-0481.0966e-0489,128,960
lm_head160.065.2681e-041,271,398,400

At 16-bit, isolated_kl is always exactly 0.0 by construction — it's the lossless reference point every other bit-width is measured against, not evidence of "no cost."

4. The floor mechanisms behind that table

The generic 100x boundary/output weight is identical in both builds — it is not what differs. What differs is which extra hard-minimum floors are stacked on top of it.

FloorGODMODE settingStratified settingReal effect
early_qkv_flooroff — 0 bits, 0 tensorson — 6-bit min, 48 tensorsForces Q/K/V input-projections in layers 0–15 above their natural KL-optimal bit. Only applies to 3 of layer 0's 8 tensors.
late_attention_flooroff — 0 bits, 0 tensorson — 5-bit min, 76 tensorsSame idea for late attention. Covers 4 of layer 63's 7 tensors (k/o/q/v_proj).
boundary_flooron — 6-bit min, 15 tensorsnot present (older plan schema)GODMODE-only mechanism; Stratified's plan predates it and relies on the generic 100x weight alone for the tensors this would cover.
lm_head_flooron — 8-bit minnot presentThe one floor that actually decides an outcome here. GODMODE's floor forces lm_head above 8-bit; the 100x weight then pushes it the rest of the way to 16. Stratified has only the weight, no floor — and under the identical real sensitivity numbers (both checkpoints carry the same KL-fallback value for lm_head), it settles at 6-bit on its own.
max_low_bit_run (run guard)off (0)on (3)Never triggered in either build's real solve (run_guard_promotions: [] in both) — empirically had zero effect on either real outcome.
Reading this precisely

"Not present" never means "no bit value" — every Stratified tensor still landed on a real, specific bit, decided by the generic 100x weight plus that build's own real sensitivity data, with no extra floor needed. In every "no floor" case except lm_head, the weight alone was already enough to land at 16-bit. lm_head is the one case where the extra floor is the actual deciding factor, not just insurance.

5. The hypothesis on the table

Your read on why the Stratified build wins, quoted directly:

"I think the stratified strategy is a winning strategy because we know that if you send a signal and you break the first layer and you quantize them too much, you create massive noise. I think the winning strategy of the stratified build was to keep as much of the first layer unquantized as possible. […] LM head, which is the last layer, is kept at 6, but the signal that goes in there is clean. I wonder if it's not what makes it win the MTP game."

The per-tensor table in Section 3 is consistent with this: every one of layer 0's 8 tensors sits at full 16-bit in the Stratified build (zero quantization noise entering the stack), versus GODMODE's 6–8 bit layer 0. Layer 63 mostly follows the same pattern before lm_head. The open question the table can't answer on its own is whether that upstream cleanliness is what lets lm_head compress safely to 6-bit without hurting MTP agreement — or whether lm_head's own compression would do just as well sitting downstream of a noisier stack, and the two effects are independent. Section 6 is the real test for that.

6. The exact command to test it

A new GODMODE plan, same real GODMODE sweep checkpoint, same target BPW — only the early/late floors changed to match the Stratified build's real values, and the lm_head floor lowered to let the (identical, real, shared) sensitivity data decide on its own instead of forcing it. This isolates the floor settings as the variable, keeping the GODMODE sweep data itself unchanged — a real test of your hypothesis, not a guess.

# Step 1 — preview only, writes nothing real (recommended first) python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/prepare_godmode_build.py \ --checkpoint /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/MASTER_HESSIAN_CHECKPOINT_QWEN_3.8_27B/hessian_hybrid_checkpoint.json \ --target-bpw 5.028075122774708 \ --early-qkv-min-bits 6 \ --late-attn-min-bits 5 \ --lm-head-min-bits 4 \ --boundary-min-bits 6 \ --output-plan /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/HESSIAN_TEST/plan_godmode_STRATIFIED_FLOORS.json \ --output-resume /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-GODMODE-v2-stratified-floors.yaqa_resume \ --dry-run # Step 2 — once the dry-run's reported bit distribution and lm_head outcome look right, drop --dry-run to actually solve the plan and build the resume-seed folder python3 /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port/prepare_godmode_build.py \ --checkpoint /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/MASTER_HESSIAN_CHECKPOINT_QWEN_3.8_27B/hessian_hybrid_checkpoint.json \ --target-bpw 5.028075122774708 \ --early-qkv-min-bits 6 \ --late-attn-min-bits 5 \ --lm-head-min-bits 4 \ --boundary-min-bits 6 \ --output-plan /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/04_LANGUAGE_PLANS/HESSIAN_TEST/plan_godmode_STRATIFIED_FLOORS.json \ --output-resume /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-GODMODE-v2-stratified-floors.yaqa_resume # Step 3 — the real correction build that consumes the plan + seed above (same shape as the original GODMODE build; confirm the exact run_full_yaqa.sh invocation you used before re-running)
Why --lm-head-min-bits 4, not 0

4 is the lowest real candidate bit-width in this plan's own candidate_bits list ([4, 5, 6, 8, 16] — no 3-bit candidate here). Setting the floor to 4 is the closest real equivalent to "no floor at all," since the solver can never choose below its lowest candidate anyway.

What this does and doesn't isolate

This changes only the floor settings — the real GODMODE sweep sensitivity data, target BPW, and candidate bits are untouched. If lm_head lands at 6-bit under this command (expected, since its real sensitivity number is identical in both checkpoints) and the resulting build's D3 speed closes the gap with the Stratified build's 56 tok/s, that's real evidence the floor settings — not the underlying sweep data — were the speed driver. It does not yet test the quality side of your hypothesis (whether a cleaner layer 0 is what protects MTP agreement specifically) — that question is answered in section 7 (GODMODE-v2 at n=200, completed 3 October 2026).

7. Result — GODMODE-v2 at n=200 (completed 2026-10-03)

GODMODE-v2 (Hessian middle network, lm_head floor 6, boundary floor 16) was run through the full n=200 suite and compared with HESSIAN-PROBE (cascaded KL, YAQA-corrected), the same protocol and seeded question set.

TaskGODMODE-v2HESSIAN-PROBELeader
MMLU (n=171)87.7%90.1%HESSIAN-PROBE
GSM8K98.0%97.0%GODMODE-v2
IFEval strict, prompt (instruction)89.5% (92.6%)90.5% (93.5%)HESSIAN-PROBE
BFCL92.0%92.5%HESSIAN-PROBE
HumanEval (n=164)95.7%95.1%GODMODE-v2
Mean of 5 core92.5893.04HESSIAN-PROBE by 0.46
D3 decode speed (tuned)55.9 tok/s (3.11x)51.4 tok/s, D2 best (2.86x)GODMODE-v2
What this answers

Speed: confirmed. With the floor settings changed, the Hessian build reaches D3 and runs at 55.9 tok/s, within 0.2 tok/s of the Stratified build (56.0). The speed gap closed.

Quality: not decided. The mean gap (0.46 points) is inside the confidence intervals (about ±3–5 points per task). Each model wins three tasks. Quality parity is the honest reading, not a win for either side.

Still missing for a full three-way verdict: V3 (exact Stratified boundary) is not built; Stratified has no n=200 run; no BF16 reference or KL-against-BF16 measurement exists for any model.