Improvement ledger · chapter 11

GODMODE-v1, GODMODE-v2 and the n=200 results

Covers 30 Sep to 3 Oct 2026: the pure-Hessian builds, their speed and acceptance, and the first n=200 comparison against HESSIAN-PROBE.

What was built

GODMODE-v1 (1–2 Oct): the solver was given Hessian curvature data only, with the lm_head and boundary rules of its plan. D3 acceptance was 84.5%, and it ran at 54.8 tok/s.

GODMODE-v2 (2 Oct): the same Hessian-only data, with a lm_head floor of 6 and a boundary floor of 16 on the first and last layers. It was verified uncapped at every depth, with no budget hits and passing quality. Acceptance: D1 95.8%; D2 97.6% and 91.3%; D3 97.6%, 95.1% and 91.9%. D3 speed 54.8 tok/s (3.04x) uncapped, 55.9 tok/s (3.11x) in the capped tune run.

n=200 comparison

Both models are YAQA-corrected. The difference is the allocation source: HESSIAN-PROBE comes from the cascaded P1 plan, GODMODE-v2 from Hessian curvature alone.

TaskHESSIAN-PROBEGODMODE-v2
MMLU (n=171)90.1%87.7%
GSM8K97.0%98.0%
IFEval strict, prompt (instruction)90.5% (93.5%)89.5% (92.6%)
BFCL92.5%92.0%
HumanEval (n=164)95.1%95.7%
Mean of 5 core93.0492.58

The mean gap of 0.46 is inside the confidence intervals. This is parity, not a win for either side. GODMODE-v2 is the only one of the two that holds acceptance at D3.

What this chapter does not settle

Files