Subject model: Qwen3.8-27B-heretic-ara, quantized on Apple Silicon (MLX). This exists because recalling benchmark numbers from memory across a long session produced inconsistent, incomplete answers more than once. This page is the single source of truth going forward — pulled directly from the real log files in quality/, not from prose summaries.
All rows are the same base model, Qwen3.8-27B-heretic-ara, quantized differently — not different models. Same protocol throughout: MMLU n=114, GSM8K/IFEval/BFCL/HumanEval n=100, --reasoning on. Real log files cited below the table.
| Model | BPW | V3.3 floor | MMLU | GSM8K | IFEval strict prompt · instr | BFCL | HumanEval | Mean (5 core) | D3 speed |
|---|---|---|---|---|---|---|---|---|---|
| Stock5 OptiQ competition | 5.0128 | n/a | 88.6% | 99.0% | 86.0% · 89.0% instr | 94.0% | 92.0% | 91.92% | ~54.0 tok/s |
| flat5 MLX competition | ~5.0 | n/a | 90.4% | 99.0% | 82.0% · 87.7% instr | 93.0% | 94.0% | 91.68% | 38.46 tok/s |
| V6 ours | 5.0004 | predates V3.3 | 89.5% | 97.0% | 90.0% · 92.6% instr | 90.0% | 94.0% | 92.10% | 55.23 tok/s |
| v6matched ours | 5.0133 | OFF | 86.8% | 99.0% | 89.0% · 92.6% instr | 91.0% | 93.0% | 91.76% | 55.02 tok/s |
| v33floor ours | 5.0133 (=v6matched) | ON | 89.5% | 100.0% | 86.0% · 91.4% instr | 93.0% | 92.0% | 92.10% | 54.03 tok/s |
| v52 / 5.2bpw ours | 5.2003 (diff. budget) | ON | 89.5% | 98.0% | 87.0% · 91.4% instr | 94.0% | 95.0% | 92.70% | 53.7 tok/s |
Mean = average of MMLU, GSM8K, IFEval strict (prompt-level), BFCL, HumanEval. Not itself proof of anything — individual deltas between rows haven't cleared statistical significance. It's here so direction is visible at a glance, computed exactly, not eyeballed.
V6 55.23 > v6matched 55.02 > v33floor 54.03 > Stock5 OptiQ ~54.0 > v52 53.7 (single run) > flat5 MLX 38.46 tok/s. All 6 models measured, all real. The four floor/matched-budget variants (V6, v6matched, v33floor, v52) cluster within ~1.5 tok/s of each other, well ahead of flat5 MLX — the floor mechanism doesn't cost meaningful speed at this budget.
v6matched vs v33floor — same 5.0133 BPW, only the V3.3 early-QKV floor toggled. All 5 core tasks now complete. v33floor leads on MMLU (+2.6pp), GSM8K (+1.0pp), and BFCL (+2.0pp); trails on IFEval strict (−3.0pp) and HumanEval (−1.0pp). Mean(5 core): v33floor 92.10% vs v6matched 91.76%, a +0.34pp edge.
Computed two-proportion z-test on every individual metric: none clear statistical significance (all |z|<1.0, all p>0.3 — largest single-metric z was −1.00 on GSM8K, p=0.32). The direction is consistently positive for the floor across 3 of 5 metrics and the mean, but at n=100–114 per task this is not distinguishable from noise. Honest verdict: the V3.3 floor shows a small, consistent, but not statistically proven advantage at this sample size — a real signal worth keeping, not yet a proven win.
Both have the floor ON, but at different budgets — v52 is 5.2003 BPW, v33floor is 5.0133 BPW. Same mechanism, two separate data points, not one result counted twice.
quality/stock5_*_n{114,100}.log (full rerun, completed 2026-08-30 01:28, after an earlier disk-full crash) · quality/v6_*_n{114,100}.log · quality/v6matched_*_n{114,100}.log · quality/v33floor_*_n{114,100}.log · quality/v52_*_n{114,100}.log · quality/flat5naive_*.log (from repro_check_matched_flat_v52.sh). Full detail and narrative also in LEDGER.md, "MASTER BENCHMARK RECAP" section.
A second, separate quantization method from everything in the table above — not another Pareto/MILP/floor variant of the same rounding, but a different way of correcting the rounding error itself. Built to test for regressions against the established results above, not yet a finished candidate.
This section's build (naive MTP, V3.3 plan, 91.34%, last of 7) has been superseded by a later, improved build at the same BPW (~5.03) — a different, cascade-KL-based trunk plan (V4) with the MTP sidecar also YAQA-corrected. That build scored 92.14%, the highest of 6 real models compared. Note: the trunk plan and MTP status both changed, so this is not an isolated MTP-only comparison against this exact row — but the matched BPW means it's a fair comparison at the same compression level. See Improvement Ledger §09 for the full, real result.
The build referred to above is now in the table below as model B (Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6): YAQA-corrected trunk and YAQA-corrected MTP sidecar, stratified-cascade P1 plan with the lm_head Q6 floor, 5.0286 BPW in the plan (about 4.67 as actually built — see the caveat below), 92.14%. The naive-MTP row is now called model A. A larger n=200 run exists too, but on a different build (HESSIAN-PROBE), not on model B — see "C: a bigger sample" at the bottom of this section.
YAQA (Tseng, Sun & De Sa, Cornell; arXiv 2505.22988) — an adaptive-rounding correction method, and that is what the paper contributes: it corrects how each tensor gets rounded, using a curvature (Hessian) measurement of the end-to-end output error, so the rounding error costs less of the model's real output quality at a given bit-width (see "effective rank" below).
What is this project's own work, not the paper's: deciding which tensor gets how many bits — the Pareto filter, the MILP solve over Hessian-weighted sensitivity scores, the structural and boundary floors, and the cascaded stratified-calibration measurement — plus the port of YAQA's rounding correction to MLX affine quantization ("YAQA for MLX") and its extension to the MTP sidecar. Every other row in the master table uses this project's allocation with plain rounding; the YAQA rows add the paper's rounding correction on top of this project's allocation.
Stratified calibration (N=24) — the real text used to measure sensitivity and build corrections, deliberately pulled in equal amounts (4 windows each) from 6 real kinds of writing this model actually sees (agent, code, instruct, prose, thought, tool) — 24 windows total. Replaces an earlier "flat" calibration set that wasn't deliberately balanced this way.
Effective rank / bit-error correction — for each tensor, a real measurement of whether its rounding risk is concentrated in a handful of directions (dangerous to round carelessly) or spread across most of the tensor (safe). YAQA uses this measurement to correct the rounding itself, not just to choose a bit-width.
Trunk vs. MTP sidecar — the "trunk" is the main model (the row below has it YAQA-corrected). The "MTP sidecar" is a small, separate speculative-decoding module; in this build it's still naive (plain rounding, not yet YAQA-corrected) — that correction is real, separate, in-progress work, documented on this project's own research pages, not reflected in the quality numbers below. (Update 2026-10-01: that correction is now built and fully benchmarked — "model B" further down.)
Cascade (P0 → P1 → …) — this project's sensitivity measurement. Instead of resetting the KL measurement at every bit-width, the plan is held fixed while sensitivity is measured forward through the network, so the measurement at the end reflects the real KL of the whole plan. A new plan P1 is derived from that sensitivity checkpoint; if it differs meaningfully from P0, another cascade starts from P1. Most of the change happened at P1; later rounds mostly oscillated.
Boundary floor (lm_head ≥ Q6) — a hard minimum of 6 bits on layer 0, layer 63 and lm_head. Added after the cascade's side effect: once every earlier layer is already quantized, the measurement at the end of the network barely moves whichever bit-width lm_head gets, so a plan with no boundary floor let the MILP drop lm_head from 8-bit to 4-bit. With only regional (early/late 25%) floors and no boundary floor, the MILP instead bumped lm_head to BF16 — consistent with the literature that lm_head is critical.
The goal isn't a new competing bit-allocation heuristic — it's testing whether correcting the rounding itself, on top of the existing MILP/floor allocation work above, buys real quality headroom the floor mechanism alone can't reach. This build exists specifically to check for regressions against the established table above before any further investment: same base model, same benchmark protocol, so the row below is a direct, apples-to-apples comparison, not a separate claim.
| Model | BPW | V3.3 floor | MMLU | GSM8K | IFEval strict prompt · instr | BFCL | HumanEval | Mean (5 core) | D3 speed |
|---|---|---|---|---|---|---|---|---|---|
| YAQA trunk, naive MTP prototype | 5.0284 | ON | 87.7% | 98.0% | 87.0% · n/a instr | 92.0% | 92.0% | 91.34% | 43.30 tok/s |
| YAQA trunk, YAQA-corrected MTP model B · P1-LMHead-Q6 | 5.0286 plan · 4.67 built | ON (lm_head ≥ Q6) | 87.7% | 98.0% | 88.0% · 92.6% instr | 94.0% | 93.0% | 92.14% | 49.58 tok/s |
| B − A | −0.36 (built) | 0.0 | 0.0 | +1.0 | +2.0 | +1.0 | +0.80pp | +14.5% | |
| v33floor ours, reference row | 5.0133 | ON | 89.5% | 100.0% | 86.0% · 91.4% instr | 93.0% | 92.0% | 92.10% | 54.03 tok/s |
| Rank | Model | Mean (5 core) | BPW | D3 speed |
|---|---|---|---|---|
| 1 | v52 / 5.2bpw ours | 92.70% | 5.2003 | 53.7 tok/s |
| 2 | YAQA trunk, YAQA MTP model B | 92.14% | 5.0286 plan · 4.67 built | 49.58 tok/s |
| 3 (tie) | V6 ours | 92.10% | 5.0004 | 55.23 tok/s |
| 3 (tie) | v33floor ours | 92.10% | 5.0133 | 54.03 tok/s |
| 5 | Stock5 OptiQ competition | 91.92% | 5.0128 | ~54.0 tok/s |
| 6 | v6matched ours | 91.76% | 5.0133 | 55.02 tok/s |
| 7 | flat5 MLX competition | 91.68% | ~5.0 | 38.46 tok/s |
| 8 | YAQA trunk, naive MTP model A · prototype | 91.34% | 5.0284 | 43.30 tok/s |
Exact, computed from the real per-task numbers above — two-decimal means, no rounding during comparison. V6 and v33floor are a genuine exact tie at 92.10% (89.5+97.0+90.0+90.0+94.0 vs 89.5+100.0+86.0+93.0+92.0, both ÷5 = 92.10 to 2dp) — not an approximation collapsing two close-but-different numbers. YAQA ranks last on quality and speed here specifically because its MTP sidecar is still naive, not because trunk-level YAQA correction underperforms — see the caveat below.
At near-identical BPW (5.0284 vs 5.0133) and the same V3.3 floor: YAQA trunk trails v33floor on Mean(5 core) by −0.76pp (91.34% vs 92.10%), and is meaningfully slower on D3 speed (43.30 vs 54.03 tok/s). Neither gap should be read as YAQA underperforming the method above — the MTP sidecar here is still naive, and this project's own earlier real, isolated comparison found correcting that sidecar closes most of the D1/D3 speed gap on its own. This row is a checkpoint against regressions, not the finished result.
The second prototype — YAQA trunk, YAQA-corrected MTP (the sidecar correction referenced above, already built and speed-tested, quality suite not yet run): real D1/D2/D3 speed comparison against this same naive-MTP row is documented on MTP_YAQA_CORRECTION_PLAN.html — D1 +34.9%, D2 +9.0%, D3 +11.4% decode speed over this row, D1/D2 now faster than the reference model, not just less slow. No MMLU/GSM8K/IFEval/BFCL/HumanEval numbers for that variant yet — speed only, so far.
The paragraph above said the YAQA-corrected-MTP variant had speed numbers only. It now has the full suite (row "model B" in the table above; same protocol: MMLU n=114, GSM8K / IFEval / BFCL / HumanEval n=100, --reasoning). B beats A on the mean by +0.80pp (92.14% vs 91.34%), is +14.5% faster at D3 (49.58 vs 43.30 tok/s; D1 +13.9%, D2 +14.8%), and its D3 draft acceptance rises from 88.6% to 95.5%. It ranks 2nd of 8 by mean, just above V6 and v33floor (both 92.10%), below only v52 (92.70%, a larger 5.2 BPW budget).
Why B exists: it is the build rebuilt with the right floor protection, so lm_head cannot be degraded to Q4 by the cascade side effect (see "Boundary floor" in the glossary), with YAQA rounding correction on every non-lm_head tensor and a YAQA-corrected MTP sidecar (7 tensors, each individually beating naive rounding). A YAQA-corrected sidecar next to a YAQA-corrected trunk means trunk and draft head agree, which is what the higher acceptance reflects.
Caveat — B’s real size is not 5.0286 BPW. That figure is the plan’s nominal budget (recomputed here from the plan’s own per-tensor bits and parameter counts: 5.0286). The built model contains 32 tensors that the P1 plan marks BF16 but that are actually quantized (21 at 5-bit, 7 at 6-bit, 4 at 8-bit) — in every one of the 32 cases at exactly the bits they had in the earlier plan/build whose cache this build was seeded from, which looks like a stale-cache leak (the mechanism is not proven). They hold 933M parameters, 3.64% of the 497-tensor parameter mass. Computing the same nominal-BPW formula from model B’s own config gives 4.67 BPW, about 0.36 BPW (≈1.1 GB) leaner than the same-plan trunk of model A (5.0284). Cross-check from file sizes: model B’s weights are 19.08 GB vs 20.19 GB for the earlier same-plan YAQA trunk, consistent with the 0.36 BPW gap. So the A → B comparison is not at matched BPW: B is smaller and still scores higher. The “same BPW (~5.03)” wording in the 2026-09-28 update above refers to the plan’s figure.
Still not an isolated MTP-only comparison: between A and B the trunk plan also changed (V3.3 plan → stratified cascade P1 plan with the lm_head Q6 floor). The isolated MTP-only result — same trunk, same plan, only naive vs YAQA-corrected sidecar — is the one on MTP_YAQA_CORRECTION_PLAN.html (D1 +34.9%, D2 +9.0%, D3 +11.4%). At n=100 no single-metric difference clears statistical significance (each score carries a ±5–6pp 95% interval); the direction is non-negative on all 5 tasks and the mean. Full detail: Improvement Ledger §09.
| Model | BPW (nominal) | MMLU (n=171) | GSM8K | IFEval strict prompt · instr | BFCL | HumanEval (n=164) | Mean (5 core) |
|---|---|---|---|---|---|---|---|
| HESSIAN-PROBE rebuild n=200 | 5.0284 plan label · 5.00 built | 90.1% | 97.0% | 90.5% · 93.5% instr | 92.5% | 95.1% | 93.04% |
| GODMODE-v2 (pure Hessian allocation, YAQA) n=200 | 5.00 built | 87.7% | 98.0% | 89.5% · 92.6% instr | 92.0% | 95.7% | 92.58% |
Both models are YAQA-corrected. HESSIAN-PROBE takes its allocation from the cascaded P1 plan. GODMODE-v2 takes its allocation only from Hessian curvature data, with fixed lm_head and boundary floors.
Draft acceptance by depth: D1 99.2% (HESSIAN-PROBE) vs 96.3% (GODMODE-v2). D2 99.2% / 92.9% vs 98.4% / 94.4%. D3 is only reached by GODMODE-v2: 96.9% / 93.8% / 90.7%, because HESSIAN-PROBE's depth-3 acceptance falls to 81.5% and its tuner rejects it. Decode speed at D2 is 51.4 vs 51.6 tok/s; GODMODE-v2 reaches 55.9 tok/s at D3.
Coding: HumanEval 95.7% (GODMODE-v2) vs 95.1% (HESSIAN-PROBE). Tool calling: BFCL 92.0% vs 92.5%, one call out of 200. Mean 5 core: 92.58 vs 93.04, inside the confidence intervals.
The HESSIAN-PROBE build started as an intermediate step to force the pipeline to re-calculate every Hessian score and weighted error at every bit-width; building a model was a side effect. Its plan targets the same nominal budget as model B (5.0284 vs 5.0286 BPW), but a direct tensor-by-tensor comparison of the two built models shows they are not the same bit allocation: 176 tensors have a different quantization assignment, 146 tensors are explicitly 8-bit (vs 7 in model B), lm_head is left unquantized (vs Q6 in model B), and the model is about 0.99 GB (+5.2%) larger on disk. (The 5.0284 is a label inherited from the base plan: the probe script overrides tensors to 8-bit without recomputing it. Computed from the built model’s own config, PROBE is 5.00 nominal BPW vs 4.67 for model B, a gap that matches the 0.99 GB size difference.) The n=200 numbers describe that build and cannot be read as "model B at a bigger sample".
Sample sizes: n=200 for GSM8K, IFEval, BFCL; MMLU is capped at 171 and HumanEval at 164 by their test sets, so those two use every available question. Intervals are tighter than n=100 (±3.3pp HumanEval, ±3.7pp BFCL, ±4.5pp MMLU), but the 93.04% mean is not directly row-comparable with the n=100 rows: a different question subset and a different build.
Open runs: (1) the classic Stratified model at n=200, running now. (2) A BF16 reference at n=200, with KL against each model. (3) The clean cascaded model (P1-LMHead-Q6) at n=200. (4) The V3 build (Hessian middle network with the exact Stratified boundary), built and tested at n=200. GODMODE-v2 at n=200 is complete and shown above.
quality/yaqa_mmlu_n114.log · quality/yaqa_gsm8k_n100.log · quality/yaqa_ifeval_n100.log · quality/yaqa_bfcl_n100.log · quality/yaqa_humaneval_n100.log (model: Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32) · D3 speed from 04_LANGUAGE_PLANS/pareto5/YAQA_MTP_NAIVE_vs_YAQA_MTP_Corrected_runtime.json.
Update 2026-10-01: the files named above were later overwritten by other benchmark runs that used the same filenames. Model A's MMLU and GSM8K logs were recovered from git as quality/modelA_fp32_naiveMTP_{mmlu_n114,gsm8k_n100}.log; its IFEval, BFCL and HumanEval raw logs could not be recovered (scores kept from the table above). Model B's logs: quality/modelB_P1LMHead_Q6_{mmlu_n114,gsm8k_n100,ifeval_n100,bfcl_n100,humaneval_n100}.log (recovered from the 2026-09-25 snapshot). n=200: quality/yaqa_bfcl_n200.log, quality/yaqa_{mmlu,gsm8k,ifeval,humaneval}_n200_run2.log (HESSIAN-PROBE). Speed: model A D1/D2/D3 = 29.01 / 47.42 / 43.30 tok/s from pareto5/Hybrid_Pareto_STRATIFIED_vs_YAQA_runtime.json (the file named above now holds a different comparison); model B = 33.05 / 54.42 / 49.58 tok/s from pareto5/YAQA_MTP_corrected_PO_vs_YAQA-5bpw-v2-P1-LMHead-Q6_runtime.json. Details: quality/README_LOG_OVERWRITE.md.
Update 2026-10-03: every IFEval cell on this page now also shows instruction-level strict next to prompt-level strict (real numbers from each row's own raw log, measured all along but not previously shown). Model A's instruction-level number is marked n/a — its raw IFEval log is one of the ones not recoverable, noted above. Only prompt-level strict feeds the Mean column; no existing mean or ranking changed.