← Back to index
Improvement ledger 09 · the first real benchmark · 2026-09-15

The stratified-calibration, cascade-corrected build's first real score

Every number on this page is read directly from a real, completed eval log in ~/HybridOptiQ_FINAL_BENCHMARK/quality/ — nothing estimated. Deterministic decoding throughout: temperature=0, confirmed directly from the eval harness's own source (optiq/eval/decode.py): "Every benchmark used to build make_sampler(temp=0.0) ... Greedy is the right call for a reproducible benchmark."

What this page is not: a claim that the gap to full precision (BF16) is known, or that this is the largest sample size this model could be tested at. Both are real, open next steps — see the closing section.
Where this build came from (added 2026-09-17)

This is the real P1 cascaded checkpoint from the live stratified-calibration cascade, rebuilt with the boundary floor after the lm_head Q4 incident (§04) and the fix that followed (§05) — Hessian data not yet involved. It remains the best real benchmark on this site because the Hessian-aware solver work in §08 hasn't been benchmarked end-to-end yet.

THE REAL BUILD

What was actually benchmarked

Verified directly against the real plan file, not assumed from the model name.

PropertyReal value
ModelQwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6
Plan04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/plan_v4_round1.json
CalibrationDomain-stratified (not flat), the calibration this project moved to specifically because it measures real, per-domain sensitivity instead of one undifferentiated pass
Sensitivity signalReal cascade, round 1 (P1) — each tensor measured against the already-quantized state of every earlier-decided tensor, not a pristine bf16 assumption
Rounding correctionFull YAQA two-sided Hessian correction on every non-lm_head tensor — the real rounding-error protection this whole project is built around
lm_head6-bit, real hard floor (confirmed in plan: bits=6)
MTP sidecarYAQA-corrected (not naive) — 7 real tensors, each individually confirmed beating naive rounding (beats_weighted: True). Source: this build's own assemble.log.
Achieved BPW5.0286
NoteThis build predates the per-tensor hard floor / lexicographic Hessian mechanism from §8 of the solver debrief — it is the cascade+floor baseline, not yet the newest solver output.
FULL RESULTS

Every real model, every real benchmark

All five real comparison models plus the naive flat-bit control, same real eval harness, same real temperature=0, same real sample sizes (MMLU n=114, everything else n=100).

ModelMMLUGSM8KIFEval strict/looseIFEval instr. strict/looseBFCLHumanEval5-test mean
YAQA (P1, this build)87.7%98.0%88.0 / 88.092.6 / 92.694.0%93.0%92.14%
v689.5%97.0%90.0 / 91.092.6 / 93.390.0%94.0%92.10%
v33floor89.5%100.0%86.0 / 87.091.4 / 92.093.0%92.0%92.10%
stock588.6%99.0%86.0 / 86.089.0 / 89.094.0%92.0%91.92%
v6matched86.8%99.0%89.0 / 90.092.6 / 93.391.0%93.0%91.76%
flat5-naive (no allocation logic)90.4%99.0%82.0 / 84.087.7 / 89.093.0%94.0%91.68%

5-test mean = simple average of MMLU, GSM8K, IFEval-strict, BFCL, HumanEval. This build's real mean (92.14%) is the highest of the six real models compared — and it is never the single worst on any one metric (the two metrics it trails on, MMLU and GSM8K, both have a different model in last place: v6matched on MMLU, v6 on GSM8K).

READING THE RESULT

What the real numbers actually show

No regression, real evidence

The real question this benchmark answers: did the stratified calibration, the cascade correction, and the real YAQA rounding-error protection break anything? Real answer: no. This build's mean beats every other real model compared, and it is competitive with or ahead of the group on every individual metric except two, where it is 2nd-of-six, never last.

Instruction-following holds up under strict grading

Real instruction-level score: 92.6% strict, 92.6% loose — identical. Two other real models share that same property (stock5 at 89.0%/89.0%), so it isn't unique to this build, but the level it holds at is: 92.6% strict exactly matches v6 and v6matched, the two best real instruction-followers in the whole comparison group. Real, structured instruction-following is exactly the kind of behavior naive quantization damages first (see below) — holding this level under strict grading is a genuine, positive signal.

The naive baseline shows what "no real allocation" actually costs

flat5-naive — no sensitivity-aware allocation at all, just uniform bits — actually scores the highest real MMLU (90.4%) of any model here. Raw factual recall barely notices naive quantization. What it does notice: IFEval strict, real score 82.0% — the lowest of any model by a wide margin (4–8 points below every allocated model). Structured, rule-following behavior is where careless quantization damage actually shows up, and it's exactly the metric this build (88.0% strict) and its instruction-level score (92.6%) hold up well on.

WHAT'S STILL OPEN

Two real next steps, not yet done

This is a real, complete first result — not the final word. Two things would make it a stronger one.

1. A larger benchmark, to see the real extent of its strengths

Every task here ran at n=100 (n=114 for MMLU) — enough for a real first read, not enough to fully resolve close calls (every real score above carries a real ±5–6pp 95% confidence interval). Real, exact command to run at a larger sample size:

optiq eval /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6 \ --task all --score --n-samples 500 \ --baseline <reference_model_path>

2. A real BF16 (full precision) comparison, to measure the actual capability gap

The real, unquantized reference model already exists on disk (the same one used as this project's KL-truth reference throughout): /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara (52GB, confirmed present). Running the exact same real suite against it directly measures how much real capability, if any, this 5.03 BPW build gives up relative to full precision — not inferred, measured:

optiq eval /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara \ --task all --score \ --baseline /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6

Real caveat: BF16 at 52GB will run meaningfully slower per-token than any of the quantized models above (no compute/memory savings) — feasible on this machine (128GB RAM, per the benchmark ledger's own hardware line), just a real, longer wall-clock run than any single row in the table above.

Author: Hakim Ghelab, VegaLaboratories LTD · Companion to IMPROVEMENT_LEDGER/07, 08 · Every score on this page is read directly from a real, completed log in ~/HybridOptiQ_FINAL_BENCHMARK/quality/ — none estimated, none illustrative.