The stratified-calibration, cascade-corrected build's first real score
Every number on this page is read directly from a real, completed eval log in
~/HybridOptiQ_FINAL_BENCHMARK/quality/ — nothing estimated. Deterministic decoding throughout:
temperature=0, confirmed directly from the eval harness's own source
(optiq/eval/decode.py): "Every benchmark used to build make_sampler(temp=0.0) ...
Greedy is the right call for a reproducible benchmark."
This is the real P1 cascaded checkpoint from the live stratified-calibration cascade, rebuilt with the boundary floor after the lm_head Q4 incident (§04) and the fix that followed (§05) — Hessian data not yet involved. It remains the best real benchmark on this site because the Hessian-aware solver work in §08 hasn't been benchmarked end-to-end yet.
What was actually benchmarked
Verified directly against the real plan file, not assumed from the model name.
| Property | Real value |
|---|---|
| Model | Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-P1-LMHead-Q6 |
| Plan | 04_LANGUAGE_PLANS/V4_CLEAN_STRATIFIED_CASCADE_LMHead_protected/plan_v4_round1.json |
| Calibration | Domain-stratified (not flat), the calibration this project moved to specifically because it measures real, per-domain sensitivity instead of one undifferentiated pass |
| Sensitivity signal | Real cascade, round 1 (P1) — each tensor measured against the already-quantized state of every earlier-decided tensor, not a pristine bf16 assumption |
| Rounding correction | Full YAQA two-sided Hessian correction on every non-lm_head tensor — the real rounding-error protection this whole project is built around |
lm_head | 6-bit, real hard floor (confirmed in plan: bits=6) |
| MTP sidecar | YAQA-corrected (not naive) — 7 real tensors, each individually confirmed beating naive rounding (beats_weighted: True). Source: this build's own assemble.log. |
| Achieved BPW | 5.0286 |
| Note | This build predates the per-tensor hard floor / lexicographic Hessian mechanism from §8 of the solver debrief — it is the cascade+floor baseline, not yet the newest solver output. |
Every real model, every real benchmark
All five real comparison models plus the naive flat-bit control, same real eval harness, same real temperature=0, same real sample sizes (MMLU n=114, everything else n=100).
| Model | MMLU | GSM8K | IFEval strict/loose | IFEval instr. strict/loose | BFCL | HumanEval | 5-test mean |
|---|---|---|---|---|---|---|---|
| YAQA (P1, this build) | 87.7% | 98.0% | 88.0 / 88.0 | 92.6 / 92.6 | 94.0% | 93.0% | 92.14% |
| v6 | 89.5% | 97.0% | 90.0 / 91.0 | 92.6 / 93.3 | 90.0% | 94.0% | 92.10% |
| v33floor | 89.5% | 100.0% | 86.0 / 87.0 | 91.4 / 92.0 | 93.0% | 92.0% | 92.10% |
| stock5 | 88.6% | 99.0% | 86.0 / 86.0 | 89.0 / 89.0 | 94.0% | 92.0% | 91.92% |
| v6matched | 86.8% | 99.0% | 89.0 / 90.0 | 92.6 / 93.3 | 91.0% | 93.0% | 91.76% |
| flat5-naive (no allocation logic) | 90.4% | 99.0% | 82.0 / 84.0 | 87.7 / 89.0 | 93.0% | 94.0% | 91.68% |
5-test mean = simple average of MMLU, GSM8K, IFEval-strict, BFCL, HumanEval. This build's real mean (92.14%) is the highest of the six real models compared — and it is never the single worst on any one metric (the two metrics it trails on, MMLU and GSM8K, both have a different model in last place: v6matched on MMLU, v6 on GSM8K).
What the real numbers actually show
The real question this benchmark answers: did the stratified calibration, the cascade correction, and the real YAQA rounding-error protection break anything? Real answer: no. This build's mean beats every other real model compared, and it is competitive with or ahead of the group on every individual metric except two, where it is 2nd-of-six, never last.
Real instruction-level score: 92.6% strict, 92.6% loose — identical. Two other real models share that same property (stock5 at 89.0%/89.0%), so it isn't unique to this build, but the level it holds at is: 92.6% strict exactly matches v6 and v6matched, the two best real instruction-followers in the whole comparison group. Real, structured instruction-following is exactly the kind of behavior naive quantization damages first (see below) — holding this level under strict grading is a genuine, positive signal.
flat5-naive — no sensitivity-aware allocation at all, just uniform bits — actually
scores the highest real MMLU (90.4%) of any model here. Raw factual recall barely notices naive
quantization. What it does notice: IFEval strict, real score 82.0% — the lowest of any model by a wide
margin (4–8 points below every allocated model). Structured, rule-following behavior is where careless
quantization damage actually shows up, and it's exactly the metric this build (88.0% strict) and its
instruction-level score (92.6%) hold up well on.
Two real next steps, not yet done
This is a real, complete first result — not the final word. Two things would make it a stronger one.
1. A larger benchmark, to see the real extent of its strengths
Every task here ran at n=100 (n=114 for MMLU) — enough for a real first read, not enough to fully resolve close calls (every real score above carries a real ±5–6pp 95% confidence interval). Real, exact command to run at a larger sample size:
2. A real BF16 (full precision) comparison, to measure the actual capability gap
The real, unquantized reference model already exists on disk (the same one used as this project's KL-truth
reference throughout): /Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara (52GB, confirmed
present). Running the exact same real suite against it directly measures how much real capability, if any,
this 5.03 BPW build gives up relative to full precision — not inferred, measured:
Real caveat: BF16 at 52GB will run meaningfully slower per-token than any of the quantized models above (no compute/memory savings) — feasible on this machine (128GB RAM, per the benchmark ledger's own hardware line), just a real, longer wall-clock run than any single row in the table above.
~/HybridOptiQ_FINAL_BENCHMARK/quality/ — none estimated, none illustrative.