Real data read directly from each model's actual safetensors header and config.json — reference, the current naive build, and the new YAQA-corrected build. No estimates.
· 2026-09-07Where each tensor sits, and its real quantization status — identical in all three builds compared here. Outlined = always BF16, by deliberate design (the "cyankiwi" policy). Purple = Q4, quantized — naive today, YAQA-corrected in the new build.
Hessian-weighted error, naive vs. YAQA-corrected, same 7 tensors, same Q4 bit-width. Log-scale improvement factor — real numbers from the smoke test run against the actual model.
Bar length is on a log scale so the 185× and 9,675× results are both visible on one chart.
Every real tensor, all three builds.
| Tensor | Role | Shape | Reference | Naive YAQA (live) | Corrected YAQA (new) |
|---|---|---|---|---|---|
| pre_fc_norm_hidden | Norm on trunk hidden state | [5120] | BF16 | BF16 | BF16 |
| pre_fc_norm_embedding | Norm on next-token embedding | [5120] | BF16 | BF16 | BF16 |
| fc | Combines the two normed inputs | [5120,10240] | BF16 | BF16 | BF16 |
| layers.0.input_layernorm | Pre-attention norm | [5120] | BF16 | BF16 | BF16 |
| self_attn.q_proj | Attention query projection | [12288,640] | Q4 | Q4 naive | Q4 YAQA |
| self_attn.k_proj | Attention key projection | [1024,640] | Q4 | Q4 naive | Q4 YAQA |
| self_attn.v_proj | Attention value projection | [1024,640] | Q4 | Q4 naive | Q4 YAQA |
| self_attn.o_proj | Attention output projection | [5120,768] | Q4 | Q4 naive | Q4 YAQA |
| self_attn.q_norm | Per-head query norm | [256] | BF16 | BF16 | BF16 |
| self_attn.k_norm | Per-head key norm | [256] | BF16 | BF16 | BF16 |
| post_attention_layernorm | Pre-MLP norm | [5120] | BF16 | BF16 | BF16 |
| mlp.gate_proj | MLP gate projection | [17408,640] | Q4 | Q4 naive | Q4 YAQA |
| mlp.up_proj | MLP up projection | [17408,640] | Q4 | Q4 naive | Q4 YAQA |
| mlp.down_proj | MLP down projection | [5120,2176] | Q4 | Q4 naive | Q4 YAQA |
| norm | Final norm before shared lm_head | [5120] | BF16 | BF16 | BF16 |
Every .scales/.biases pair rides alongside its .weight row above, same treatment — not listed separately. 15 tensors always BF16, 7 always quantized, identical set in all three builds, confirmed directly from each real safetensors header.
There was never a discrepancy in which tensors get quantized versus kept full precision — both the reference and the naive YAQA build declare the identical real config: {"bits": 4, "group_size": 64, "mode": "affine", "policy": "cyankiwi"}. Traced directly in optiq/runtime/mtp/mtp_patch.py::_quantize_mtp_module — "cyankiwi" is the real, named policy that excludes fc/pre_fc_norm_*/norm and applies one uniform bit-width to the 7 real projections. No per-tensor bit variation exists anywhere in this. The real gap was never which tensors are touched — it was how the 7 already-correctly-scoped Q4 tensors get rounded: naive in both the reference and the old YAQA build, YAQA-corrected in the new one.