← Back to index
MTP sidecar · real tensor match

Every real tensor, what it is, and what changed.

Real data read directly from each model's actual safetensors header and config.json — reference, the current naive build, and the new YAQA-corrected build. No estimates.

· 2026-09-07

The real MTP forward pass

Where each tensor sits, and its real quantization status — identical in all three builds compared here. Outlined = always BF16, by deliberate design (the "cyankiwi" policy). Purple = Q4, quantized — naive today, YAQA-corrected in the new build.

Trunk hidden state post-norm, from the already-quantized trunk Embedding of token i+1 shared with the trunk's own embed_tokens pre_fc_norm_hidden RMSNorm · [5120] · BF16 pre_fc_norm_embedding RMSNorm · [5120] · BF16 fc — concatenate + project Linear · [5120, 10240] · BF16 · excluded from quantization by policy ONE REAL DECODER LAYER mlx_lm.models.qwen3_5.DecoderLayer — the exact same class the trunk's own 64 layers are built from input_layernorm · [5120] · BF16 self-attention — 4 real quantized tensors q_proj [12288,640] Q4 k_proj [1024,640] Q4 v_proj [1024,640] Q4 o_proj [5120,768] Q4 q_norm / k_norm · [256] each · BF16 post_attention_layernorm · [5120] · BF16 MLP — 3 real quantized tensors gate_proj [17408,640] Q4 up_proj [17408,640] Q4 down_proj [5120,2176] Q4 7 tensors inside this layer are Q4 — naive today, YAQA-corrected in the new build. Everything else stays BF16, always. norm · final RMSNorm · [5120] · BF16 lm_head · shared with the trunk · tie_word_embeddings=False Predicted token i+2
Always BF16 — 15 tensors, identical in all 3 builds Q4 — naive today, YAQA-corrected in the new build (same 7 tensors) External / shared with the trunk

The real correction result

Hessian-weighted error, naive vs. YAQA-corrected, same 7 tensors, same Q4 bit-width. Log-scale improvement factor — real numbers from the smoke test run against the actual model.

7/7
Safety Gate
Every tensor passed on the first attempt
185×
Worst-case improvement
k_proj — still 185x better than naive
9,675×
Best-case improvement
o_proj — the largest single gap closed
o_proj
9,675×
better
up_proj
5,404×
better
down_proj
3,327×
better
gate_proj
3,075×
better
q_proj
850×
better
v_proj
226×
better
k_proj
185×
better

Bar length is on a log scale so the 185× and 9,675× results are both visible on one chart.

The exact table

Every real tensor, all three builds.

TensorRoleShapeReferenceNaive YAQA (live)Corrected YAQA (new)
pre_fc_norm_hiddenNorm on trunk hidden state[5120]BF16BF16BF16
pre_fc_norm_embeddingNorm on next-token embedding[5120]BF16BF16BF16
fcCombines the two normed inputs[5120,10240]BF16BF16BF16
layers.0.input_layernormPre-attention norm[5120]BF16BF16BF16
self_attn.q_projAttention query projection[12288,640]Q4Q4 naiveQ4 YAQA
self_attn.k_projAttention key projection[1024,640]Q4Q4 naiveQ4 YAQA
self_attn.v_projAttention value projection[1024,640]Q4Q4 naiveQ4 YAQA
self_attn.o_projAttention output projection[5120,768]Q4Q4 naiveQ4 YAQA
self_attn.q_normPer-head query norm[256]BF16BF16BF16
self_attn.k_normPer-head key norm[256]BF16BF16BF16
post_attention_layernormPre-MLP norm[5120]BF16BF16BF16
mlp.gate_projMLP gate projection[17408,640]Q4Q4 naiveQ4 YAQA
mlp.up_projMLP up projection[17408,640]Q4Q4 naiveQ4 YAQA
mlp.down_projMLP down projection[5120,2176]Q4Q4 naiveQ4 YAQA
normFinal norm before shared lm_head[5120]BF16BF16BF16

Every .scales/.biases pair rides alongside its .weight row above, same treatment — not listed separately. 15 tensors always BF16, 7 always quantized, identical set in all three builds, confirmed directly from each real safetensors header.

What "match the reference" actually means here

There was never a discrepancy in which tensors get quantized versus kept full precision — both the reference and the naive YAQA build declare the identical real config: {"bits": 4, "group_size": 64, "mode": "affine", "policy": "cyankiwi"}. Traced directly in optiq/runtime/mtp/mtp_patch.py::_quantize_mtp_module — "cyankiwi" is the real, named policy that excludes fc/pre_fc_norm_*/norm and applies one uniform bit-width to the 7 real projections. No per-tensor bit variation exists anywhere in this. The real gap was never which tensors are touched — it was how the 7 already-correctly-scoped Q4 tensors get rounded: naive in both the reference and the old YAQA build, YAQA-corrected in the new one.