← Back to index ← Back to research index
Improvement ledger 06 · trunk / hidden state / sidecar · 2026-09-13

Why changing 12 tensors out of 497 forced the whole MTP sidecar to be re-tuned from scratch

Your reading of what the code actually did is exactly right. This page exists to make the mechanism itself visible — not just described — so "the trunk's hidden state changed" stops being a phrase and becomes something you can watch happen.

Real structure · 9 of 497 language-model layers shown, representative not literal
unchanged trunk layer
tensor whose bit-width changed under the corrected plan
hidden state, flowing
New trunk, freshly re-tuned sidecar: recalibrated against this exact trunk’s real hidden states. Aligned, and the connection is clean.
Drag to rotate · watch the signal travel from the trunk into lm_head, then into the sidecar

01Three things this page has to define before anything else makes sense

No assumed knowledge. If any of the three words in the title were fuzzy, they won't be after this section.

☰The trunk

The main stack of transformer decoder layers — for this model, 497 real weight tensors across roughly 60 layers. Every token you feed the model passes UP through this stack, layer by layer, each layer transforming the token's internal representation a little further toward "the next word." This is the part everyone means when they say "the model."

≈The hidden state

At every layer, the token isn't a word anymore — it's a real vector of numbers (thousands of them), the model's running internal summary of "what this token means so far, given everything before it." Each layer reads the PREVIOUS layer's hidden state and writes a new one. It is a real, literal array of floating-point numbers — not a metaphor, an actual tensor you can print.

⚡The MTP sidecar

Multi-Token Prediction: one extra, smaller decoder layer, bolted on the side. While the trunk is busy computing the very next token, the sidecar takes the trunk's own hidden state and tries to guess the token AFTER that — one step further ahead. If its guess turns out right, that extra token is free: no separate trunk pass needed for it. This is what makes speculative decoding faster.

%Quantization, in one line

Every real weight starts as a BF16 number (about 8 million possible values). Quantizing to 4-bit rounds it to one of only 16 possible values. That rounding is a real, measurable error — small on any one weight, but it changes that layer's real output by a small, nonzero amount, every single time.

02Why a small change upstream becomes a real change downstream

This is the actual mechanism behind "the trunk's hidden state changed" — not an analogy, the real data path.

Step 1A tensor's bits changeA layer's weight matrix gets rounded to a different bit-width than before — say 4-bit instead of the original allocation. The rounding error on that one matrix is real and nonzero.
Step 2That layer's OUTPUT changesThe hidden state THIS layer hands to the next one is computed FROM the now-slightly-different weights. It is measurably different from what it would have been.
Step 3Every layer above inherits itLayer 4 has no idea layer 3's weights changed — it just receives whatever hidden state arrives. It computes on top of the new one, compounding forward, layer after layer, all the way to the top.
Step 4The sidecar receives the resultBy the time the hidden state reaches the MTP sidecar, it carries the accumulated effect of every real bit-width change beneath it — even ones several layers away.

In this project's own real incident: the corrected plan (built after fixing the lm_head protection bug documented in ledger 04) changed 12 of the model's 497 real language-model tensors compared to the plan the previous build used — 2.4% of them. Here is exactly which ones, and how their bits moved:

Real diff, this project's own two plans
language_model.model.layers.1.…changed→re-optimized under the fixed solver
language_model.model.layers.3.…changed→re-optimized under the fixed solver
language_model.lm_head4-bit→6-bit
…and 9 more, 12 total (2.4% of 497)—full real diff in ledger 04

Layers 1 and 3 sit near the very BOTTOM of the trunk — meaning their hidden-state effect had the entire rest of the stack, dozens of layers, to compound through before reaching the top. That is precisely why the code cannot assume "only 12 tensors changed, so the effect is small": it isn't the count that matters, it's where the changes sit in the stack.

03What "re-tuning the sidecar" actually measures

The sidecar's own correction isn't calibrated against the plan, or the bits, or any abstraction — it's calibrated against the real, literal hidden-state numbers the trunk hands it.

YAQA correction (for any tensor, trunk or sidecar) works by minimizing real error against a real Hessian measured from real forward/backward passes over real calibration data. For the sidecar specifically, that calibration data is the trunk's own real hidden states — not a fixed reference, not the plan, the ACTUAL numbers produced by running real tokens through the ACTUAL trunk that will ship. Change the trunk, even in 12 places, and the real numbers the sidecar was calibrated against are simply not the real numbers the new trunk produces anymore. A correction tuned against the old numbers is tuned against a target that no longer exists.

The real code, not paraphrased

From 05_full_model_quantize.py, immediately before the sidecar correction call — the script reloads the model it JUST finished saving to disk, specifically so calibration is measured against exactly what will ship, not an in-memory guess:

# MTP correction is a new, separate process that must run
# regardless of how the trunk was obtained.

That comment is the whole answer in one sentence: the sidecar's correction was never wired into the trunk's own resume-cache, on purpose, because doing so safely would require proving the trunk's real output hadn't moved — and the simplest, safest thing the code can do instead is just always remeasure it fresh.

04Your question, answered directly

Confirmed, precisely

Yes. The source BF16 model was identical. The overall bits-per-weight budget was identical. But the PLAN — the specific bits assigned to specific tensors — was a genuinely different allocation (the cascaded, lm_head-protected P1 plan, not the original flat plan). Because that allocation changed real tensors partway up the trunk, the trunk's real hidden states changed too, for every token, at every layer above the change.

The MTP sidecar's own prior correction — the one already sitting in the previous, non-cascaded build, already YAQA-corrected from naive — was calibrated against the OLD trunk's hidden states. It doesn't matter that it was already corrected, or that the source model never moved: it was tuned to a target (the old real hidden-state distribution) that the new trunk no longer produces. So irrespective of the sidecar's own prior correction status, the real forward/backward calibration had to be redone — not because the code failed to find a cache, but because no cache could have been valid here without first proving the trunk hadn't moved, which nothing in the pipeline currently checks for.

05What this run actually produced

Real numbers from this build's own completed log, not estimates.

QuantityReal value
Real MTP calibration batches run24 / 24
Real MTP-side tensors in the sidecar29
Of those, genuinely re-corrected (not naive)7
Sidecar quantization4-bit, group_size=64
Trunk tensors reused from the seeded prior build236 / 343 (69%)
Trunk tensors genuinely recomputed this run97–107 (final count in the real manifest)

The trunk's own per-tensor correction largely reused what the prior build had already done — 236 of 343 real targets matched by name, bits, group_size, and curvature version, and were skipped outright (see ledger 05 and RUNNING_GUIDE.md section 9 for the full mechanism). The sidecar could not participate in that same reuse for the structural reason above — and that real, unavoidable 24-batch recalibration is what re-tuned it to agree with the corrected trunk.

Real outcome you reported, pending exact figures

You've observed this freshly re-tuned build performing better on real inference metrics — faster acceptance rate (AR) and faster depth-1/depth-2 (D1/D2) speculative-decoding throughput — than the earlier non-cascaded, naive-then-YAQA-corrected sidecar build. That is the real, concrete payoff for the recalibration cost documented above: a sidecar that is actually in agreement with the trunk it ships next to, rather than one tuned against a trunk that no longer exists. Send over the exact AR/D1/D2 numbers from your benchmark run and this section gets a real comparison table, not a qualitative statement.

Author: Hakim Ghelab, VegaLaboratories LTD · Companion to IMPROVEMENT_LEDGER/04_LMHEAD_PROTECTION_INCIDENT.html and 05_LMHEAD_FIX_NEXT_STEPS.html · Every number on this page is either read directly from this build's own real log output or cited from the plan JSON — none estimated.