Why changing 12 tensors out of 497 forced the whole MTP sidecar to be re-tuned from scratch
Your reading of what the code actually did is exactly right. This page exists to make the mechanism itself visible — not just described — so "the trunk's hidden state changed" stops being a phrase and becomes something you can watch happen.
lm_head, then into the sidecar01Three things this page has to define before anything else makes sense
No assumed knowledge. If any of the three words in the title were fuzzy, they won't be after this section.
☰The trunk
The main stack of transformer decoder layers — for this model, 497 real weight tensors across roughly 60 layers. Every token you feed the model passes UP through this stack, layer by layer, each layer transforming the token's internal representation a little further toward "the next word." This is the part everyone means when they say "the model."
≈The hidden state
At every layer, the token isn't a word anymore — it's a real vector of numbers (thousands of them), the model's running internal summary of "what this token means so far, given everything before it." Each layer reads the PREVIOUS layer's hidden state and writes a new one. It is a real, literal array of floating-point numbers — not a metaphor, an actual tensor you can print.
⚡The MTP sidecar
Multi-Token Prediction: one extra, smaller decoder layer, bolted on the side. While the trunk is busy computing the very next token, the sidecar takes the trunk's own hidden state and tries to guess the token AFTER that — one step further ahead. If its guess turns out right, that extra token is free: no separate trunk pass needed for it. This is what makes speculative decoding faster.
%Quantization, in one line
Every real weight starts as a BF16 number (about 8 million possible values). Quantizing to 4-bit rounds it to one of only 16 possible values. That rounding is a real, measurable error — small on any one weight, but it changes that layer's real output by a small, nonzero amount, every single time.
02Why a small change upstream becomes a real change downstream
This is the actual mechanism behind "the trunk's hidden state changed" — not an analogy, the real data path.
In this project's own real incident: the corrected plan (built after fixing the lm_head protection bug documented in ledger 04) changed 12 of the model's 497 real language-model tensors compared to the plan the previous build used — 2.4% of them. Here is exactly which ones, and how their bits moved:
Layers 1 and 3 sit near the very BOTTOM of the trunk — meaning their hidden-state effect had the entire rest of the stack, dozens of layers, to compound through before reaching the top. That is precisely why the code cannot assume "only 12 tensors changed, so the effect is small": it isn't the count that matters, it's where the changes sit in the stack.
03What "re-tuning the sidecar" actually measures
The sidecar's own correction isn't calibrated against the plan, or the bits, or any abstraction — it's calibrated against the real, literal hidden-state numbers the trunk hands it.
YAQA correction (for any tensor, trunk or sidecar) works by minimizing real error against a real Hessian measured from real forward/backward passes over real calibration data. For the sidecar specifically, that calibration data is the trunk's own real hidden states — not a fixed reference, not the plan, the ACTUAL numbers produced by running real tokens through the ACTUAL trunk that will ship. Change the trunk, even in 12 places, and the real numbers the sidecar was calibrated against are simply not the real numbers the new trunk produces anymore. A correction tuned against the old numbers is tuned against a target that no longer exists.
From 05_full_model_quantize.py, immediately before the sidecar correction call — the script
reloads the model it JUST finished saving to disk, specifically so calibration is measured against exactly
what will ship, not an in-memory guess:
# MTP correction is a new, separate process that must run
# regardless of how the trunk was obtained.
That comment is the whole answer in one sentence: the sidecar's correction was never wired into the trunk's own resume-cache, on purpose, because doing so safely would require proving the trunk's real output hadn't moved — and the simplest, safest thing the code can do instead is just always remeasure it fresh.
04Your question, answered directly
Yes. The source BF16 model was identical. The overall bits-per-weight budget was identical. But the PLAN — the specific bits assigned to specific tensors — was a genuinely different allocation (the cascaded, lm_head-protected P1 plan, not the original flat plan). Because that allocation changed real tensors partway up the trunk, the trunk's real hidden states changed too, for every token, at every layer above the change.
The MTP sidecar's own prior correction — the one already sitting in the previous, non-cascaded build, already YAQA-corrected from naive — was calibrated against the OLD trunk's hidden states. It doesn't matter that it was already corrected, or that the source model never moved: it was tuned to a target (the old real hidden-state distribution) that the new trunk no longer produces. So irrespective of the sidecar's own prior correction status, the real forward/backward calibration had to be redone — not because the code failed to find a cache, but because no cache could have been valid here without first proving the trunk hadn't moved, which nothing in the pipeline currently checks for.
05What this run actually produced
Real numbers from this build's own completed log, not estimates.
| Quantity | Real value |
|---|---|
| Real MTP calibration batches run | 24 / 24 |
| Real MTP-side tensors in the sidecar | 29 |
| Of those, genuinely re-corrected (not naive) | 7 |
| Sidecar quantization | 4-bit, group_size=64 |
| Trunk tensors reused from the seeded prior build | 236 / 343 (69%) |
| Trunk tensors genuinely recomputed this run | 97–107 (final count in the real manifest) |
The trunk's own per-tensor correction largely reused what the prior build had already done — 236 of 343
real targets matched by name, bits, group_size, and curvature version, and were skipped outright (see
ledger 05 and RUNNING_GUIDE.md section 9 for
the full mechanism). The sidecar could not participate in that same reuse for the structural reason above — and
that real, unavoidable 24-batch recalibration is what re-tuned it to agree with the corrected trunk.
You've observed this freshly re-tuned build performing better on real inference metrics — faster acceptance rate (AR) and faster depth-1/depth-2 (D1/D2) speculative-decoding throughput — than the earlier non-cascaded, naive-then-YAQA-corrected sidecar build. That is the real, concrete payoff for the recalibration cost documented above: a sidecar that is actually in agreement with the trunk it ships next to, rather than one tuned against a trunk that no longer exists. Send over the exact AR/D1/D2 numbers from your benchmark run and this section gets a real comparison table, not a qualitative statement.