← Back to index ← Back to research index
Draft — not live
Improvement Ledger · Draft · not yet promoted to the live site · 2026-09-10

Your real MTP-YAQA build log, decoded line by line

This is the exact real terminal output from your ./run_full_yaqa.sh "$NEW" --n-calibration 4 --incoherence none --correct-mtp --reuse-cached-fallback run, broken into its real sections, each one explained in plain language directly underneath — including the one question you asked directly: why does the log show a naive MTP build before the YAQA correction. Read 01_ZERO_TO_EXPERT_REDO.html first if you haven't — this page reuses "effective rank," "H_I/H_O," and "Hessian" without redefining them.

0. Your exact question, answered first

Not a bug, not a contradiction

Your log really does show two steps for the MTP sidecar: a naive build, then a YAQA correction that overwrites it. That's deliberate, and it's spelled out in the actual code (05_full_model_quantize.py, lines 1597–1636):

Step A (runs always, reused code): reattach_vision_and_mtp() attaches the MTP sidecar using — in the code's own words — "the exact same real, already-trusted mechanism 08_gptq_apply_plan.py uses". That mechanism is a plain, naive mx.quantize at whatever bit-width the plan says (Q4/G64 here). This is old, already-tested code being reused, not new logic — it exists independently of YAQA and was never meant to be the final answer on its own.

Step B (runs only when you pass --correct-mtp): correct_mtp_sidecar() then reloads the model fresh from the file Step A just wrote to disk — the code's own comment explains why: "the honest, unambiguous way to guarantee MTP is calibrated against exactly what will actually ship, not an in-memory object." It then runs the real YAQA correction and overwrites the same optiq/mtp.safetensors file with the corrected version.

So: the naive file is real, but it's a scaffold that exists for a few seconds mid-run. The file that's actually still on disk when the script finishes is the YAQA-corrected one. Your intended design — trunk YAQA-corrected, MTP also YAQA-corrected — is exactly what you got. The log just shows you the real two-step mechanics of building it, which happen to reuse an older naive-build step as the scaffold.

1. The header — what you actually launched

hghelab@MacBook-Pro-2 yaqa_port % ./run_full_yaqa.sh "$NEW" --n-calibration 4 --incoherence none --correct-mtp --reuse-cached-fallback === run_full_yaqa.sh === output: .../Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected resume: .../Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected.yaqa_resume
Plain language

run_full_yaqa.sh is the one command that runs this project's entire correction pipeline end to end. Each flag is a real, separate choice:

  • --n-calibration 4 — use 4 real text samples per "domain" (code, prose, agent conversation, etc.) when building calibration data for this run's own steps. (The Hessian data itself was already built earlier, separately, at N=24 — this flag controls calibration for this run's own fresh measurements, like the MTP correction later in the log.)
  • --incoherence none — don't apply a Hadamard rotation to the weights before quantizing. (Rotation is a separate technique this project tested and is not using for the live model — see RESEARCH_Hadamard_Blowup.html for why.)
  • --correct-mtp — the flag from Section 0 above: run the real YAQA correction pass on the MTP sidecar, not just the trunk.
  • --reuse-cached-fallback — if a tensor was already corrected in a previous run and cached, use that cached result instead of recomputing it from scratch.

resume is a folder where every real tensor's correction result gets saved as it's computed, so a multi-hour run can be safely stopped and picked back up without redoing finished work.

2. Loading the plan — what "bits=X, group_size=64" means

Real plan loaded from .../plan_Qwen38_V3_3_5.028075122774708_BPW.json 363 real tensors need real quantization (bits != 16) -- FULL real plan language_model.lm_head: bits=6, group_size=64 language_model.model.layers.63.mlp.up_proj: bits=8, group_size=64 language_model.model.layers.62.mlp.up_proj: bits=4, group_size=64
Plain language

The "plan" is the real bit-allocation JSON produced earlier by the MILP/Pareto solver (the tool that decided which tensor gets how many bits, from the isolated-KL sensitivity sweep). bits=4 means: every weight number in that tensor gets compressed down to 4 bits instead of the original 16 (BF16). bits=16 (not shown in this excerpt, but present for 134 tensors) means: leave this tensor at full original precision, don't compress it at all — reserved for the most sensitive tensors.

group_size=64: quantization doesn't give every single weight number its own personal scale factor — that would need too much extra storage. Instead, weights are chopped into groups of 64 numbers, and each group of 64 shares one scale/bias pair (the "decoder key" from 01_ZERO_TO_EXPERT_REDO.html, Section 3). Smaller groups are more accurate but cost more storage for the extra scale/bias numbers; 64 is this project's chosen real trade-off.

363 tensors need real compression out of the full model — the rest (embeddings, norms, and a few protected tensors) stay at bits=16.

3. Splitting into batches — why 42, and why it matters

362 real targets split into 42 real batch(es) (budget=8.0GB/batch) -- 42 separate real forward+backward passes over the calibration data will run, one per batch.
Plain language

Correcting every one of the 362 real tensors (lm_head is handled separately, hence 362 not 363) needs the model to actually run — a real forward pass — on real calibration text, to see how each tensor actually behaves. Doing this for the whole ~27-billion-parameter model at once would need more memory than the machine has. So the tensors are split into 42 smaller groups ("batches"), each one sized to fit in an 8GB memory budget, and each batch gets its own real forward pass. This is purely a memory-management step — it doesn't change the math, just how much of it happens at once.

4. Trunk correction result — the summary line

Part 5b: 363/363 real tensors corrected successfully, 0 failed (363/363 ran without rotation -- no valid Hadamard factor). 362/363 tensors: YAQA reduces the real trace-based error vs. naive.
Plain language

All 363 tensors that needed correction got one, and none of them crashed or produced invalid (NaN/Inf) numbers. "Ran without rotation" means the Hadamard-rotation option from the launch flags wasn't used (matches --incoherence none above). The second line is the real headline result: for 362 out of 363 tensors, YAQA's corrected version genuinely has less real, Hessian-weighted error than plain naive quantization would have — only 1 tensor (the known exception, lm_head, excluded from two-sided correction for memory reasons and handled by its own separate GPTQ pass instead) isn't counted in that 362.

5. Attaching the MTP sidecar — Step A, the naive scaffold

MTP: native rule (matches this project's established V5/V5-HB baseline) -- Q4/G64 [mtp] Found 15 MTP tensors in source across 1 shard(s). [mtp] Extracted 15 tensors, 424,699,392 params. [mtp] Quantizing projections at 4-bit (group_size=64); norms + mtp.fc.weight stay at BF16. [mtp] Wrote optiq/mtp.safetensors (299.7 MB).
Plain language

This is Step A from Section 0 — the naive scaffold. The MTP ("multi-token prediction") module is a small, separate extra piece of the model that guesses several tokens ahead at once, to speed up generation. It has 15 real tensor pieces. This step just compresses those 15 tensors down to 4-bit with plain, uncorrected quantization (the norms and one combining layer, mtp.fc.weight, are left at full BF16 — they're small and not part of what gets YAQA-corrected). This creates a real, valid, loadable 299.7MB file — but it's about to be overwritten in Step B below.

6. Building real calibration data specifically for MTP correction

Building a real stratified calibration batch for MTP correction (k_per_domain=4, seq_len=128)... [stratified] 24 windows across 6 domains (4 target per domain): agent 56744 tokens -> 4 windows code 9168 tokens -> 4 windows instruct 1283 tokens -> 4 windows prose 1417 tokens -> 4 windows thought 9595 tokens -> 4 windows tool 9498 tokens -> 4 windows
Plain language

Before Step B can measure how sensitive the MTP tensors really are, it needs real text to run through the model. "Stratified" means the text is deliberately pulled evenly from 6 different real kinds of writing this model actually sees (agent transcripts, code, instructions, prose, reasoning/"thought," and tool calls) — 4 separate 128-token windows from each of the 6 domains, 24 windows total. This matters because a calibration set that's accidentally all one kind of text (e.g. only code) would measure sensitivity that doesn't represent the model's real, varied use — this is the same "N=24 stratified" data you've been asking about elsewhere in this investigation, just built fresh here specifically for MTP.

7. Step B — the real YAQA correction pass, one tensor decoded completely

This is the densest, most jargon-heavy part of the log. Below is the complete, real block for one tensor — mtp.layers.0.self_attn.q_proj — with every field explained against the actual safety_gate() code in yaqa_core.py.

mtp.layers.0.self_attn.q_proj: real effective rank -- H_I=1.7/5120, H_O=2.9/12288 mtp.layers.0.self_attn.q_proj/Hin: block_LDL succeeded on attempt 1, final sigma_reg=1.0000 mtp.layers.0.self_attn.q_proj/Hout(sigma_O=1.0): block_LDL succeeded on attempt 1, final sigma_reg=1.0000 mtp.layers.0.self_attn.q_proj: sigma_O=1e+00 weighted=0.0041x naive frob=1.019x naive SAFE=True
Line 1 — effective rank

Exactly the same effective-rank idea from 01_ZERO_TO_EXPERT_REDO.html Section 2, computed fresh for this MTP tensor: H_I's real risk concentrates in about 1.7 effective directions out of 5,120 possible; H_O in about 2.9 out of 12,288. Both close to 1 — a narrow, concentrated weak spot, same interpretation as before.

Lines 2–3 — block_LDL and sigma_reg

Before YAQA can use a Hessian matrix to compute a correction, that matrix has to be numerically "factored" in a stable way — the real algorithm used is called block LDL decomposition. Raw Hessians can be near-singular (numerically unstable to work with directly), so a small amount of regularization (sigma_reg, starting at 1.0) is added to stabilize the math before factoring. "Succeeded on attempt 1" means the gentlest regularization already worked — the code would automatically try progressively stronger values only if attempt 1 failed, and it didn't need to here.

Line 4 — the real safety-gate summary

Straight from yaqa_core.py's own safety_gate() function, checked against 4 real, independent conditions before any corrected tensor is accepted:

ConditionWhat it checksThis tensor
1. finiteNo NaN/Infinity in the corrected weightsTrue
2. beats_weightedReal Hessian-weighted error is lower than naive quantization'sweighted=0.0041× naive → True
3. within_frobRaw (unweighted) magnitude of the change stays ≤1.25× naive's own raw change — a sanity check so condition 2 can't be gamed by a numerically extreme correctionfrob=1.019× naive → True (1.019 ≤ 1.25)
4. within_magnitudeNo single corrected weight balloons past 2× the largest original weightnot printed in the short line, but confirmed True below

weighted=0.0041x naive is the standout number: this tensor's real, Hessian-weighted rounding error after YAQA correction is only 0.41% of what naive rounding would have produced — a ~244x reduction in real, curvature-weighted error, for this one tensor. SAFE=True means all 4 conditions passed, so this corrected version is the one that actually gets used.

[MTP-YAQA] self_attn.q_proj: {'effective_rank_in': 1.693, 'effective_rank_out': 2.935, 'trials': [{'finite': True, 'weighted_err': 1.5066e-06, 'naive_weighted_err': 0.0003672, 'beats_weighted': True, 'frob': 11.9165, 'naive_frob': 11.6912, 'within_frob': True, 'max_abs_hatW': 0.4791, 'max_abs_W': 0.4746, 'within_magnitude': True, 'safe': True, 'sigma_O': 1.0}]}
The full real record — the exact same 4 checks, with their raw numbers

This is the same tensor's complete safety-gate record, hand-verifiable:

  • weighted_err (YAQA's real error) = 0.0000015066, vs naive_weighted_err (plain rounding's real error) = 0.0003672 — YAQA's is smaller, so beats_weighted = True. (0.0000015066 / 0.0003672 = 0.0041, matching the "0.0041x naive" from the short line above.)
  • frob = 11.9165 vs naive_frob = 11.6912. Check: is 11.9165 ≤ 1.25 × 11.6912 (=14.614)? Yes → within_frob = True.
  • max_abs_hatW (largest weight after correction) = 0.4791 vs max_abs_W (largest original weight) = 0.4746. Check: is 0.4791 ≤ 2.0 × 0.4746 (=0.9492)? Yes → within_magnitude = True.
  • All 4 conditions true → safe = True, exactly matching SAFE=True in the short summary line.

8. The final line — what actually shipped

[MTP-YAQA] Wrote real YAQA-corrected sidecar to .../optiq/mtp.safetensors (29 tensors, 7 real corrected).
Plain language

The file from Step A (Section 5) is overwritten here with the real, YAQA-corrected version. "29 tensors" is everything in the MTP module (the 7 correctable linear projections you saw decoded above, plus norms, the combining layer, and other small pieces that don't go through YAQA correction). "7 real corrected" are specifically self_attn.{q,k,v,o}_proj and mlp.{gate,up,down}_proj — the same 7 tensor types the trunk's own DecoderLayer uses, confirming the MTP module really is architecturally one more decoder layer, corrected the same way as all 362 trunk tensors above it.

Hakim Ghelab, VegaLaboratories LTD · improvement ledger draft · every log line above is real, unedited terminal output from a real run; every explanation traced to the actual source (05_full_model_quantize.py, yaqa_core.py) · not yet promoted to the live site