Author: Hakim Ghelab, VegaLaboratories LTD
Real, chronological log of every change made to this codebase. This is the short what; PORT_LEDGER.md has the detailed why, the real evidence behind each fix, and the full engineering narrative.
HESSIAN_PRIMARY_EXPLAINED.html section 11, built by
hessian_primary_page.py from 18 real solver runs at the
live coverage (six floor layouts: none, top 5% only, second tier at
5-bit, current 8/6, add a 5-bit layer to 30%, add a 5-bit layer to 50%;
times three plans: A2, C2, C3). Per run: Hessian danger deficit over
scored tensors, raw summed KL, tensors at 16-bit, mean bits of the 10%
most Hessian-critical and 10% most KL-critical tensors; a Pareto scatter
with frontier; per-plan “which layout does each signal prefer” and “what
the floors themselves do”; the rank correlation and critical-set overlap
between the two signals; a table of the unknowns and the benchmark that
would settle each.0.05:8,0.15:6,0.30:5) makes both
proxies worse on all three plans (deficit +48 to +55, raw KL +0.97 to
+1.77, 7 to 18 fewer 16-bit tensors) because it spends budget that would
otherwise reach 16-bit..page_cache/live_scores_build.json) so
every solver run and fact in that build sees the same data;
check notes when the live file has moved on since the
build.research_hadamard_blowup/hessian_primary_page.py +
HESSIAN_PRIMARY_EXPLAINED.template.html replace four
scripts and two files that lived only in a session scratch folder. One
script now builds, watches, controls
(start/status/logs/stop/restart)
and checks the page. The page is a Jinja2 rendering: every number, table
row and diagram is recomputed from the live Hessian score file and real
solver runs (plans A2, C2, C3), and sentences that depend on an outcome
switch wording when the measured outcome changes. Plan B2 (its exact
command is unknown) and the production plan stay frozen and are
labelled.hessian_primary_page.py check proves the page honest:
no unresolved markers, balanced tags, the facts embedded in the page
equal a fresh recomputation, the independent phase-1 re-solve equals the
real solver plans, and the frozen snapshot self-check;
--deep re-runs the real solver recipes on the snapshot.
Regression: a snapshot-mode build reproduces all 122 numbers of the
earlier verified page.research_hadamard_blowup/_orphaned/ (moved,
not deleted): refresh_hessian_primary_page_live.py,
run_hessian_primary_live_watcher.sh,
verify_hessian_primary_page_numbers.py, and the scratch
builder/body kept as reference. The background watcher uses the same pid
file and log, so hessian_primary_page.py restart 300
replaces the process already running.research_hadamard_blowup/refresh_hessian_primary_page_live.py
(--once, --watch --interval N,
--force): recomputes the live figures on
HESSIAN_PRIMARY_EXPLAINED.html from the live Hessian score
file (coverage, unscored composition, what plain primary and primary +
KL fallback would do right now, the worked example’s live rank, sweep
process status), first self-checking that its independent phase-1
re-solve still reproduces the saved snapshot plans C2 and C3 exactly
(live numbers are withheld if it does not). Rewrites only the marked
live box and the data-live figures, atomically, and only
when the score file, solver or sweep status changed.research_hadamard_blowup/run_hessian_primary_live_watcher.sh
(start [seconds], status, logs,
stop): runs it detached in the background with a PID file;
never touches any other process.HESSIAN_PRIMARY_EXPLAINED.html: new section “The exact
command, first” (the full recommended command, a plain-words table of
every line, and its real terminal output), and a “Keeping the live
numbers honest” subsection with the exact commands.RUNNING_GUIDE.md and
YAQA_COMMAND_REFERENCE.md: the complete fallback command as
a full copy-pasteable block (previously only described as an added
flag).--hessian-unscored-fallback kl: tensors with no Hessian
score no longer read as “safest” under
--hessian-primary02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py: new
opt-in flag --hessian-unscored-fallback {zero,kl} (default
zero = behaviour unchanged) and
compute_kl_fallback_weight(). With kl, each
tensor absent from the Hessian score file gets the percentile rank of
its own isolated KL at the lowest candidate bit, over all tensors, as
its --hessian-primary danger weight (1.0 = most dangerous,
same direction and i / (n - 1) formula as the Hessian
weight). Scored tensors and all floors are untouched. Plan JSON records
hessian_unscored_fallback.04_LANGUAGE_PLANS/HESSIAN_TEST/plan_c3_unscored_kl_fallback_score407_2026-09-19.json
and hess_scores_snapshot_407_for_C3_2026-09-19.json: the
real solver run and the 407-tensor score snapshot it used (new files;
nothing overwritten).HESSIAN_PRIMARY_EXPLAINED.html section 5.6, the C3 bar
in the unscored figure and the scoreboard, a C3 command in section 8;
122 numbers now checked by the verify script.audit_site_design.py: audits every page in the
PAGES manifest for the back-to-index pill (link and CSS),
the right “On this page” table of contents, the left scroll-progress
rail, the Cmd+K block, the index sidebar link, the index body card and a
written date; with --layout it loads each page in headless
Chrome at 1300/1500/1700/2000/2400px and fails on sideways scrolling or
any content overlapping a rail. Exit status 0 only when every page
passes.fix_page_rails.py: installs or repairs the rails and
the back pill on one hand-built page, measuring that page’s own
container width and setting its own breakpoint (container + 400px), with
--toc-only --bp N for pages that already have a left
sidebar.min-width:0.--hessian-primary versus the floor)research_hadamard_blowup/exports/HESSIAN_PRIMARY_EXPLAINED.html:
a written-from-zero mechanism page for the solver’s use of the Hessian
score. It separates the three numbers the code calls “danger” (raw
score, percentile rank, danger weight), walks the floor tiers and the
two-phase --hessian-primary solve line by line, and carries
ten diagrams and every step’s math with a plain-words line under each
equation.research_hadamard_blowup/verify_hessian_primary_page_numbers.py:
recomputes every number on that page by calling the solver’s own
compute_hessian_danger_weight and
compute_hessian_floor_map, re-solves phase 1 independently,
and with --check fails if any quoted number is missing from
the page (112 numbers checked).MILP_AWARE_HESSIAN_GUIDE.html,
HESSIAN_MASTER_REFERENCE.html,
HESSIAN_SCORE_ORIGIN.html,
DANGER_SCORE_GEOMETRY.html, RUNNING_GUIDE.md
§11.8 and YAQA_COMMAND_REFERENCE.md; site index card and
Cmd+K entry (manifest now 44).--hessian-primary the 68 tensors that reach
16-bit in plan C2 are exactly the 68 with the highest danger per million
parameters (lowest among them 0.0242, highest among the rest 0.0240).
The Phase-1 objective has no size term but the budget does, so the
solver buys protection where it is cheapest per unit of danger: a
nearly-safe 245,760-parameter tensor gets 16-bit while the single most
dangerous tensor stays at its 8-bit floor.layers.2,
layers.10 and layers.13
linear_attn.in_proj_qkv share the score 0.000283203125 but
receive floors of 8, 6 and 6. Reported, not changed.--pareto meaningful
(the default) silently blocks Hessian floor requirements; soft weight
confirmed to do nothingReal, direct debugging of a genuine
HiGHS Status 8: Infeasible failure on
02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py with
--hessian-primary + --hessian-floor-tiers
active. Root cause found by testing, not guessed:
--pareto meaningful (the real default) filters candidate
bit-widths using each tensor’s isolated-KL dominance – a filter
designed for a world where KL was the only signal, with no awareness
that a Hessian floor is about to demand a specific bit-width for a
completely different reason. Real, checked at 405/497 coverage: 55 of
497 tensors had their Hessian-floor-required bit filtered out, several
forced to jump to 16-bit instead of their real tier, and the resulting
collision with --hessian-primary’s own tight re-pinning
produced genuine infeasibility – worse as real coverage grows, not
better.
--pareto none whenever
--hessian-score-file drives allocation. Confirmed:
previously-failing real configurations (Plan A/B/C) all solve cleanly
once added.--hessian-weight-strength directly, isolatedly tested
(identical command, only that flag toggled between its default
2.0 and 0, --pareto none held
constant): bit-for-bit identical output. The soft
weight changes nothing once the Hessian floor and/or
--hessian-primary are active – it was only ever inflating
the reported OptiQ-weighted ΣKL statistic (up to 3× per
tensor), never the real allocation. Real fix: set it to
0..callout/.callout.red CSS
added to build_html_docs.sh’s shared pandoc template
(previously these docs had no styled warning box at all, just plain
text).RUNNING_GUIDE.html,
YAQA_COMMAND_REFERENCE.html, and every other doc it builds
now get real on-page wayfinding for the first time. Breakpoint (1348px)
computed and verified (zero collisions, 300-4000px) for this template’s
real 980px container, not copied from another page.MILP_AWARE_HESSIAN_GUIDE.html,
RUNNING_GUIDE.md §11.8,
YAQA_COMMAND_REFERENCE.md,
IMPROVEMENT_LEDGER/08_HESSIAN_HYBRID_SOLVER_DEBRIEF.html.Real Metal
[metal::malloc] Resource limit (499000) exceeded crash
root-caused and fixed in yaqa_core.py’s
ldlq2hess_quantize() –
wq_full/scales_full/ biases_full
were never flushed by mx.eval(), only hatWr
was, letting a Metal command buffer’s resource count balloon past the
hard limit on large tensors.
Fixed two real multi-job bugs in
monitor.sh/generate_dashboard.py: a fixed
dashboard port shared across every simultaneously-watched job caused
real job-flipping (fixed via a per-job hashed port), and
find_pid() grabbed the first system-wide process match
instead of filtering on the real --resume-dir path (fixed).
Dashboards now show which job they belong to.
watch.sh – lists every real *.yaqa_resume
job, prompts interactively for selection, execs monitor.sh
on the resolved job. Replaces the deleted watch_probe.sh
and other scattered per-job launch scripts.08_extract_hessian_scores.py consolidated into one file
with three subcommands (extract / watch /
hybrid), absorbing and deleting
research_hadamard_blowup/08b_merge_live_hessian_scores.py
and 08c_build_hessian_hybrid_checkpoint.py.
hybrid merges a real godmode multi-bit sensitivity sweep
into the same sensitivities[bit] field the MILP solver
(02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py) already
reads, tagging each tensor cost_source: "godmode" or
"kl_fallback" – no solver code changes needed.hessian_story_lib.py.Five new pages built and integrated into the Cmd+K manifest and exports/index.html sidebar/cards: HESSIAN_MASTER_REFERENCE.html (start here), HESSIAN_SCORE_ORIGIN.html, HESSIAN_VS_KL_TENSOR_EXPLORER.html, YAQA_PAPER_VS_CODE.html, HESSIAN_HYBRID_CHECKPOINT_DESIGN.html.
Not yet done: the Hessian Hybrid checkpoint has never been built into a real model or benchmarked end to end – whether Hessian-driven allocation is actually competitive with cascaded-KL-driven allocation remains untested.
Full real narrative, start here: research_hadamard_blowup/exports/HESSIAN_MASTER_REFERENCE.html.
Real godmode multi-bit sensitivity sweep added
(--godmode-multi-bit-checkpoint --godmode-candidate-bits 3,4,5,6,8),
testing every candidate bit-width per tensor and writing
godmode_sensitivity_checkpoint.jsonl – closes the real
coverage gap left by tensors this project’s own protection rules kept at
full precision (and so never got a Hessian score from the normal
correction pass).
02_optimize_plan_HYBRID_OPTIQ_PARETO_v3_3.py: a real
per-tensor hard Hessian floor (--hessian-floor-tiers), and
a lexicographic Hessian-first solve (--hessian-primary) –
built after proving the existing soft Hessian weight saturates (even
500x stronger changes nothing for the most dangerous tensors).Full real narrative: research_hadamard_blowup/HESSIAN_STRENGTH_SWEEP_2026-09-14/HESSIAN_PROBE_COMPLETE_GUIDE.html and …/MILP_AWARE_HESSIAN_GUIDE.html. Raw sweep data + logs: SWEEP_SUMMARY.html.
Correction (2026-09-16, twice): this entry was first dated
“2026-09-09/10” in this changelog with no verification, which was wrong.
A second pass re-dated it from internal content clues (fix comments,
“REAL BUG FOUND” markers) – closer, but still wrong for several pages,
because a date mentioned in a page’s own body text is sometimes about
the data being analyzed, not the page’s own authorship. Final, real
dates below come from the actual Write tool-call timestamps
in this project’s own Claude Code session transcript – the ground truth,
not inferred from either source.
Real, live free-signal analysis of the Hessian’s own structure as a candidate proxy for expensive cascaded-KL sensitivity measurement – first pass at what would become the Sep 14-16 MILP-aware Hessian investigation above.
IMPROVEMENT_LEDGER/00-09, each its own real
page. Every date below is the page’s first-written date
(its first real Write call in this project’s own session
history), not a “last touched” date – a minor edit does not reset it,
only a real, substantial revision earns a separate “updated” note: - 00
· Absolute Zero Start Here – 2026-09-10. - 01
· Zero to Expert, Redone – 2026-09-10. - 02
· MTP-YAQA Log Walkthrough – 2026-09-10. - 03
· The Brainstorm – 2026-09-11. - 04
· The lm_head Q4 Incident – written 2026-09-12,
updated 2026-09-13. - 05
· lm_head Fix, Next Steps – written 2026-09-12,
updated 2026-09-13. - 06
· Trunk, Hidden State, Sidecar – 2026-09-13. - 07
· The Hessian Signal – written 2026-09-13, updated
2026-09-15. - 08
· Full Solver Debrief – written 2026-09-14, updated
2026-09-15. - 09
· First Real Benchmark Results – 2026-09-15.
The lm_head Q4 protection incident and its fix, the trunk/hidden-state/sidecar story, the Hessian signal discovery, the full hybrid-solver debrief, and the project’s first real, complete MMLU/GSM8K/IFEval/BFCL/HumanEval benchmark results – highest 5-test mean of the group, never worst on any single metric.
Real implementation of the plan below:
--correct-mtp/--mtp-bits/--mtp-group-size
flags added to 05_full_model_quantize.py, reusing the exact
same Hessian-collection mechanism
(YaqaCatcher/mx.custom_function vjp) already
proven on the trunk, applied to the MTP module’s one real
mlx_lm.models.qwen3_5.DecoderLayer instance – confirmed via
inject_mtp_support() that this is the same class the
trunk’s own layers are built from, so no new correction math was needed.
New forward-pass recipe verified directly against
mtp_patch.py’s own real _mtp_core(): trunk
hidden state + real next-token embedding -> MTP layer -> real
next-NEXT-token cross-entropy (MTP predicts one token further out than
the trunk).
test_mtp_ yaqa_smoke.py), not by review: (1)
model(inputs) returns real logits, not hidden states –
needed text_model.model(inputs), the inner decoder stack,
instead; (2) lin.module.weight – lin was
already the unwrapped nn.Linear (captured in
targets before wrapping), never gained a
.module attribute.All 7 MTP tensors corrected, every one passed the safety gate on the
first attempt (no damping search needed – easier to correct than some
trunk tensors were). Real Hessian-weighted error vs. naive:
q_proj ~850x better, k_proj ~185x,
v_proj ~226x, o_proj ~9,675x,
gate_proj ~3,075x, up_proj ~5,404x,
down_proj ~3,327x. Corrected sidecar has the identical
29-tensor structure as the naive one – only the 7 quantized tensors’
values changed. Full regression suite still green.
Not yet built into a full official output model or benchmarked end to end – the smoke test validated the mechanism against a symlinked-trunk test directory, not a real official build. Full plan and real evidence: research_hadamard_blowup/exports/MTP_YAQA_CORRECTION_PLAN.html.
Real production launch, reusing the existing fully-completed trunk
resume cache (363/363 tensors + lm_head already corrected in the prior
-v2-fp32 build) rather than recomputing it, so this run
only re-assembles the model and additionally runs the new
--correct-mtp step. The existing -v2-fp32
model is left completely untouched – the resume cache was copied under a
new name first, specifically so this first real run of the new code path
can’t corrupt the only known-good copy if something goes wrong
mid-write.
Exact commands:
OLD="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32"
NEW="/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32-mtpcorrected"
cp -a "${OLD}.yaqa_resume" "${NEW}.yaqa_resume"
cd /Users/hghelab/ai-employee-build/projects/MLX_OptiQ/qwen38-27b-heretic-ara/scripts/yaqa_port
nohup ./run_full_yaqa.sh "$NEW" --n-calibration 4 --incoherence none --correct-mtp \
> "${NEW}.yaqa_resume/orchestrator_top.log" 2>&1 &(--mtp-bits/--mtp-group-size deliberately
omitted – defaults to the plan’s own native MTP bit-width, so this run
is bit-for-bit naive-vs-corrected, not confounded by also changing the
target bit-width.)
Real gaps found and fixed while launching this, both real bugs not
previously exercised because no prior run had a live process outside
05_full_model_quantize.py itself at monitoring time: -
monitor.sh’s find_pid() only pattern-matched
05_full_model_quantize.py, so it reported “NOT RUNNING
(finished or crashed)” for the entire real duration of the
06_lm_head_gptq.py step even though that process was
genuinely alive (confirmed directly via ps) – fixed to
match both script names. - generate_dashboard.py’s
real_assembly_progress() had no milestone at all for the
new MTP-correction step – neither dashboard could show it happening.
Added a mtp_yaqa step keyed on
correct_mtp_sidecar()’s own real log markers
([MTP-YAQA] Wrote real YAQA-corrected sidecar to... /
[MTP-YAQA] Source declares no MTP tensors...), shown as
skipped (not a false failure) on any run that didn’t pass
--correct-mtp at all. Verified by actually running the
generator against the live resume dir and grepping the real output HTML
for the new label – present in both dashboard.html and
brand/live-status.html (they share the same render
function).
--reuse-cached-fallback skip
option for lm_headReal, direct user request during the launch above: lm_head’s
one-sided GPTQ damping search is deterministic against a fixed real
Hessian, so a prior run that already exhausted every damping ratio and
fell back to naive (method: "naive_fallback" in the
manifest) will fail identically on retry – wasting ~15-20 real minutes
to reach the same known result. Added
--reuse-cached-fallback to
06_lm_head_gptq.py’s existing resume-skip check (previously
only accepted a prior "gptq" success) so it also accepts a
matching cached naive_fallback, still checked against the
current plan’s real bits/group_size. Off by default.
Real bug found and fixed immediately after adding it:
run_full_yaqa.sh forwards its full EXTRA_ARGS
to every invocation, including 05_full_model_quantize.py
(trunk batches, --print-num-batches,
--assemble-from-resume) – which does not define this flag,
so passing it at all crashed every one of those calls with “unrecognized
arguments” before 06_lm_head_gptq.py ever got to use it.
Fixed by adding the same flag to
05_full_model_quantize.py’s argparse as a real, documented
no-op there.
cp -a into an
already-existing dirRe-running Step 1 of the launch procedure
(cp -a "${OLD}.yaqa_resume" "${NEW}.yaqa_resume") a second
time, after ${NEW}.yaqa_resume already existed from the
first run, silently nested the entire source directory one level inside
the destination instead of refreshing it – cp -a src dst
copies src’s contents into a new dst, but copies src
itself into dst as a subdirectory when dst already exists.
Real, measured effect: resume-cache size doubled from 16GB to 32GB with
no error or warning. Did not affect the running build (it reads the
correct top-level files regardless), but is real wasted disk space.
Documented in YAQA_COMMAND_REFERENCE.md’s Step
1 as a check-before-you-copy warning.
Real gap, found while reviewing today’s full benchmark
result: the YAQA build’s MTP speculative-decode draft head is
quantized with plain mx.quantize
(optiq.runtime. mtp_convert.preserve_mtp() ->
_quantize_mtp()), zero correction – confirmed 0/363
MTP-related entries in the real resume manifest. This is the same code
path the reference baseline uses too, so it looked at first like a
shared, non-differentiating limitation.
Real correction to that read: checked the reference
model’s own real build manifest (V5_BUILD_MANIFEST.json)
directly – "elapsed_build_s": 35.77 for the whole trunk,
"language_trunk_source": "original_source + exact plan predicate",
and zero Hessian/GPTQ code anywhere in
05_build_final_model_V5.py, the script that actually built
it. The reference trunk is naive too. So the reference is naive-trunk +
naive-sidecar (internally consistent); YAQA is corrected-trunk +
naive-sidecar (internally inconsistent). Today’s entire
YAQA-vs-reference benchmark was never a clean apples-to-apples
comparison – a real, plausible contributor to the D1/D3
speculative-decode degradation found earlier, since deeper depths depend
most on trunk/draft agreement and only one side of that pair moved.
Fix, not a workaround: extend real YAQA correction
to the MTP sidecar’s 7 real quantizable tensors, at the same bit width
the naive sidecar already uses (Q4, confirmed from the real build
manifest), so the comparison becomes naive-Q4-sidecar vs.
YAQA-corrected-Q4-sidecar – same tensors, same bits, only the rounding
method differs. Real feasibility check:
optiq.runtime.mtp.mtp_patch.inject_mtp_support() builds the
MTP module from mlx_lm.models.qwen3_5.DecoderLayer – the
exact same class the trunk’s own decoder layers use – so YAQA’s existing
Hessian-collection and correction machinery, already proven on 362 real
trunk DecoderLayer instances, should attach to this one
directly rather than needing new correction math.
Full plan, evidence, and design requirements: research_hadamard_blowup/exports/MTP_YAQA_CORRECTION_PLAN.html,
with the real trunk-vs-MTP-sidecar tensor-name cross-reference in MTP_TENSOR_MATCH_TABLE.html.
Implementation as a new flag on 05_full_model_quantize.py
(not a new standalone script) in progress.
build_html_docs.sh was
showing its own title twiceReal, visible bug, found by directly clicking through every
link in exports/index.html: every page
built through this project’s own build_html_docs.sh
rendered its title twice – once in pandoc’s styled title-block header,
once again immediately underneath as a plain, unstyled heading, same
text both times.
Real root cause: the script derives each doc’s title
with TITLE="$(grep -m1 '^# ' "$INPUT_MD" | ...)" – the
first # Heading line found anywhere in the file, which is
also that same doc’s real first body heading. Passing that to
pandoc as --metadata title="$TITLE" makes pandoc render it
once in the generated title-block header and pandoc separately
renders the body’s own identical heading as a normal
<h1> right after it – same text, twice, stacked on
the page. Confirmed present on 12 of the 13 real docs
reachable from exports/index.html (README, RUNNING_GUIDE, YAQA_COMMAND_REFERENCE, PORT_LEDGER, CHANGELOG, RESEARCH_Hadamard_Blowup,
Adaptive_Damping_Power,
BF16_Curvature_Bug_Plain_English,
YAQA_UMA_Hessian_Precision_Root_Cause,
GPTQ_LM_Head_Safety_Gate_Plan,
AFFINE_MODE_CODE_REVIEW,
GPTQ_MLX_Integration)
– every hand-authored HTML doc (no title-block-header,
e.g. this project’s dashboards and the RCA report above) was unaffected,
confirmed by direct check, not assumed.
build_html_docs.sh: added a real post-processing step
(same Python block that already unescapes real Mermaid code fences) that
strips the duplicate <h1> immediately following
</header>, but only when its text exactly matches the
title-block header’s own text – a body that genuinely opens with a
different first heading is left untouched. Fixes every future
doc built with this tool by default, not just the 12 found today.OUT_DIR is
always yaqa_port/exports/; docs that belong in
research_hadamard_blowup/exports/ or
HybridOptiQ_FINAL_BENCHMARK/research/quantization_methods_2026/exports/
were built then mv’d there – no stray copies left in
yaqa_port/exports/, verified directly). .pdf
regenerated alongside each .html.BF16_CURVATURE_BUG_PLAIN_ENGLISH.md,
FIX_CHAT_GPT/YAQA_UMA_ROOT_CAUSE_FLOAT64_CURVATURE_FIX.md)
instead, so they survive this and every future rebuild rather than being
silently overwritten by the next pandoc pass.Real, measured gap, after the mode: affine fix
above was already live: mtplx tune still showed
the YAQA build 21-55% slower than the reference across AR/D1/D2/D3, and
peak memory +17.2% (27.85GB vs 23.77GB). The mode fix did
not touch this – it’s a separate bug.
Real root cause #1, confirmed by reading every tensor’s real
safetensors header on disk (not estimated): every one of 727
.scales/.biases tensors in the YAQA build was
stored as float32 instead of bfloat16
– double the per-group metadata bytes MLX’s quantized-matmul kernel
reads on every forward pass.
bits/group_size/mode matched the
reference on all 1912 tensors (0 mismatches) – never a bit-width bug.
Traced to 5 save points across 05_full_model_quantize.py
and 06_lm_head_gptq.py: the correction math deliberately
runs in float32 for numerical precision (correct), but nothing ever
downcast the resulting scales/biases back to bf16 before saving, unlike
the reference pipeline which runs bf16 throughout.
Real root cause #2 (found while checking for
like-for-like mtplx tune parity): embed_tokens
isn’t in either model’s real plan file (it’s a lookup table, not a
YAQA/GPTQ correction target), but this script’s generic “outside the
plan” fallback branch was quantizing it anyway (4-bit, U32, 636MB) while
the reference leaves it uncompressed (BF16, 2.54GB) – a real confound
for any like-for-like comparison, not itself the speed cause.
05_full_model_quantize.py: all 5 real save points
(live-correction rotated path, live-correction unrotated path, the
--assemble-from-resume path – confirmed the one both real
production builds actually used, and both out-of-plan fallback paths)
now cast scales/biases to
mx.bfloat16 explicitly before the array is assigned/saved.
Packed uint32 weight codes are dtype-invariant and
untouched.06_lm_head_gptq.py: same fix at its one save
point.05_full_model_quantize.py: any
nn.Embedding-type leaf outside the plan (i.e.
embed_tokens) is now left at full precision instead of
falling into the generic fallback-bits branch, matching the reference
build exactly.09_patch_scales_bf16.py (new): standalone tool to apply
both fixes to an already-built model without recomputing the correction
– casts scales/biases and swaps embed_tokens for the real
unquantized source weight, writing to a new directory. Not the path used
for the current live model (see below), kept for future use. Verified
end-to-end on an isolated synthetic fixture before use.Rebuilt via --assemble-from-resume (26h correction not
re-run; already-corrected values just re-saved at the right width) with
--fallback-bits 4 added to match the reference’s fallback
convention:
/Users/hghelab/.local/share/uv/tools/mlx-optiq/bin/python3 -u 05_full_model_quantize.py \
--assemble-from-resume \
--resume-dir "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.yaqa_resume" \
--output "/Users/hghelab/.mtplx/models/Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32" \
--n-calibration 4 --incoherence none --fallback-bits 4
Prior broken build moved aside first (not deleted) to
Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32.OLD-before-fallback4-fix.
Verified against the real rebuilt model on disk, independent of the
manifest’s own self-reported numbers: - Tensor forensics re-run: 0/1910
tensors differ from the reference in dtype, bits, group_size, mode, or
byte size – BF16 and U32 totals match the reference to the byte. -
Weight-provenance check: dequantized the same real tensor 4 ways
(original source, naive quantization, resume-cache correction output,
live shipped model). Live matched the resume cache to bf16-rounding
precision (~0.1-0.6%) and differed substantially from naive (1.9-4.5%
absolute) – the real correction ran and its exact output shipped, not
naive quantization standing in for it. - Resume manifest analysis:
362/363 tensors used real YAQA correction (lm_head is the
one known naive_fallback); real Hessian-weighted error
reduction vs. naive ranged 97.49-100%, median 99.67%, zero tensors worse
than naive. - Re-benchmarked (mtplx tune): AR 17.889 vs
reference 17.914 (-0.14%, parity), D2 47.422 vs 47.931 (-1.06%, parity),
peak memory 23.765GB == reference exactly. D1 (-20.19%) and D3 (-19.93%)
still lag, with real acceptance-rate degradation at deeper speculative
depths (D3 depth-3: 94.74% -> 80.60%). No localized failure found in
the manifest to explain this – working hypothesis is that YAQA’s
correction and the reference’s own correction algorithm land on
different, individually-small residual errors that compound differently
through depth; not yet confirmed by an independent output-quality
benchmark.
Full detail, real evidence, and the full comprehensive tensor table: research_hadamard_blowup/exports/RCA_scales_f32_bug.html.
Real, measured gap (direct user finding via
mtplx tune on both models): the real production YAQA build
ran substantially slower than the reference
HybridPareto5bpw-V3.3_Stratified model at every real
benchmarked depth – AR 14.06 vs 17.91 tok/s (~27% slower), best-depth
28.83 vs 54.08 tok/s (~1.87x slower). Confirmed NOT caused by a
different plan: real per-tensor bit assignments matched almost exactly
(178/130/45/11 vs 178/130/44/11 tensors at 4/5/6/8-bit – off by one
tensor) and real total file size was nearly identical (20.22GB vs
20.19GB).
Real root cause, confirmed by direct comparison of both
models’ real config.json: every per-tensor entry
this script ever wrote to plan_config[name] only ever
carried {"bits", "group_size"} – the real
"mode" key (MLX’s own quantization-mode field) was never
included anywhere in 05_full_model_quantize.py, at either
the per-tensor or top-level quantization dict. The
reference model’s config declares mode: "affine" on all 363
real entries plus the top level; the YAQA build’s config had it on 0 of
364. Confirmed the underlying packed tensor bytes were never wrong –
every real mx.quantize()/to_quantized() call
in this codebase always used MLX’s own real default
(mode="affine", confirmed directly from
mlx.core.quantize’s real signature) – this was purely a
missing config declaration, not corrupted data.
config.json directly (backed up first as
config.json.bak-before-mode-fix): added
"mode": "affine" to all 364 real per-tensor entries plus
the top-level quantization.mode field. Verified the patched
model still lazy-loads correctly via
mlx_lm.utils.load().05_full_model_quantize.py
so no future run repeats this: all four places a per-tensor quantization
config dict gets built (the two direct
{"bits","group_size"} literals, the real plan-derived
cfg = full_plan[name], and the fallback-bits dict for
leaves outside the plan) now explicitly include
"mode": "affine", plus the top-level
config["quantization"]["mode"] is now set to match the
trusted reference builders’ own real structure.language_model.model.embed_tokens fell into the real
“outside the plan” fallback path and was quantized at the fallback
default (6-bit for this run) vs the reference’s 4-bit – real, but likely
smaller impact on measured decode speed since embedding lookup is a
per-token gather, not a full matmul (unlike lm_head, which
was confirmed identical – 6-bit – in both models).The real full run finished successfully:
Part 5c PASS, real output model saved to
Qwen3.8-27B-heretic-ara-YAQA-5bpw-v2-fp32 — 363/363 tensors
real-corrected (362 YAQA + 1 GPTQ), 134 kept at full precision by the
real plan, 6.015 real bits/weight, real vision (921MB) and MTP (299.7MB)
sidecars reattached. Confirmed the source BF16 model’s memory was
genuinely released (no process running, ~34.6GB real free RAM
afterward).
vocab.json,
merges.txt, preprocessor_config.json, and
video_preprocessor_config.json all exist in the real source
model but were silently absent from the real production output —
confirmed the cause by reading mlx_lm.utils.save()’s actual
source: its entire extra-file-copy logic is hardcoded to
["*.py", "generation_config.json"], and this project’s own
reattach_vision_and_mtp()
(08_gptq_apply_plan.py, shared with the OptiQ/HybridPareto
pipeline) only ever handled real .safetensors weight
sidecars — neither ever looked at what else the real source directory
actually contained. The missing preprocessor configs broke real vision
inference on the built model (confirmed: an MTPLX vision-chat request
returned a real HTTP 500). Real fix, genuinely dynamic
(not a hardcoded per-model filename list, since different model families
ship different real sidecar files):
reattach_vision_and_mtp() now scans the real source
directory’s own top-level files at the end of its run and copies over
anything not already written to the destination and not on a small,
explicit exclusion list of known repo/download bookkeeping
(.DS_Store, .gitattributes,
crc32.txt, model-download.log,
resume-model-download.sh) — generalizes to whatever a
future model actually ships. Verified with a real isolated test (temp
source/dest dirs): correctly copies real missing sidecars, correctly
skips .safetensors/dotfiles/ excluded names, and
critically, never overwrites a file the pipeline already wrote. The four
missing files were also copied directly into the already-built real
output model so it works today, not just on the next run.Real, measured gap missing from this changelog until
2026-09-07 — this fix has full documentation (5 real docs:
root-cause research, plain-English writeup, the math/technical
companion, the lm_head GPTQ safety-gate companion fix, and
the adaptive-damping writeup, all in
research_hadamard_blowup/exports/) but never got a
changelog entry when it actually happened. Recorded here now, in its
real chronological place, not folded into a later entry.
Real root cause: production was constructing YAQA’s
Hessian curvature (the matrix the two-sided correction is built on) in
bfloat16 instead of float32. A Hessian needs enough numerical precision
to stay genuinely positive-semidefinite — bf16’s ~3 decimal digits of
precision was not enough, and this broke that guarantee for at least one
real tensor (layers.14.linear_attn.in_proj_qkv), which the
correction made measurably worse than naive quantization
instead of better. Found via a full Define/Hypothesize/Research/Test/
Measure/Validate investigation that first ruled out Hadamard rotation
and calibration size as causes before finding the real one.
lm_head’s separate GPTQ correction path had the same
precision issue and no safety net — fixed with a real safety gate (falls
back to naive if the corrected result isn’t actually better) plus the
same precision fix; see research_hadamard_blowup/exports/GPTQ_LM_Head_Safety_Gate_Plan.html.Zero-assumed-knowledge companion explainer for this whole investigation: HESSIAN_ZERO_TO_EXPERT.html.
Not to be confused with the 2026-09-07 F32-scales/biases bug below — that one is about the output metadata being saved too wide (a speed bug); this one is about the input curvature computation being too narrow (a correction-quality bug). Different stage, different symptom, different fix. Cross-referenced directly in both docs.
07_tensor_health_check.py: real
per-tensor integrity auditor. Scans every real .safetensors
file already saved in a resume dir against the manifest — checks the
file exists, has the right keys, weight is real uint32
packed codes, scales/biases are all-finite (no
NaN/Inf), and shapes are consistent. Regression-tested
(test_regression.py test [16]) by injecting a real NaN into
scales and confirming it’s caught, not silently passed —
proves the checker actually works, not just that it runs. Live result at
time of writing: 109/109 real cached tensors clean.outreach/ folder: local-only
launch/outreach kit (X, LinkedIn, Reddit r/LocalLLaMA, Hacker News,
Hugging Face model card, Apple outreach email, YouTube/Facebook video
scripts, guerrilla marketing playbook, IP protection notes), rolled up
into LAUNCH_KIT.html. Every
factual claim in it is sourced from this project’s own verified work —
no invented numbers. Nothing has been posted anywhere; these are drafts.
outreach/IP_PROTECTION_NOTES.md flags a real, unresolved
decision point: consult a real IP attorney on patent timing
before any public launch, since public disclosure can affect
patent eligibility depending on jurisdiction.monitor.sh running stale in-memory
code: the customer-facing dashboard
(brand/live-status.html) stopped updating for ~4.5 hours
while the internal dashboard.html kept updating fine. Root
cause: a long-running bash process doesn’t re-read its own script file —
monitor.sh had been running since before the
live-status.html generation logic was added to it, so it
never picked up the change. Fixed by killing and relaunching the
process; confirmed genuinely cycling again (timestamps advancing across
repeated checks, not a one-off)..pillars grid was hardcoded to repeat(3,1fr),
orphaning the 4th card in both the Timeline and Impact sections (which
have 4 cards each) onto its own mostly-empty row. Changed to
repeat(auto-fit,minmax(240px,1fr)) so the grid sizes itself
to however many cards actually exist. Also fixed inconsistent section
padding (Timeline section had 0 top padding, the only section on the
page missing it) and an overexposed Three.js gem (point-light
intensities were ~10x too high for the scale, blowing every facet to
white except one).IcosahedronGeometry, real mint/cyan point lights), after
two earlier real bugs in the 2D version — a highlight line that rendered
as a solid white bar instead of a subtle glint, and a near-black
pavilion fill nearly invisible against the page background.plan_Qwen38_V3_3_5.028075122774708_BPW.json): 497 total
assignments, 134 marked bits: 16 (the plan’s own “don’t
quantize this” decision), 363 real remaining targets. Not a bug — the
final model will still have all 497 layers, 134 just pass through at
full precision by design.~/.mtplx/models/ — checked all 14 model folders
there, only one (an unrelated third-party download) has the optional
mtplx_runtime.json sidecar, confirming it’s not a
requirement for MTPLX to load a model.05_full_model_quantize.py):
model.leaf_modules() stops treating a wrapped position as a
leaf once it has a .module child, silently breaking the
restore step. Every real save before this fix was corrupted (wrong keys,
wrong bits, no real YAQA correction). Fixed by reusing the pre-wrap leaf
key list for both wrap and restore, with a hard runtime check that fails
loudly if it regresses. Regression-tested
(test_regression.py test [9], toy-model reproduction).mx.quantize non-idempotency
(yaqa_core.py): re-quantizing the assembled correction
discarded real accuracy (max diff 3.7e-3 on a synthetic test, not float
noise). ldlq2hess_quantize now returns the packed codes
captured directly from each block’s own quantize call during the sweep.
Regression-tested (test [4b]).--assemble-from-resume printed “0/N tensors reduce error”
even when every tensor really did, because the manifest write happened
before the real metrics were computed. Fixed by moving the write after
metric computation; verified with a synthetic manifest (no real model
load needed).monitor.sh required manually tracking the live
batch number: rewritten to auto-detect and follow the newest
*.log file in a given directory, switching automatically as
the run moves from batch to batch.yaqa_core.YaqaQuantizedLinear: real,
derived, end-to-end-verified rotation-aware deployment module
(mx.hadamard_transform + mx.quantized_matmul).
Forward-pass formula hand-derived from
rotate_weight/inverse_rotate_weight, verified
to 2.98e-07 against the unrotated ground truth, and to 3.58e-07 through
the full real rotate→correct→pack→forward pipeline (test [8]). Not yet
loadable by plain mlx_lm.load() – research branch, not the
production default.--incoherence {none,hadamard} flag:
none is the production default (unrotated, real
mlx_lm-loadable, matches
08_gptq_apply_plan.py’s save mechanism exactly);
hadamard is the verified research branch above.--hessian-batch-budget-gb splits real targets into batches
(default 8GB) instead of wrapping all 363 simultaneously (557GB
requirement, confirmed by direct calculation, real reference needs
multi-GPU FSDP + CPU offload to do the same thing that this project
cannot replicate on one Mac).lm_head exclusion: its 248,320-wide
output means its own H_O alone needs 246.7GB. Excluded from
YAQA’s two-sided correction (falls to plain native
mx.quantize at its real plan bits instead), matching the
real reference’s own precedent of not running its per-layer correction
loop on this tensor either.run_full_yaqa.sh: real orchestrator. A
second real crash appeared after the memory fix –
RuntimeError: [metal::malloc] Resource limit (499000) exceeded
when running multiple batches inside one process.
mx.clear_cache() was tried and directly confirmed NOT to
fix it (re-tested, same crash). Real fix: each batch now runs as its own
separate OS process, writing results to a resume-cache directory; a
final --assemble-from-resume pass loads every batch’s
result and performs the real save.--resume-dir, --batch-only N,
--print-num-batches,
--assemble-from-resume: new CLI plumbing on
05_full_model_quantize.py supporting the orchestrator
above.--hessian-batch-budget-gb correctly
skips every tensor already completed (checked by name against the real
manifest, not by batch index), added after real elevated memory/swap
pressure during the live full run made a mid-run budget change worth
doing safely.hessian_llama/get_hess_llama.py): it also holds every
layer simultaneously in one pass, made feasible there by multi-GPU
sharding this project doesn’t have – so full single-pass collection was
never achievable here regardless of porting fidelity, but arbitrary
(non-layer-aligned) batching was a real, separate, fixable mistake.interpret_manifest.py: real-time
interpretation of a run’s manifest.json – explains Metric A
(Hessian-weighted error, NOT KL divergence) vs. Metric B (plain
Frobenius) in plain terms, breaks down real progress by tensor kind and
bit-width.requirements.txt: pinned real
dependency versions for a clean install elsewhere.mx.custom_function/vjp capture mechanism (test
[10]). Directly answers whether batching corrupts cross-tensor
measurements: no.mlx_lm.load() and run through
a real forward pass – correctly predicted “Paris” for “The capital of
France is.”--n-calibration 4, N=24 stratified calibration) launched
and progressing with no crashes since the fixes above landed.06_lm_head_gptq.py: real, scoped
one-sided GPTQ correction for language_model.lm_head, the
one tensor YAQA’s two-sided correction structurally cannot reach. Direct
question from the user (“if the OOM fix works for batches, why not
lm_head too?”) led to the real distinction: batching fixes “too many
tensors together,” not “one tensor’s own Hessian is too big” –
lm_head’s H_O alone is 246.7GB, even in a
batch of one. GPTQ only ever needs the small input-side Hessian, so it
was never going to hit that wall – confirmed directly that
08_gptq_apply_plan.py already corrects lm_head
with zero special-casing. Reusing that function’s own
gptq_quantize_with_plan() as-is was checked and rejected
(wraps every real Linear/SwitchLinear at once, needs 125.93GB – too
close to this machine’s 128GB). Wraps only lm_head, reuses
Apple’s own real Catcher class, packs via
mx.quantize for manifest-format consistency with the
YAQA-corrected tensors. Streamlined into run_full_yaqa.sh
as an automatic step (not a manual one to remember), made
resume-aware.--batch-only 21 out of range (0..20)). No real work was
lost. Fixed by always requesting the next undone chunk and re-querying
the real remaining count each iteration until it hits zero.--resume-dir, caught immediately by checking the
printed count rather than trusting the fix — would have hit the same
crash again once all real work was truly done.--source, --plan,
--lm-head-name: real CLI overrides replacing every
remaining hardcoded, model-specific constant. This codebase is no longer
Qwen3.5/3.8-specific.YaqaQuantizedLinear (needed before
--incoherence hadamard is actually runnable).--lm-head-name auto-detection was a reflexive
naming guess (.lm_head suffix match only) –
correctly flagged as not genuinely universal. Real fix: two independent
signals cross-checked – (1) a real vocab_size found by recursively
searching the model’s config at ANY nesting depth (confirmed necessary:
this model’s real vocab_size lives under
config['text_config'], not the top level), matched against
which plan candidate’s real output shape equals it; (2) the
.lm_head naming convention, used only as a cross-check or
lower-confidence fallback. REAL AMBIGUITY FOUND AND RESOLVED: shape
alone collides with a weight-tied embed_tokens (identical
shape, different module type) – resolved because the real plan’s own
candidate names never include embeddings (verified against the actual
497-entry plan). Raises loudly with both signals’ real findings shown if
detection is ever ambiguous, never guesses silently. Regression-tested
(test_regression.py test [11], the exact real collision
above).<pre class="mermaid"><code>ESCAPED</code></pre>);
real Mermaid 10.9.1 does not decode those entities first, so real
<br/>/quote characters inside labels broke its parser
– rendering a real “Syntax error in text” SVG that still has a real
<svg> tag and data-processed="true", so
a shallow check for either falsely reports success.--print-to-pdf snapshots the page
before Mermaid’s own async render finishes, regardless of
--virtual-time-budget (tried up to 8000ms) – produced a
completely blank diagram in the PDF specifically, while the HTML
rendered correctly live.htmlLabels:true renders label text inside a
<foreignObject> – displays and screenshots correctly,
but Chrome’s print pipeline does not render foreignObject
content at all, so the PDF was still blank. Isolated directly
(screenshot vs. print of the identical file); fixed with
flowchart.htmlLabels:false, which forces real SVG
<text>/<tspan> labels that print
correctly (real consequence: node labels use literal newlines for line
breaks now, not <br/>, which only works inside real
HTML content).bake_mermaid.py: the one real, shared,
reusable tool that renders a real diagram once (headless Chrome +
vendored vendor/mermaid.min.js) and bakes the result into
the shipped file as static SVG – no live JS dependency for anyone
opening the file, no CDN, no sidecar folder. Verifies real, positive
text content (not just “an <svg> tag exists” – that
check already once passed on a page that was rendering a visible error)
before writing anything. Used by both the markdown pipeline below AND
directly on hand-authored HTML files.build_html_docs.sh: real, reusable
HTML+PDF export for this project’s markdown docs, replacing the ad-hoc
one-off pandoc/Chrome commands used earlier this session. pandoc ->
unescape real mermaid blocks -> bake_mermaid.py ->
print PDF from the now fully-static file.vendor/mermaid.min.js: one real,
vendored copy inside this project (previously borrowed ad hoc from an
unrelated project’s browser-saved sidecar folder).| File | Status |
|---|---|
| scripts/yaqa_port/PORT_LEDGER.md | Fixed (diagram source uses real newlines, not
<br/>) |
scripts/yaqa_port/exports/PORT_LEDGER.html +
.pdf |
Fixed and visually verified (both HTML and printed PDF page) |
| 00_DOCS/HTML/yaqa_batching_architecture.html | Fixed in place (real diagram baked in,
.eyebrow/.callout .tag fonts raised from
11.5px/10.5px to 13px, below this project’s stated 12-13px floor) |
HybridOptiQ_FINAL_BENCHMARK/.../YAQA-to-MLX Port Ledger_dark.html |
Still broken, not yet touched – a stray browser-saved copy, not a canonical source; should be regenerated from the fixed exports/PORT_LEDGER.html rather than hand-patched separately |
| README.md, RUNNING_GUIDE.md, YAQA_COMMAND_REFERENCE.md | Never affected – confirmed zero Mermaid blocks in any of them |
Full detail in PORT_LEDGER.md’s Step
1-4 and early Step 5 sections. Summary: the gradient-capture mechanism,
Sketch B’s real Hessian formula, the LDLQ_2hess
block-sequential rounding algorithm, and the complete pipeline on one
real layer were each independently built and verified (Steps 1-4). Step
5 generalized this to many real tensors at once, found and fixed a
Hessian batch-accumulation formula bug, a missing Hadamard rotation
stage, a diagonal-zeroing bug, and an RNG-reuse bug across two rounds of
independent code review, before this session’s Part 5c work above.
© 2026 Hakim Ghelab, VegaLaboratories LTD. All rights reserved.