2026-10-01
Who this page is for: you have a BF16 model and want a quantized MLX build, but you don’t have (and don’t want to build) this project’s own real per-tensor sensitivity plan — the stratified-calibration → KL-sensitivity → MILP-allocation pipeline that normally decides which bit width every individual tensor gets. Flat mode skips all of that. You pick one bit width, and every real trunk tensor in the model gets it — except a small, automatically-protected set of tensors that get a real hard floor instead, so quality-critical layers are never silently flattened down to your chosen width.
This is the primary path for anyone outside this
project — nobody else has this project’s own private
Hessian-sweep data, so --plan (a full custom multi-bit
allocation) only works for someone who has run that whole separate
pipeline themselves. Flat mode needs nothing but a source model.
Flat mode is implemented by one shared module,
flat_plan.py, called from two otherwise-independent build
scripts — YAQA (05_full_model_quantize.py)
and GPTQ (08_gptq_apply_plan.py). Both
take the exact same --flat-* flags and produce the exact
same kind of plan; only the actual correction math that gets applied
afterward differs (YAQA’s two-sided Hessian correction vs. GPTQ’s
one-sided correction — see README.md for why this project
built both).
Given a source model and a bit width, Flat mode:
nn.Linear/SwitchLinear tensor — the same
leaf-type set the build scripts actually quantize).--flat-n-protect real transformer layer indices (default:
just layer 0 and the very last layer).--flat-bits.
Boundary and lm_head tensors get
max(--flat-bits, their own floor) — so asking for a
higher --flat-bits than the floor never silently
downgrades a protected tensor; it only ever raises it.plan.json to disk (same schema every
other plan file in this project already uses) and points the build
script’s own --plan at it — every existing, already-trusted
call site in either script then runs completely unchanged. Flat mode
never touches the actual correction logic, only which bit width gets
requested.Why boundary gets a real hard floor at all: in a
normal MILP-built plan, boundary protection comes from a soft cost
weighting the solver can freely trade against a real token budget. Flat
mode has no solver and no budget to trade against — it’s one direct
assignment — so it can’t reuse “no hard floor” as if it meant “no
protection needed.” The default floor (6-bit) is the same one already
given to lm_head, justified by real data from this
project’s own Hessian sweep: the last boundary layer contains the single
highest-danger tensor of all 497 scored in the model.
| Flag | Default | Meaning |
|---|---|---|
--flat-bits N |
(none — required to turn Flat mode on) | The one bit width every trunk tensor gets. Must be one of
4, 5, 6, 8 — the
exact set MLX’s native quantizer and this project’s GPTQ packing both
support. |
--flat-group-size N |
64 |
Quantization group size applied to every tensor under Flat mode. |
--flat-boundary-min-bits N |
6 |
Hard floor for boundary tensors (see above). A higher
--flat-bits can raise this; a lower one never lowers
it. |
--flat-lm-head-min-bits N |
6 |
Hard floor for lm_head, same logic. |
--flat-n-protect N |
1 |
How many layer indices at each end (first and
last) count as “boundary” — not how many tensors. 1 means
every tensor belonging to layer 0 and every tensor belonging to the
final layer, not one representative tensor per layer. |
--output PATH |
(none — required) | Where the build is written. Also required by Flat mode specifically,
because the synthesized plan.json is placed next to it
(Path(args.output).parent). Both scripts raise a clear
error (ap.error(...)) if you give --flat-bits
without --output. |
--plan and --flat-bits are mutually
exclusive in spirit: if you pass both, --flat-bits wins and
its synthesized plan silently replaces whatever --plan
pointed at.
On this project’s own 64-layer hybrid model, –flat-n-protect
1 protects every tensor in layer 0 and layer 63
— all 15 of them, not one tensor named “layer 0”:
Layer 0 (8 tensors -- a GatedDeltaNet layer):
language_model.model.layers.0.linear_attn.in_proj_a
language_model.model.layers.0.linear_attn.in_proj_b
language_model.model.layers.0.linear_attn.in_proj_qkv
language_model.model.layers.0.linear_attn.in_proj_z
language_model.model.layers.0.linear_attn.out_proj
language_model.model.layers.0.mlp.down_proj
language_model.model.layers.0.mlp.gate_proj
language_model.model.layers.0.mlp.up_proj
Layer 63 (7 tensors -- a standard attention layer):
language_model.model.layers.63.mlp.down_proj
language_model.model.layers.63.mlp.gate_proj
language_model.model.layers.63.mlp.up_proj
language_model.model.layers.63.self_attn.k_proj
language_model.model.layers.63.self_attn.o_proj
language_model.model.layers.63.self_attn.q_proj
language_model.model.layers.63.self_attn.v_proj
All 15 of those get max(–flat-bits,
–flat-boundary-min-bits) — exactly matching the real
[flat-mode] Synthesized plan: … 15 boundary tensor(s) floored to
>= Q6 line from an actual tested run of this project. Layer
counts differ by architecture (GatedDeltaNet layers like layer 0 have no
self_attn at all, which is why it’s 8 tensors instead of 7)
— run your own model through mtplx inspect or check your
plan file to see its real per-layer names.
08_gptq_apply_plan.py (GPTQ only) also has a separate
–calibration-mode flag with a flat choice
(stratified is the default). That controls how
calibration text gets sampled — nothing to do with bit-width
assignment. Flat mode (this page,
–flat-bits) picks the quantization bit width with no plan
file. Flat calibration (–calibration-mode
flat) picks how calibration text is drawn, and is a completely
separate axis — you can combine either with either.
05_full_model_quantize.py (YAQA) has no
–calibration-mode flag at all; it always uses stratified
calibration.
Quick test first (3 example tensors, a couple of minutes, needs your own real BF16 model):
"$PYBIN" 05_full_model_quantize.py --source /Users/yourname/Downloads/my-model \
--flat-bits 5 --subset default --output /Users/yourname/Downloads/my-model-quicktestWhat success looks like: a line like
[flat-mode] Synthesized plan: 497 trunk tensors at Q5, 15 boundary tensor(s) floored to >= Q6, 1 lm_head tensor floored to >= Q6
(your own real counts will differ by model), then the model loads, a
real stratified calibration batch builds across all 6 content domains,
Part 5a PASS: real Hessian collection works end-to-end ... no NaN,
and finally a real per-tensor error-reduction number for each tensor
(90%+ is expected and is success).
05_full_model_quantize.py above is for the quick test only.
It is not safe to call directly for a real full run — a
real full run (every tensor, hours) risks a real, confirmed Metal
resource-handle crash. A real full run must go through
./run_full_yaqa.sh instead, which runs each memory-bounded
batch as its own fresh process:
./run_full_yaqa.sh /Users/yourname/Downloads/my-model-quantized \
--source /Users/yourname/Downloads/my-model --flat-bits 5 --n-calibration 4(--n-calibration 4 is this project’s own validated
choice for real production builds, not a hard requirement — the bare
code default is 1.) If interrupted for any reason, run the
exact same command again; run_full_yaqa.sh finds the
.yaqa_resume folder it already wrote next to your output
and continues from there instead of starting over.
Watch a real run live:
./monitor.sh <output>.yaqa_resume/logs/batch_<N>.log,
or open the live HTML dashboard it also serves
(generate_dashboard.py).
GPTQ is the other, independently-built full-model path in this
project — a simpler, one-sided correction (versus YAQA’s two-sided
H_I+H_O correction), still real and still
Hessian-weighted. Unlike YAQA, 08_gptq_apply_plan.py has
its own built-in resumability and internal calibration batching, so
there’s no separate wrapper script to call — you run it directly, for a
quick test or a real full run alike:
"$PYBIN" 08_gptq_apply_plan.py --source /Users/yourname/Downloads/my-model \
--flat-bits 5 --output /Users/yourname/Downloads/my-model-gptq-flatIf interrupted, re-running the same command resumes from
<output>.gptq_resume automatically, the same idea as
YAQA’s .yaqa_resume.
Before committing to a real multi-hour run, --inspect
reports real architecture facts about --source (and
--plan, if given) as JSON and exits immediately — no model
loading, no GPU work, no calibration:
"$PYBIN" 08_gptq_apply_plan.py --source /Users/yourname/Downloads/my-model --inspectThis project’s own real, benchmarked production build uses
YAQA for the trunk (497 tensors) and
GPTQ scoped narrowly to just lm_head (the
one tensor YAQA structurally can’t reach — its 248,320-wide output
dimension alone needs 246.7GB of Hessian memory, more than this
project’s own correction loop can hold). That combination — not a
full-model GPTQ build — is what’s actually shipped and benchmarked
(92.14% mean across MMLU/GSM8K/IFEval/BFCL/HumanEval, highest of six
real models compared).
GPTQ’s own --flat-bits path on the full trunk is real
and available — useful if you want a faster, simpler build and are
comfortable trading away YAQA’s extra correction quality — but it’s not
what this project’s own numbers above were measured on. See
README.md’s “Why two implementations” section for the full
real comparison.