← Back to index

Flat Mode — Quantize Without a Sensitivity Plan

Hakim Ghelab, VegaLaboratories LTD

2026-10-01

Who this page is for: you have a BF16 model and want a quantized MLX build, but you don’t have (and don’t want to build) this project’s own real per-tensor sensitivity plan — the stratified-calibration → KL-sensitivity → MILP-allocation pipeline that normally decides which bit width every individual tensor gets. Flat mode skips all of that. You pick one bit width, and every real trunk tensor in the model gets it — except a small, automatically-protected set of tensors that get a real hard floor instead, so quality-critical layers are never silently flattened down to your chosen width.

This is the primary path for anyone outside this project — nobody else has this project’s own private Hessian-sweep data, so --plan (a full custom multi-bit allocation) only works for someone who has run that whole separate pipeline themselves. Flat mode needs nothing but a source model.

What actually happens when you run it

Flat mode is implemented by one shared module, flat_plan.py, called from two otherwise-independent build scripts — YAQA (05_full_model_quantize.py) and GPTQ (08_gptq_apply_plan.py). Both take the exact same --flat-* flags and produce the exact same kind of plan; only the actual correction math that gets applied afterward differs (YAQA’s two-sided Hessian correction vs. GPTQ’s one-sided correction — see README.md for why this project built both).

Given a source model and a bit width, Flat mode:

  1. Loads the model’s real module structure (no weight data, just the list of every quantizable nn.Linear/SwitchLinear tensor — the same leaf-type set the build scripts actually quantize).
  2. Classifies each tensor into one of three groups:
  3. Assigns bits: trunk tensors get exactly --flat-bits. Boundary and lm_head tensors get max(--flat-bits, their own floor) — so asking for a higher --flat-bits than the floor never silently downgrades a protected tensor; it only ever raises it.
  4. Writes a real plan.json to disk (same schema every other plan file in this project already uses) and points the build script’s own --plan at it — every existing, already-trusted call site in either script then runs completely unchanged. Flat mode never touches the actual correction logic, only which bit width gets requested.

Why boundary gets a real hard floor at all: in a normal MILP-built plan, boundary protection comes from a soft cost weighting the solver can freely trade against a real token budget. Flat mode has no solver and no budget to trade against — it’s one direct assignment — so it can’t reuse “no hard floor” as if it meant “no protection needed.” The default floor (6-bit) is the same one already given to lm_head, justified by real data from this project’s own Hessian sweep: the last boundary layer contains the single highest-danger tensor of all 497 scored in the model.

The flags (identical on both scripts)

Flag Default Meaning
--flat-bits N (none — required to turn Flat mode on) The one bit width every trunk tensor gets. Must be one of 4, 5, 6, 8 — the exact set MLX’s native quantizer and this project’s GPTQ packing both support.
--flat-group-size N 64 Quantization group size applied to every tensor under Flat mode.
--flat-boundary-min-bits N 6 Hard floor for boundary tensors (see above). A higher --flat-bits can raise this; a lower one never lowers it.
--flat-lm-head-min-bits N 6 Hard floor for lm_head, same logic.
--flat-n-protect N 1 How many layer indices at each end (first and last) count as “boundary” — not how many tensors. 1 means every tensor belonging to layer 0 and every tensor belonging to the final layer, not one representative tensor per layer.
--output PATH (none — required) Where the build is written. Also required by Flat mode specifically, because the synthesized plan.json is placed next to it (Path(args.output).parent). Both scripts raise a clear error (ap.error(...)) if you give --flat-bits without --output.

--plan and --flat-bits are mutually exclusive in spirit: if you pass both, --flat-bits wins and its synthesized plan silently replaces whatever --plan pointed at.

Concrete example, real tensor names

On this project’s own 64-layer hybrid model, –flat-n-protect 1 protects every tensor in layer 0 and layer 63 — all 15 of them, not one tensor named “layer 0”:

Layer 0 (8 tensors -- a GatedDeltaNet layer):
  language_model.model.layers.0.linear_attn.in_proj_a
  language_model.model.layers.0.linear_attn.in_proj_b
  language_model.model.layers.0.linear_attn.in_proj_qkv
  language_model.model.layers.0.linear_attn.in_proj_z
  language_model.model.layers.0.linear_attn.out_proj
  language_model.model.layers.0.mlp.down_proj
  language_model.model.layers.0.mlp.gate_proj
  language_model.model.layers.0.mlp.up_proj

Layer 63 (7 tensors -- a standard attention layer):
  language_model.model.layers.63.mlp.down_proj
  language_model.model.layers.63.mlp.gate_proj
  language_model.model.layers.63.mlp.up_proj
  language_model.model.layers.63.self_attn.k_proj
  language_model.model.layers.63.self_attn.o_proj
  language_model.model.layers.63.self_attn.q_proj
  language_model.model.layers.63.self_attn.v_proj

All 15 of those get max(–flat-bits, –flat-boundary-min-bits) — exactly matching the real [flat-mode] Synthesized plan: … 15 boundary tensor(s) floored to >= Q6 line from an actual tested run of this project. Layer counts differ by architecture (GatedDeltaNet layers like layer 0 have no self_attn at all, which is why it’s 8 tensors instead of 7) — run your own model through mtplx inspect or check your plan file to see its real per-layer names.

Don’t confuse these two

08_gptq_apply_plan.py (GPTQ only) also has a separate –calibration-mode flag with a flat choice (stratified is the default). That controls how calibration text gets sampled — nothing to do with bit-width assignment. Flat mode (this page, –flat-bits) picks the quantization bit width with no plan file. Flat calibration (–calibration-mode flat) picks how calibration text is drawn, and is a completely separate axis — you can combine either with either. 05_full_model_quantize.py (YAQA) has no –calibration-mode flag at all; it always uses stratified calibration.

Running it — YAQA

Quick test first (3 example tensors, a couple of minutes, needs your own real BF16 model):

"$PYBIN" 05_full_model_quantize.py --source /Users/yourname/Downloads/my-model \
    --flat-bits 5 --subset default --output /Users/yourname/Downloads/my-model-quicktest

What success looks like: a line like [flat-mode] Synthesized plan: 497 trunk tensors at Q5, 15 boundary tensor(s) floored to >= Q6, 1 lm_head tensor floored to >= Q6 (your own real counts will differ by model), then the model loads, a real stratified calibration batch builds across all 6 content domains, Part 5a PASS: real Hessian collection works end-to-end ... no NaN, and finally a real per-tensor error-reduction number for each tensor (90%+ is expected and is success).

One rule with no exception

05_full_model_quantize.py above is for the quick test only. It is not safe to call directly for a real full run — a real full run (every tensor, hours) risks a real, confirmed Metal resource-handle crash. A real full run must go through ./run_full_yaqa.sh instead, which runs each memory-bounded batch as its own fresh process:

./run_full_yaqa.sh /Users/yourname/Downloads/my-model-quantized \
    --source /Users/yourname/Downloads/my-model --flat-bits 5 --n-calibration 4

(--n-calibration 4 is this project’s own validated choice for real production builds, not a hard requirement — the bare code default is 1.) If interrupted for any reason, run the exact same command again; run_full_yaqa.sh finds the .yaqa_resume folder it already wrote next to your output and continues from there instead of starting over.

Watch a real run live: ./monitor.sh <output>.yaqa_resume/logs/batch_<N>.log, or open the live HTML dashboard it also serves (generate_dashboard.py).

Running it — GPTQ

GPTQ is the other, independently-built full-model path in this project — a simpler, one-sided correction (versus YAQA’s two-sided H_I+H_O correction), still real and still Hessian-weighted. Unlike YAQA, 08_gptq_apply_plan.py has its own built-in resumability and internal calibration batching, so there’s no separate wrapper script to call — you run it directly, for a quick test or a real full run alike:

"$PYBIN" 08_gptq_apply_plan.py --source /Users/yourname/Downloads/my-model \
    --flat-bits 5 --output /Users/yourname/Downloads/my-model-gptq-flat

If interrupted, re-running the same command resumes from <output>.gptq_resume automatically, the same idea as YAQA’s .yaqa_resume.

Before committing to a real multi-hour run, --inspect reports real architecture facts about --source (and --plan, if given) as JSON and exits immediately — no model loading, no GPU work, no calibration:

"$PYBIN" 08_gptq_apply_plan.py --source /Users/yourname/Downloads/my-model --inspect

Which one should I use?

This project’s own real, benchmarked production build uses YAQA for the trunk (497 tensors) and GPTQ scoped narrowly to just lm_head (the one tensor YAQA structurally can’t reach — its 248,320-wide output dimension alone needs 246.7GB of Hessian memory, more than this project’s own correction loop can hold). That combination — not a full-model GPTQ build — is what’s actually shipped and benchmarked (92.14% mean across MMLU/GSM8K/IFEval/BFCL/HumanEval, highest of six real models compared).

GPTQ’s own --flat-bits path on the full trunk is real and available — useful if you want a faster, simpler build and are comfortable trading away YAQA’s extra correction quality — but it’s not what this project’s own numbers above were measured on. See README.md’s “Why two implementations” section for the full real comparison.