VEGA / MLX OPTIQ LAB
Independent applied AI systems research · Apple Silicon · MLX

Beyond greedy
mixed precision.

A reproducible research pipeline that measures BF16 tensor sensitivity, removes dominated precision choices, globally allocates Q4/Q5/Q6/Q8/BF16 with MILP, and builds auditable MLX models with native MTP speculative decoding and preserved multimodal sidecars.

Built on mlx-optiq's sensitivity-driven workflow and structural safeguards. This project adds Pareto filtering, global allocation, source-only production tooling and controlled runtime validation.

25.62Bmeasured trunk parameters
429 / 497assignments unchanged vs Stock
0meaningfully dominated Hybrid choices
Q4 / G64native MTP projections
BF16vision sidecar preserved
01 — Measured result

Same budget class. Better allocation.

The important comparison is not “high bits versus low bits.” It is where the bit budget is spent when both methods see the same Qwen3.8 BF16-reference sensitivity matrix.

Stock OptiQ · target 5.0
5.0128nominal BPW
1.91157Σ isolated KL
Q4 42.3%Q5 40.6%Q6 13.0%Q8 1.9%BF16 2.3%
Hybrid Pareto / global MILP · target 5.0
5.0004nominal BPW
1.81528Σ isolated KL
Q4 45.1%Q5 39.3%Q6 11.2%Q8 1.7%BF16 2.6%
5.04% lower planning proxy · slightly lower nominal budget
01

Pareto cleanup matters—but it is not the whole gain.

Stock contains 10 meaningfully dominated choices. Correcting only those explains roughly one quarter of the measured KL improvement. The rest comes from global reallocation across the model.

02

More BF16 tensors does not mean “a high-bit model.”

The optimizer can leave many tiny sensitive tensors in BF16 while spending more Q4 mass on large tolerant tensors. At 5 BPW, BF16 rises only from 2.28% to 2.61% of parameter mass.

03

The protected boundary set is not sacrificed.

The key output/boundary choices match the Stock control. The measured improvement comes from reallocating the remaining degrees of freedom, not deleting safeguards.

02 — Physical artifact → runtime

The plan survived contact with the machine.

The V5 artifact was rebuilt from the original source, physically audited, loaded with native MTP, then tuned through MTPLX. Representative fresh runs repeatedly land around 54.5–54.65 tok/s at D3.

AR
17.90
17.89
D1
34.38
37.95
D2
49.04
48.64
D3
53.96
54.65
Stock OptiQ 5 BPWHybrid Pareto/MILP V5
54.65tok/s D3Representative V5 run; AR 17.89 tok/s.
~3.0×over autoregressive decodeD3 is the winning depth in the validated 5-BPW regime.
source-onlyprovenanceNo Stock quant is used as a donor in final V5 ancestry.
Publication boundary: tok/s measures runtime throughput. ΣKL measures a calibration planning proxy. Neither number, alone, means “5% smarter.” Coding, tools, long context and vision remain separate regression questions.
03 — Method

A global resource-allocation problem, not a quantization preset.

The core move is simple: measure first, then spend precision where its measured marginal return is highest.

01

BF16 reference

Freeze the source. Measure 497 language tensors against the high-precision reference.

02

Sensitivity matrix

Test Q4 / Q5 / Q6 / Q8 / BF16 one tensor at a time. Record isolated KL + parameter cost.

03

Meaningful Pareto

Remove higher-bit options that are materially worse than cheaper measured alternatives.

04

Structural safeguards

Retain OptiQ-style output / boundary priority, precision floors and low-bit-run protection.

05

Global MILP

Minimize weighted isolated KL under one explicit BPW ceiling instead of greedily upgrading locally.

06

V5 production build

Original source + exact plan. Rebuild native MTP and vision sidecars from source, audit every physical assignment, hash, smoke-test, atomically finalize.

07

MTPLX validation

Measure AR / D1 / D2 / D3 repeatedly. Runtime decides whether the mathematically attractive plan is operationally useful.

“Important-looking modules are not automatically the best place to spend precision. Measure the marginal return, preserve the invariants, then optimize the whole budget.”
04 — Release strategy

Two builds, not a ladder of bigger numbers.

The public line is intentionally narrow: an Efficiency 4.5-BPW target and a Balanced 5-BPW flagship. Higher budget is not automatically better for the complete speculative-decoding system.

EFFICIENCY4.5 BPW

The quality-biased knee.

On the Qwen3.6 measured frontier, ~4.5 BPW sits just to the quality side of the geometric knee while remaining predominantly Q4/Q5.

44.2%flat-Q4 isolated-KL error removed
60.97%Q4 parameter mass
0.60%BF16 parameter mass
Hugging Face release · publishing slot
FLAGSHIP / VALIDATED5.0 BPW

The highest-confidence public build.

Qwen3.8 V5 is physically built and runtime-validated: global mixed precision, native Q4/G64 MTP, BF16 vision sidecar, no Stock donor ancestry.

5.0004selected nominal BPW
54.65representative D3 tok/s
−5.04%Σ isolated KL vs Stock 5
Open Hugging Face profile ↗
Why 4.5 and 5?

Precision returns flatten; system behavior matters.

The historical Qwen3.6 frontier gives an intuitive view: the first extra half-bit buys a large reduction in isolated KL, then returns progressively diminish. Qwen3.8 runtime adds a second constraint: the ~6-BPW regime did not preserve the strong D3-over-D2 scaling seen near 5 BPW.

4.5 · ΣKL 1.22985.0 · ΣKL 1.0064 4.04.55.05.56.06.5Target / achieved BPW

4.5 data shown here is the measured Qwen3.6 frontier story. A Qwen3.8 4.5 release should receive its own final V5 build + runtime regression before its model card inherits performance claims.

05 — A failed hypothesis worth publishing

Matching MTP precision did not restore D3.

A useful research project records what fails. The ~6-BPW D3 issue appeared in genuine Stock OptiQ and in the Hybrid trunk. So the next hypothesis was: perhaps a higher-precision trunk needs a higher-precision MTP sidecar.

HYPOTHESIS~6 BPW trunk + Q5/Q6 MTP → D3 scaling returns
→
RESULTNot supportedQ4, Q5 and Q6 MTP variants did not recover the strong D3-over-D2 scaling.
Fixed ~6 BPW Hybrid trunkD2 tok/sD3 tok/sWinner
MTP Q4 / G6439.8438.57D2
MTP Q5 / G6439.92*36.98*D2*
MTP Q6 / G6439.5536.68D2

*Representative controlled comparison run. Q5 showed one noisy run where D3 won, but repeated runs did not reproduce 5-BPW-style linear D3 scaling. The narrow conclusion is that MTP projection bit-width alone does not explain the ~6-BPW discontinuity.

What this does not prove: it does not establish a universal law that “higher precision breaks D3.” It establishes a reproducible Qwen3.8 / MTPLX operating-regime discontinuity that survived a targeted MTP-precision counter-experiment.
06 — Start here if you are new

The definitions I wish I had on day one.

You should not need to already be a quantization engineer to understand the experiment. Open any term below for the exact meaning used on this site.

BPW+

Bits per weight. Here it is the parameter-weighted average precision of the measured language trunk—not a promise that every tensor has the same bit-width.

Sensitivity+

A controlled one-tensor-at-a-time test: change one weight matrix to Q4/Q5/Q6/Q8, keep the rest at the reference state, and measure how much the output distribution moves.

KL+

Kullback–Leibler divergence from the BF16 reference distribution. Lower is better for this planning proxy. It is not an intelligence score or a benchmark result.

Pareto pruning+

Do not pay more bits for an option that measures materially worse than a cheaper one. Dominated precision choices are removed before global allocation.

MILP+

Mixed-Integer Linear Programming. Instead of upgrading tensors greedily, solve the whole 497-choice allocation problem together under one global bit budget.

MTP+

Multi-Token Prediction head: a compact drafter that proposes future tokens. The main trunk verifies them; accepted drafts can make generation much faster.

D1 / D2 / D3+

Speculative depths. D3 attempts a deeper draft chain than D2. More depth only helps when the extra draft/verification work pays for itself.

V5 builder+

The production path: original source + exact plan → mixed-precision trunk, native OptiQ MTP, source vision/audio sidecars, physical audits, hashes and load verification. Stock models are controls, never donors.

MAIN TRUNKSenior engineerExpensive, trusted, final decision.
→
MTPFast junior assistantDrafts the next 1–3 tokens cheaply.
→
MTPLXApproval workflowKeeps drafts the senior agrees with.
07 — Reproducibility

A model result should have an ancestry, not a story.

V5 was designed to make the production artifact independently reconstructable and to fail rather than silently lose the MTP head, vision sidecar or planned quantization map.

ORIGINAL BF16 SOURCE├─ language → exact Pareto/MILP plan → mixed trunk├─ mtp.* → installed OptiQ preserve_mtp() → native sidecar└─ vision/audio → installed OptiQ attach_sidecars()↓V5 FINAL MODELStock artifact dependency: NONE
mlx-optiq0.4.22
mlx-lm0.31.3
mlx0.32.0
MTP15 logical · 7 packed / 8 protected
Vision333 tensors · BF16
Physical plan497 exact predicate hits
08 — Researcher

Hakim Ghelab

Independent Applied AI Systems Researcher

I work at the boundary between local inference, model optimization and production tooling: form a falsifiable hypothesis, instrument the system, build the control, preserve provenance, then let measured behavior—not intuition—decide what survives.

This project grew from learning MLX and quantization from first principles into a reproducible optimization lab spanning sensitivity analysis, integer optimization, speculative decoding, multimodal preservation, runtime instrumentation and release engineering.

Inference optimizationApplied AI systemsMLX / Apple SiliconTechnical strategySolutions architecture
Cite / credit

Built in the open, credited precisely.

This is independent research built around mlx-optiq, MLX and MTPLX. The allocator and V5 tooling described here are project extensions; upstream projects retain their own authorship and licenses. Model releases should preserve the source model's license and attribution requirements.

@misc{ghelab2026mlxoptiq,
  author = {Hakim Ghelab},
  title  = {Beyond Greedy Mixed Precision: Pareto/MILP Allocation for MLX},
  year   = {2026},
  note   = {Independent applied AI systems research; artifact and benchmark links in release model cards}
}