VEGA / MLX OPTIQ LAB
Independent applied AI systems research · MLX · Apple Silicon

THE WRONG BITS ARE EXPENSIVE.

What if quantization quality depends less on how many bits you use — and more on where you spend them?

I turned MLX OptiQ's BF16 sensitivity measurements into a global resource-allocation problem: remove irrational precision choices, protect fragile structures, then spend the remaining budget where measured distortion falls fastest.
01 / THE ASSUMPTION
The inciting incident

More precision should help. It didn't always.

The BF16-reference sensitivity matrix broke the simple Q4 → Q5 → Q6 → Q8 quality ladder. Some higher-bit candidates cost more memory and measured worse.

One measured example is enough to make the problem concrete. For layers.3.mlp.gate_proj, Q5 had lower isolated KL than Q6.

Stock OptiQ selected Q6. Q5 was both cheaper and lower-distortion. That is a Pareto-dominated purchase.

If precision has a price, don't buy a worse result for more bits.

Across the Qwen3.8 target-5 Stock plan, 10 selections were meaningfully dominated. The Hybrid plan contains zero.

Measured sensitivity · one real tensor

The precision ladder is not monotonic

Lower isolated KL is better
0.0170.0130.0090.0040.000Q40.0152Q50.0094Q60.0145Q80.0119+1 bit, worse KL
Q5: 0.009370 KL · Q6: 0.014453 KL. The higher-bit candidate is ≈54% worse on this isolated measurement.
02 / THE CLIFF
The discovery hidden at the left edge

The dangerous part of the frontier is the beginning.

Near the minimum feasible budget, raw distortion rises gradually — but structurally weighted distortion can explode. A few hundredths of a bit can move the model into a different regime.

Raw KL asks how much error.
Weighted KL asks where it landed.

At Qwen3.6 target 4.40, raw ΣKL is ≈1.53 while weighted ΣKL is ≈4.05. Add only 0.05 BPW and the weighted objective collapses to ≈2.09 even though raw KL barely moves.

Qwen3.6 · 4.40→4.45−48.4%weighted KL for +0.05 BPW
Qwen3.8 · 4.35→4.506.49→2.63weighted KL across the transition
Run-guard pressure34→0Q36 promotions by 4.55 BPW
Interactive frontier explorer

Move the bit budget. Watch the structural penalty change.

Same V3.1 policy · model-relative BF16 reference
STRUCTURAL DANGER ZONE0.01.42.84.25.67.04.55.06.07.08.0weighted ΣKLraw ΣKL
Q4Q5Q6Q8BF16
Selected point
5.00 BPW
feasible · no run-guard promotions
raw 1.8153 · weighted 1.8920
Infeasible requests are shown explicitly instead of being silently omitted. Qwen3.8's 4.45 anomaly reflects the path-dependent run guard in the frozen V3.1 allocator.
03 / THE SYSTEM
One checkpoint, three jobs

A useful model is more than a quantized trunk.

The optimization only matters if the final artifact still sees, reasons and speculates correctly. Qwen3.8-27B here has three systems that must remain coherent.

◫
Language trunk

The brain

497 measured tensor assignments. This is where mixed Q4/Q5/Q6/Q8/BF16 precision is globally allocated.

experimental variable~25.6B measured params
+
◎
Vision tower

The eyes

Multimodal sidecar preserved in BF16. A fast text model with a broken vision tower is still a broken VLM.

333 tensorsbyte-identical control
+
↠
Native MTP

The fast junior

The native multi-token-prediction head drafts future tokens; the target model verifies them before they are committed.

15 logical tensors7 Q4/G64 projections
Speculative decoding

Senior engineer + fast junior assistant

Speed comes from accepted work, not guesses.
Target model

Senior engineer

Expensive and authoritative. Produces the hidden state and checks the drafter's work.

Native MTP

Fast junior

Drafts D1 / D2 / D3 future tokens cheaply from the target hidden state.

D1
D2
D3
Target verification

Approve the prefix

The target checks the drafted window and commits only the longest valid prefix.

✓
✓
✓
The fastest drafter is useless if the verifier rejects it.

This is why MTP stayed a separate experimental axis. Trunk precision can change the hidden states fed into the drafter; MTP precision can change draft quality; runtime scheduling can change the cost of verification.

The project eventually discovered that OptiQ's native MTP packaging — not a custom MTP sensitivity allocator — was the correct production rule for this model.

Fixed-state MTP proxies were still useful scientifically, but they did not predict live D3 behavior well enough to become the production optimizer.

04 / THE METHOD
The contribution

OptiQ measures the market. Pareto removes bad purchases. MILP spends the budget.

This work does not replace OptiQ's sensitivity or architecture-aware safeguards. It builds quality on top of them by changing the global allocation decision.

01

BF16 reference

Measure Q4/Q5/Q6/Q8/BF16 isolated KL for every candidate tensor.

02

Meaningful Pareto

Remove higher-bit choices that are materially worse than cheaper measured alternatives.

03

Structural safeguards

Keep OptiQ's protected output/boundary priority, attention floor and low-bit-run guard.

04

Global MILP

Solve all 497 choices together under one explicit BPW ceiling.

05

Source-only V5

Rebuild from the original source, then physically audit the artifact before release.

The problem is not “pick a quant.” It is allocate a scarce resource.

Each tensor has a menu of possible precisions, a measured distortion at each precision, a parameter mass, and possibly a structural priority. The whole model must fit under one budget.

A greedy local upgrade order can make individually plausible decisions that are globally suboptimal. The MILP evaluates the complete allocation simultaneously.

minimize   Σᵢ Σ_b xᵢ,b · KLᵢ,b · wᵢ
subject to   Σᵢ Σ_b xᵢ,b · paramsᵢ · b / Σparams ≤ target + slack
         Σ_b xᵢ,b = 1   for every measured tensor

where   Pareto-pruned candidate menus + architecture-aware constraints define what is legal.
05 / THE CONTROL
The result that matters

Same Qwen3.8 source. Same budget class. Better allocation.

The controlled comparison asks one narrow question: if the MTP and vision sidecars are held byte-for-byte constant, can a different language-trunk precision allocation reduce the measured distortion objective?

Hybrid changed only 68 of 497 assignments — but removed every meaningful Pareto-dominated Stock selection.

CONTROL SEALED ✓
Stock OptiQ · target 5
5.012811
nominal BPW
1.911571
raw Σ isolated KL
1.988341
weighted ΣKL
Hybrid Pareto/MILP · target 5
5.000409
nominal BPW · slightly lower budget
1.815278
raw Σ isolated KL
1.892047
weighted ΣKL
−5.04%raw ΣKL
−4.84%weighted objective
10 → 0meaningfully dominated choices
7.39%parameter mass changed

Where the allocator actually intervened

Each cell is one measured tensor assignment
unchangedHybrid spent more precisionHybrid spent less precision
Native MTP · identical
e3054bcb1b6fa35f482f6cca7cb3032ba9aa31a3cfb4d1e358d5ebb20b78d0f8
BF16 vision · identical
554ed429570a377ec99a75cfcf71d45fde9e64a207324a5cec8e6ad61c8cdaeb

Stock precision mass

Parameter-mass share
Q4 42.29%Q5 40.62%Q6 12.95%Q8 1.86%BF16 2.28%

Hybrid precision mass

Different topology, same budget class
Q4 45.10%Q5 39.33%Q6 11.22%Q8 1.74%BF16 2.61%
06 / THE ARTIFACT
Mathematics is not enough

The plan had to become a model that could be physically audited.

V5 is a source-only production builder. It consumes the frozen plan, rebuilds from the original checkpoint, preserves native MTP and BF16 vision, verifies exact assignment hits, hashes provenance, runs a load smoke, then finalizes atomically.

Original source

BF16 Qwen3.8-27B

Single ancestry root. No Stock donor model and no crossover artifact in the production build.

→
Frozen plan + native sidecars

V3.1 → V5 build

Mixed Q4/Q5/Q6/Q8/BF16 language trunk + OptiQ-native MTP + preserved BF16 vision.

→
Final artifact

Audited before rename

Physical packed shapes, sidecars, runtime load, MTP attach, hashes and manifest all checked before finalization.

497 / 497planned trunk assignments physically hit
15logical native MTP tensors audited
7native Q4/G64 MTP projections
333BF16 vision tensors preserved
The strongest experimental control is not a chart. It is two matching SHA256 hashes.
07 / GENERALITY
Same allocator, different model

Qwen3.6 and Qwen3.8 show the same shape of opportunity.

Absolute KL values are model-relative and must not be read as an intelligence ranking. What can be compared is how each model responds to the same V3.1 resource-allocation policy.

Same allocator · same +0.5 BPW move

4.5 → 5.0 BPW

Raw isolated-KL reduction
Qwen3.6
4.5 BPW1.4956
−24.53%
5.0 BPW1.1287
Qwen3.8
4.5 BPW2.4401
−25.61%
5.0 BPW1.8153
Absolute KL is model-relative. The comparable signal here is the fractional improvement produced by the same V3.1 budget move.
24.53% and 25.61%.

Despite different absolute sensitivity surfaces, both models recover almost the same fraction of raw isolated-KL by moving from 4.5 to 5.0 BPW under the same allocator.

That suggests the “extra half-bit” is doing a broadly similar job: buying the allocator enough freedom to move out of the low-budget danger regime and redistribute precision toward higher-return locations.

This is comparative allocation behavior, not a claim that one source model is intrinsically better than the other.

08 / RUNTIME PAYOFF
Paired MTPLX tune · MacBook Pro M3 Max · 128 GB unified memory

Same AR. Completely different speculative scaling.

This is the runtime result that connects the allocator to the system. Plain autoregressive decode is essentially tied: 17.58 vs 17.82 tok/s (+1.37%). Once MTP speculation is enabled, the curves separate. The flat 5-bit model peaks at D2 and then regresses. Hybrid V5 keeps scaling into D3.

AR baseline · near tie
17.58
↔
17.82
tok/s
Flat ↔ Hybrid · only +1.37%
Best tuned throughput
54.29
vs
45.18
tok/s
Hybrid best mode D3 vs Flat best mode D2
+20.15% peak tuned throughput
Same-depth D3
+36.15%
54.29 vs 39.87 tok/s · decode throughput
Speculative depth scaling

The curves tell the story.

Flat 5-bit MTPLXHybrid OptiQ-Pareto5 V5
20304050 AR 17.58 17.82 D1 29.66 38.04 D2 45.18 48.89 D3 39.87 54.29 FLAT: D2→D3 −11.75% HYBRID: D2→D3 +11.04%
Flat quantization
45.18 → 39.87 tok/s

D2 is the flat model's best mode. Increasing speculative depth to D3 makes it 11.75% slower.

Sensitivity-aware allocation
48.89 → 54.29 tok/s

Hybrid V5 continues to benefit from deeper speculation: D3 is 11.04% faster than its own D2.

The result is not “our weights are simply faster.” AR is nearly tied. The advantage appears when the runtime asks the model to speculate deeply.

That is the systems-level payoff of the allocation work. In this paired tune comparison, Hybrid D3 is not only faster; it accepts a larger fraction of drafts, needs fewer verification cycles, rejects fewer speculative tokens, and reaches a depth that the flat quant cannot use efficiently.

D3 acceptance
89.50% → 93.84%

+4.34 percentage points overall.

Depth-3 acceptance
82.71% → 88.66%

+5.95 pp at the deepest speculative step.

Verify calls
135 → 104

−22.96% verification cycles.

Rejected drafts
23 → 11

−52.17% rejected speculative drafts.

Correction tokens
24 → 12

−50.00% corrective work.

Draft time
2.070s → 1.069s

−48.37% total draft time.

Hidden verify cost
69.86 → 61.38

−12.13% hidden verify ms/call.

Memory trade-off
22.60 → 23.69 GB

+4.83% peak memory in speculative modes.

MacBook Pro M3 Max 128 GB unified memory long_warm_code_continuation 512-token budget AR · D1 · D2 · D3
Machine-readable paired report: Flat_5BIT_MTPLX_vs_HybridOptiQ-Pareto5_V5_runtime.json
Publication boundary: this is a paired MTPLX tune microbenchmark, not a capability benchmark. It demonstrates runtime/speculative-scaling behavior under this workload. The final three-model publication benchmark still freezes runtime, artifact verification and evaluation conditions across every model.
09 / FALSIFICATION
The experiment that refused to cooperate

Matching MTP precision did not restore the 6-BPW D3 scaling.

A useful research project records the hypothesis that failed. At ~6 BPW, strong D3-over-D2 scaling disappeared in both Stock and Hybrid regimes. Holding the Hybrid trunk fixed and raising MTP projections Q4 → Q5 → Q6 did not bring it back.

Hypothesis: “the trunk got more precise, so the MTP head should too.” Falsified.

The MTP projection bit-width alone is not responsible for the observed 6-BPW D3 discontinuity. The exact mechanism remains open.

This matters because it prevented a convenient but unsupported causal story from entering the production design.

Fixed ~6-BPW trunk

D3 vs D2 after changing MTP precision
D3 = D2MTP Q4MTP Q5MTP Q6−3.1%−7.4%−7.3%
All three variants leave D3 below D2 on the fixed ~6-BPW trunk. This falsifies simple MTP precision matching as the fix.
10 / RELEASE STRATEGY
Two operating points, two jobs

Not a ladder of bigger numbers. A deliberate frontier.

The low-BPW transition gives the release strategy a scientific reason: 4.5 sits just above the structural danger zone; 5.0 is the balanced flagship with a complete V5 physical validation chain.

Efficiency candidate · artifact validation pending
4.5

Push compression without stepping back into the cliff.

Qwen3.8 V3.1 plan is measured and feasible. The public runtime claim waits for its own source-only V5 build and regression.

2.440122raw ΣKL
2.626749weighted ΣKL
70.34%Q4 parameter mass
7run-guard promotions
Validated flagship
5.0

The balanced public artifact.

Frozen V3.1 plan → source-only V5 build → native OptiQ MTP → BF16 vision → physical audit → runtime load and repeated validation.

1.815278raw ΣKL
1.892047weighted ΣKL
5.000409nominal BPW
497/497physical assignment hits
11 / REPRODUCIBILITY
Research should have ancestry

Every headline number should resolve to an artifact.

The project keeps the research history separate from the public narrative, but does not erase it. Allocator versions, failed experiments, frontier CSVs, plan JSONs, sidecar hashes, builder code and benchmark outputs remain traceable.

Sensitivity science

BF16-reference Q4/Q5/Q6/Q8/BF16 isolated-KL matrix instead of guessed layer importance.

Optimization

Meaningful Pareto filtering plus global mixed-integer allocation under explicit resource ceilings.

Systems reverse engineering

Native MTP packaging, MLX physical quantization maps, sidecars and runtime contracts inspected and reproduced.

Artifact engineering

Source-only deterministic V5 build with exact predicate-hit audit, hashes, manifest and atomic finalization.

Experimental design

Stock controls, sidecar hash controls, crossover experiments, falsification tests and version-stamped runtime boundaries.

Technical communication

Beginner mental models and publication-grade claim boundaries live alongside the machine-readable evidence.

@misc{ghelab2026beyondgreedy,
  author = {Hakim Ghelab},
  title  = {Beyond Greedy Mixed Precision: Pareto-Pruned Global Bit Allocation for MLX on Apple Silicon},
  year   = {2026},
  note   = {Independent applied AI systems research; reproducible artifact and benchmark release}
}
Researcher / builder

Hakim Ghelab.

Independent Applied AI Systems Researcher focused on local inference, model optimization and reproducible experimentation. This project combines sensitivity analysis, global optimization, reverse engineering, multimodal preservation, runtime instrumentation and release engineering into one auditable systems result.

What this project demonstrates

Research framingmeasurement → hypothesis → control → falsification
OptimizationPareto + global MILP
ML systemsMLX / OptiQ / native MTP
MultimodalityBF16 vision preservation
Artifact discipline497/497 + SHA controls
Publication stanceclaim boundaries explicit