THE WRONG BITS ARE EXPENSIVE.
What if quantization quality depends less on how many bits you use — and more on where you spend them?
More precision should help. It didn't always.
The BF16-reference sensitivity matrix broke the simple Q4 → Q5 → Q6 → Q8 quality ladder. Some higher-bit candidates cost more memory and measured worse.
One measured example is enough to make the problem concrete. For layers.3.mlp.gate_proj, Q5 had lower isolated KL than Q6.
Stock OptiQ selected Q6. Q5 was both cheaper and lower-distortion. That is a Pareto-dominated purchase.
Across the Qwen3.8 target-5 Stock plan, 10 selections were meaningfully dominated. The Hybrid plan contains zero.
The precision ladder is not monotonic
The dangerous part of the frontier is the beginning.
Near the minimum feasible budget, raw distortion rises gradually — but structurally weighted distortion can explode. A few hundredths of a bit can move the model into a different regime.
Weighted KL asks where it landed.
At Qwen3.6 target 4.40, raw ΣKL is ≈1.53 while weighted ΣKL is ≈4.05. Add only 0.05 BPW and the weighted objective collapses to ≈2.09 even though raw KL barely moves.
Move the bit budget. Watch the structural penalty change.
A useful model is more than a quantized trunk.
The optimization only matters if the final artifact still sees, reasons and speculates correctly. Qwen3.8-27B here has three systems that must remain coherent.
The brain
497 measured tensor assignments. This is where mixed Q4/Q5/Q6/Q8/BF16 precision is globally allocated.
The eyes
Multimodal sidecar preserved in BF16. A fast text model with a broken vision tower is still a broken VLM.
The fast junior
The native multi-token-prediction head drafts future tokens; the target model verifies them before they are committed.
Senior engineer + fast junior assistant
Senior engineer
Expensive and authoritative. Produces the hidden state and checks the drafter's work.
Fast junior
Drafts D1 / D2 / D3 future tokens cheaply from the target hidden state.
Approve the prefix
The target checks the drafted window and commits only the longest valid prefix.
This is why MTP stayed a separate experimental axis. Trunk precision can change the hidden states fed into the drafter; MTP precision can change draft quality; runtime scheduling can change the cost of verification.
The project eventually discovered that OptiQ's native MTP packaging — not a custom MTP sensitivity allocator — was the correct production rule for this model.
Fixed-state MTP proxies were still useful scientifically, but they did not predict live D3 behavior well enough to become the production optimizer.
OptiQ measures the market. Pareto removes bad purchases. MILP spends the budget.
This work does not replace OptiQ's sensitivity or architecture-aware safeguards. It builds quality on top of them by changing the global allocation decision.
BF16 reference
Measure Q4/Q5/Q6/Q8/BF16 isolated KL for every candidate tensor.
Meaningful Pareto
Remove higher-bit choices that are materially worse than cheaper measured alternatives.
Structural safeguards
Keep OptiQ's protected output/boundary priority, attention floor and low-bit-run guard.
Global MILP
Solve all 497 choices together under one explicit BPW ceiling.
Source-only V5
Rebuild from the original source, then physically audit the artifact before release.
Each tensor has a menu of possible precisions, a measured distortion at each precision, a parameter mass, and possibly a structural priority. The whole model must fit under one budget.
A greedy local upgrade order can make individually plausible decisions that are globally suboptimal. The MILP evaluates the complete allocation simultaneously.
subject to Σᵢ Σ_b xᵢ,b · paramsᵢ · b / Σparams ≤ target + slack
Σ_b xᵢ,b = 1 for every measured tensor
where Pareto-pruned candidate menus + architecture-aware constraints define what is legal.
Same Qwen3.8 source. Same budget class. Better allocation.
The controlled comparison asks one narrow question: if the MTP and vision sidecars are held byte-for-byte constant, can a different language-trunk precision allocation reduce the measured distortion objective?
Hybrid changed only 68 of 497 assignments — but removed every meaningful Pareto-dominated Stock selection.
Where the allocator actually intervened
Each cell is one measured tensor assignmente3054bcb1b6fa35f482f6cca7cb3032ba9aa31a3cfb4d1e358d5ebb20b78d0f8554ed429570a377ec99a75cfcf71d45fde9e64a207324a5cec8e6ad61c8cdaebStock precision mass
Parameter-mass shareHybrid precision mass
Different topology, same budget classThe plan had to become a model that could be physically audited.
V5 is a source-only production builder. It consumes the frozen plan, rebuilds from the original checkpoint, preserves native MTP and BF16 vision, verifies exact assignment hits, hashes provenance, runs a load smoke, then finalizes atomically.
BF16 Qwen3.8-27B
Single ancestry root. No Stock donor model and no crossover artifact in the production build.
V3.1 → V5 build
Mixed Q4/Q5/Q6/Q8/BF16 language trunk + OptiQ-native MTP + preserved BF16 vision.
Audited before rename
Physical packed shapes, sidecars, runtime load, MTP attach, hashes and manifest all checked before finalization.
Qwen3.6 and Qwen3.8 show the same shape of opportunity.
Absolute KL values are model-relative and must not be read as an intelligence ranking. What can be compared is how each model responds to the same V3.1 resource-allocation policy.
4.5 → 5.0 BPW
Despite different absolute sensitivity surfaces, both models recover almost the same fraction of raw isolated-KL by moving from 4.5 to 5.0 BPW under the same allocator.
That suggests the “extra half-bit” is doing a broadly similar job: buying the allocator enough freedom to move out of the low-budget danger regime and redistribute precision toward higher-return locations.
This is comparative allocation behavior, not a claim that one source model is intrinsically better than the other.
Same AR. Completely different speculative scaling.
This is the runtime result that connects the allocator to the system. Plain autoregressive decode is essentially tied: 17.58 vs 17.82 tok/s (+1.37%). Once MTP speculation is enabled, the curves separate. The flat 5-bit model peaks at D2 and then regresses. Hybrid V5 keeps scaling into D3.
The curves tell the story.
D2 is the flat model's best mode. Increasing speculative depth to D3 makes it 11.75% slower.
Hybrid V5 continues to benefit from deeper speculation: D3 is 11.04% faster than its own D2.
That is the systems-level payoff of the allocation work. In this paired tune comparison, Hybrid D3 is not only faster; it accepts a larger fraction of drafts, needs fewer verification cycles, rejects fewer speculative tokens, and reaches a depth that the flat quant cannot use efficiently.
+4.34 percentage points overall.
+5.95 pp at the deepest speculative step.
−22.96% verification cycles.
−52.17% rejected speculative drafts.
−50.00% corrective work.
−48.37% total draft time.
−12.13% hidden verify ms/call.
+4.83% peak memory in speculative modes.
Matching MTP precision did not restore the 6-BPW D3 scaling.
A useful research project records the hypothesis that failed. At ~6 BPW, strong D3-over-D2 scaling disappeared in both Stock and Hybrid regimes. Holding the Hybrid trunk fixed and raising MTP projections Q4 → Q5 → Q6 did not bring it back.
The MTP projection bit-width alone is not responsible for the observed 6-BPW D3 discontinuity. The exact mechanism remains open.
This matters because it prevented a convenient but unsupported causal story from entering the production design.
Fixed ~6-BPW trunk
D3 vs D2 after changing MTP precisionNot a ladder of bigger numbers. A deliberate frontier.
The low-BPW transition gives the release strategy a scientific reason: 4.5 sits just above the structural danger zone; 5.0 is the balanced flagship with a complete V5 physical validation chain.
Push compression without stepping back into the cliff.
Qwen3.8 V3.1 plan is measured and feasible. The public runtime claim waits for its own source-only V5 build and regression.
The balanced public artifact.
Frozen V3.1 plan → source-only V5 build → native OptiQ MTP → BF16 vision → physical audit → runtime load and repeated validation.
Every headline number should resolve to an artifact.
The project keeps the research history separate from the public narrative, but does not erase it. Allocator versions, failed experiments, frontier CSVs, plan JSONs, sidecar hashes, builder code and benchmark outputs remain traceable.
BF16-reference Q4/Q5/Q6/Q8/BF16 isolated-KL matrix instead of guessed layer importance.
Meaningful Pareto filtering plus global mixed-integer allocation under explicit resource ceilings.
Native MTP packaging, MLX physical quantization maps, sidecars and runtime contracts inspected and reproduced.
Source-only deterministic V5 build with exact predicate-hit audit, hashes, manifest and atomic finalization.
Stock controls, sidecar hash controls, crossover experiments, falsification tests and version-stamped runtime boundaries.
Beginner mental models and publication-grade claim boundaries live alongside the machine-readable evidence.
@misc{ghelab2026beyondgreedy,
author = {Hakim Ghelab},
title = {Beyond Greedy Mixed Precision: Pareto-Pruned Global Bit Allocation for MLX on Apple Silicon},
year = {2026},
note = {Independent applied AI systems research; reproducible artifact and benchmark release}
}Hakim Ghelab.
Independent Applied AI Systems Researcher focused on local inference, model optimization and reproducible experimentation. This project combines sensitivity analysis, global optimization, reverse engineering, multimodal preservation, runtime instrumentation and release engineering into one auditable systems result.