Fewer experts, faster answers on a Mac?
A one-line change makes MoE models generate faster on Apple silicon. We tested Qwen3 and DeepSeek at bf16 and 4-bit: the speed gains travelled better than the quality results.
Chapters
When a local assistant is writing an answer, a few more tokens per second can make the wait shorter. But if a faster answer needs correcting, the time saved may simply move from the computer to you.
A recent paper suggested a tempting adjustment: ask a Mixture-of-Experts model to consult fewer experts for each token. No retraining, no new model download. Just change how many of its existing experts take part. Chen and colleagues reported retaining about 99% of the evaluated quality on average, with faster generation on the four models and two serving stacks they timed.1
We tried the idea in MLX on Apple silicon, using Qwen3-30B-A3B and DeepSeek-V2-Lite-Chat. Both got faster. At 4-bit, with one request at a time, the first reduction improved generation throughput by about 8% for Qwen3 and 11% for DeepSeek on our M5 Max. Their answers told different stories: Qwen3 had no statistically detectable loss on our full maths test at that setting; DeepSeek lost about four percentage points.
The useful question became: how much work can we skip before the answer stops being worth the shorter wait?
What an expert does
In a Mixture-of-Experts, or MoE, model, some layers contain a collection of small neural networks called experts. A router scores them for each token and selects a few. Their outputs are weighted and combined before computation continues. “Expert” is a name for a learned component, not a promise that one knows astronomy and another writes good SQL.
Qwen3-30B-A3B has 128 routed experts per MoE layer and normally selects eight for each token. We asked it to use six, then five. DeepSeek-V2-Lite-Chat normally selects six from 64 routed experts; we tried four, then three. Its two shared experts remain active.
The paper's first setting keeps two thirds of the usual selection, rounded up. That makes Qwen3's eight-to-six change a 25% reduction, while DeepSeek's six-to-four change removes a third. The router still chooses which experts to use for each token.
This is different from the top-k sampling setting used to choose the next output token. Here we change the computation that produces those token probabilities. All the expert weights remain loaded, so this does not make the model smaller or free up memory.
What changed on our Macs
We tested the models at bf16, a 16-bit floating-point format, and with 4-bit quantized weights. The main comparisons below come from an idle M5 Max with 128 GB of unified memory, using MLX 0.32.3 and mlx-lm 0.32.0. We also ran supporting tests on an M5 Pro and an M1 Max.2
For each speed comparison, we used six fresh processes per setting and interleaved the settings rather than running every baseline first. The single-request test generated 512 tokens from a 944-token prompt. A separate eight-request test used shorter prompts and generated 256 tokens per request. Within each comparison, the inputs and output lengths were fixed.
The single-request results show the tradeoff most directly:
| Model | Weights | Experts | Before | After | Gain |
|---|---|---|---|---|---|
| Qwen3 | bf16 | 8 → 6 | 70.3 | 79.0 | 12.3% |
| Qwen3 | 4-bit | 8 → 6 | 144.4 | 155.6 | 7.8% |
| DeepSeek | bf16 | 6 → 4 | 90.6 | 105.2 | 16.1% |
| DeepSeek | Our 4-bit conversion | 6 → 4 | 194.3 | 215.2 | 10.8% |
Before and after are generated tokens per second on the M5 Max. “Gain” is the increase in throughput, not the percentage reduction in elapsed time.
The 4-bit models were already much faster in absolute terms, but gained less from skipping experts. Those are two different comparisons. A smaller percentage improvement does not mean that bf16 is the faster way to run the model.
Count the bytes, then measure
Generating one token at a time often leaves a GPU waiting for weights to arrive from memory. Apple describes this memory-bandwidth constraint in its work on MLX inference.3 If fewer experts take part, fewer expert weights need to be read. That suggests a way to estimate the opportunity.
Suppose routed experts account for a fraction of the bytes read in a step, and we keep a fraction of those experts. If everything else stays constant, the remaining bytes are proportional to:
If time were proportional to those bytes, the speedup would be the reciprocal. For example, when routed experts account for 56% of the bytes, keeping six instead of eight gives , or about 1.16 times the throughput. Removing a quarter of the experts has not removed a quarter of the whole step.
We calculated tensor storage from the models' weight-file headers and quantization formats, including quantization metadata.4 Routed experts represented roughly 53–58% of the modelled bytes per step across the two models and precisions. Our initial explanation had been that 4-bit must leave a much smaller expert share to remove. Our own byte count ruled that explanation out.
For the reductions in this figure, bf16 realised about 72% of the improvement suggested by the byte calculation. The 4-bit models realised about 48–50%. The more aggressive reductions gave a similar pattern at batch one: around 70% for bf16 and around half for 4-bit.
The distinction matters when judging an estimate. A measured 1.08-times speedup is close to an estimated 1.16-times speedup if we divide the two ratios. But the improvement is eight percentage points rather than sixteen: about half as much extra throughput. We initially used the former comparison and made the estimate look more accurate than it was.
The byte calculation describes an opportunity under an assumption, not a prediction or a universal ceiling. At eight simultaneous requests, Qwen3 bf16 actually exceeded the estimate, realising 144–184% of the estimated improvement across its two reductions. DeepSeek did not repeat that result.
At batch eight, the estimate also assumes that requests choose experts independently and uniformly. Real requests can share more experts than that assumption allows for, making the distinct bytes read harder to estimate. We did not establish how much this accounts for the gap. The batch-eight workload also used different prompt and output lengths, so it is a separate operating regime rather than a pure test of batch size alone.
Why does 4-bit realise less of the estimated gain? Unpacking quantized weights, launching GPU kernels and other per-step work are plausible contributors. We did not profile their individual costs, so the measurements do not choose between those explanations. They do tell us not to carry a bf16 percentage gain over to a 4-bit deployment without measuring it.
Faster at answering what?
For quality, we used all 1 319 problems in the GSM8K maths test, with greedy decoding. Each reduced setting answered the same problems as its own baseline. The comparison is paired: it tracks which individual answers changed, not just two unrelated overall scores.
Qwen3's bf16 accuracy went from 94.09% at eight experts to 93.48% at six. That is a drop of 0.61 percentage points. Its paired 95% confidence interval ran from a loss of 1.67 points to a gain of 0.38. The 4-bit change was smaller, at minus 0.15 points, with an interval from minus 1.14 to plus 0.83.
Those intervals include zero, but also include losses. “No statistically detectable loss” does not establish that the answers are equivalent.
DeepSeek's bf16 accuracy fell from 69.67% at six experts to 65.50% at four, a loss of 4.17 points. Its confidence interval stayed below zero. Taking away another expert widened the gap further.
The same rule for choosing fewer experts produced different quality costs in the two models. Qwen3 tolerated the first reduction on this test better than DeepSeek did. Neither result gives us a blanket setting for local models.
There is an implementation difference worth investigating. Qwen3 renormalizes the router weights after selecting the retained experts. DeepSeek-V2-Lite does not, so removing experts also reduces the total weight of their routed contribution. That may matter, but we did not run the experiment needed to establish it as the cause.
The paper's results provide a useful comparison. Its Qwen3-30B-A3B-Instruct scored 94.4% on GSM8K at eight experts and 95.0% at six. Its DeepSeek-V2-Lite-Chat score fell from 63.1% to 61.7% at six versus four experts: a 1.4-point loss, compared with our 4.2 points at bf16.1
Our Qwen3 checkpoint was the original hybrid model running without thinking, rather than the paper's Instruct checkpoint. DeepSeek's prompts, decoding and answer extraction also differed. This is a test of how the idea travels to our setup, not an identical reproduction of every published condition.
Check the baseline before changing it
Our first quality probe gave a less encouraging answer for 4-bit Qwen3. It used 200 problems, sampled generation with five seeds, batch one and an M1 Max. Six experts lost 1.5 percentage points and failed the acceptance band we had set before the test.
The later full test did not reproduce that drop. But we had changed the sampler, batch and machine as well as the number of questions. We cannot credit the different result to a larger sample alone. The original probe remains a failed probe under its original rules.
DeepSeek exposed a separate problem. Our own 4-bit conversion had already lost
about 9.9 percentage points of GSM8K accuracy before we removed a single
expert. It also produced formatting problems that made its code-test
comparisons unusable. Qwen3's 4-bit baseline, a conversion from mlx-community,
lost about 1.1 points on GSM8K, a much smaller change.
The DeepSeek comparisons still tell us what reducing experts did within that conversion. They do not characterize every 4-bit version of DeepSeek. Validate the model you downloaded or converted before testing an optimization on top of it. Otherwise, two sources of error arrive wearing the same coat.
We also tested code generation with HumanEval. Its 164 tasks gave broad intervals, and the DeepSeek 4-bit formatting failures limited what we could conclude. A result on maths problems is not a substitute for testing the code, documents or conversations you actually need.
Tokens per second is not time to a useful answer
The fixed-length tests isolate generation speed, but a model can change how much it writes when its computation changes. Across the reduced settings, our models generated roughly 2–11% more tokens on GSM8K.
Consider 4-bit Qwen3 at eight versus six experts:
| Measurement | Change |
|---|---|
| Fixed-length batch-eight throughput | 12.8% higher |
| Decode time for the same amount of output, implied by that ratio | 11.4% lower |
| Tokens generated in the full GSM8K run | 5.2% more |
| Elapsed time for the full GSM8K run | 8.3% lower |
The second row uses the reciprocal: . The final two rows come from a different workload, with variable-length answers, prompt processing and batches in which finished rows wait. Each quality setting had one full run, so its wall time is indicative rather than a repeated timing result.
Longer answers plausibly account for some of the difference, but we did not separate their effect from the other workload changes. What matters for a local assistant is visible even without that attribution: a faster stream of tokens is only one part of finishing the job sooner.
Trying the setting in MLX
With mlx-lm 0.32.0, the model configuration can override the selected expert count when loading.5 For the Qwen3 checkpoint we tested, the relevant change is:
from mlx_lm import load
model, tokenizer = load(
"path/to/your/qwen3-model",
model_config={"num_experts_per_tok": 6},
)
The normal setting for that model is eight. This changes routing without retraining or rewriting the weight files. Use separate fresh processes for your baseline and changed setting, and keep the model files, prompt, generation settings and output limit fixed when timing them.
Then test actual tasks as well as speed. Save the answers, compare them problem by problem and record how much the model generated. A quiet machine matters: our busy M1 Max could not resolve the small speed gain cleanly under the protocol's noise rule, while the idle machines could.
The measurements here cover two models, short prompts and controlled benchmarks. They do not establish quality for long-context coding agents or open-ended conversations. The practical experiment is small enough to try; the evidence needed to keep the setting depends on the work it will do.
Read the full technical report (PDF, 10 pages, 294 kB) for the methods, complete results and limitations. The reproduction repository contains the protocol, runnable code and individual measurement records.
Conclusion: fewer experts, still useful answers
Reducing the expert count worked as a speed adjustment in MLX. On our M5 Max, the first reduction made single-request 4-bit generation about 8–11% faster. Qwen3's full maths test showed no statistically detectable loss at that point; DeepSeek's showed a loss of about four percentage points.
Counting bytes helped us see what work could disappear. Measuring showed how much speed actually arrived. Reading and scoring the answers decided whether that speed was useful. None of those steps could stand in for the next.
For someone building a local assistant, this is a setting worth experimenting with, not one to turn down automatically. Keep it when your own tasks support the tradeoff. The goal is not to consult the fewest experts. It is to spend less time waiting for an answer you can use.
Footnotes
-
Chen et al., You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs, September 2026. The average retention figure describes their evaluation, not every model or task. See the paper for its checkpoints, evaluation and serving configurations. ↩ ↩2
-
Magnus Lundstedt, Fewer Experts per Token on Apple Silicon: Speed Opportunity, Realised Speedup and Quality Cost, October 2026. See the report linked above and the protocol and measurements at
report-2026-10, commit902fb96. Paired GSM8K intervals use 10 000 bootstrap resamples. ↩ -
Apple Machine Learning Research, Exploring LLMs with MLX and the Neural Accelerators in the M5 GPU, November 2025. ↩
-
For our own DeepSeek 4-bit conversion, we derived storage from the bf16 header using 4.5 bits per weight for matrices quantized by mlx-lm, retaining bf16 storage elsewhere. The converted model measured 4.503 bits per weight overall. See section 3.5 of the technical report. ↩
-
The mlx-lm 0.32.0 loader accepts
model_config; its model-loading path merges those values into the checkpoint configuration before constructing the model. ↩