What we learn along the way

Notes on engineering, research and building a company.

← All articles

M6 Mac mini, first look: the two speeds of local AI

Our M6 Mac mini arrived on release day. In two local LLMs, processing the prompt gained more than generating the answer. Here is what changed against our M5 MacBook Air.

Chapters

There are two familiar pauses when you ask a local language model a question. First, you wait while it processes your prompt (prefill). Then you watch the answer appear, a few pieces at a time (decode). A faster computer can shorten both waits, but not necessarily by the same amount.

Our M6 Mac mini arrived on 22 September, the new model's first day of availability. We had a benchmark ready and ran two small language models through it, from short prompts to 32 768 tokens of context.1

The question was practical: where does the new machine save time? A model inside a coding agent may read source files, write a patch, run tests and read the results before trying again. There can be long inputs and long outputs, with both waits repeated along the way. If we want local AI to help us build things, we need to understand both.

In our first measurements, the mini's prompt-processing rate was 1.85 to 3.37 times that of our M5 MacBook Air. Its answer-generation rate was 1.41 to 1.63 times the Air's. Those ranges cover two particular models and five prompt lengths, rather than every workload either machine can run.2

Two waits behind one speed number

Consider Qwen3.5-4B with a 2 048-token prompt. Tokens are the pieces of text a model works with, often words or parts of words. We asked it to produce exactly 128 new tokens on each machine.

Measured phaseM5 AirM6 mini
Read the prompt10.95 s4.67 s
Next 127 tokens11.70 s7.82 s

These are mean measured durations. The second row covers the 127 intervals between 128 output tokens. Model loading is excluded, and this is the benchmark's timing, not a complete chat application's response time.3

Processing the existing prompt is called prefill. It builds the model's state from the text already supplied. Generating the continuation is called decode: each new token depends on what came before it, including the token just produced.

Prefill accounts for much of the first wait in this experiment. In an application, time to the first visible token can also include loading, tokenising text, queueing and other work. Decode throughput describes how quickly the answer continues after generation has started. Neither number alone tells us how long the whole interaction takes.

For this Qwen case, the M6 mini processed the prompt at 438.97 tokens per second, compared with 187.14 on the Air. Decode was 16.24 versus 10.86 tokens per second. The throughput gains were about 2.35× for prefill and 1.50× for decode. A single speedup would hide which wait became shorter.

What changed between our two Macs?

Our mini has an M6 with 12 GPU cores and 32 GB of unified memory. The control is an M5 MacBook Air with 10 GPU cores and the same memory capacity. We used the same model packages and input fixtures on both.45

This compares two complete systems. Their cooling and GPU configurations differ, and the recorded macOS builds differ too. The measurements cannot tell us how much of the gain comes from any one of those differences. They do tell us what happened when these two machines ran the same benchmark.

Two charts on the same ratio scale compare M6 Mac mini throughput with M5 MacBook Air throughput. Prefill ratios range from 1.85 to 3.37; decode ratios range from 1.41 to 1.63. Both models are shown with session-bootstrap uncertainty bands.

M6 mini throughput divided by M5 Air throughput. The line at 1× means equal rates. Both panels use the same vertical scale; shaded bands show descriptive 95% session-bootstrap intervals.

The prefill point estimates exceed the decode point estimates at every tested prompt length. The size of that difference varies. In particular, MiniCPM5-2B prefill varied substantially between the M5 sessions, which widens the interval around its comparison. The largest ratio, 3.37×, has an interval of roughly 2.51× to 4.11×. The band is part of the result.

Three sessions on one machine do not describe the variation across all machines of that model. These intervals help show repeatability within our test; they are not a guarantee of what another person's Mac will deliver.

Across these two models and five prompt lengths, our M6 mini delivered 1.85–3.37× the prefill throughput and 1.41–1.63× the decode throughput of our M5 Air.

A longer prompt changes the picture

We tested 256, 512, 2 048, 8 192 and 32 768 prompt tokens, always generating 128 tokens afterward. Every request started with a fresh prompt cache. That keeps the comparison focused on processing the supplied context, without a previous request doing some of the work for it.

The two models were Q4 packages of MiniCPM5-2B and Qwen3.5-4B. Q4 refers to four-bit quantised weights in the package; it does not mean every value or operation uses four bits. The exact conversion and package identity matter when reproducing the numbers.5

M6 Mac mini prefill throughput for MiniCPM5-2B and Qwen3.5-4B across five prompt lengths. MiniCPM peaks near 866 tokens per second at 512 tokens and reaches 369 at 32K; Qwen peaks near 442 and reaches 331 at 32K.

Prompt processing on the M6 mini, batch size one. Points are equal-weight means of three session means. Each model has 12 measured requests per prompt length, except 32K, which has six. The narrow shaded intervals remain present even where they are difficult to distinguish from the lines.

MiniCPM processed short prompts more quickly than Qwen in this setup. At 32K, their prefill rates were closer: 369.04 and 330.53 tokens per second, respectively. Both rates were lower than at 2K. There is more text to process, and in these measurements each additional token also takes more time on average.

The answer-generation curves change with context too:

M6 Mac mini decode throughput declines as prompt length increases. MiniCPM5-2B falls from 35.83 to 20.25 tokens per second between 256 and 32K prompt tokens; Qwen3.5-4B falls from 16.58 to 12.62.

Decode throughput across 128 generated tokens, measured from the first token to the last. The first token's production is outside this interval; the numerator is the remaining 127 tokens. Session and repetition counts match the prefill figure.

On the M6, both models generated tokens more slowly with longer prompts. MiniCPM's decode rate fell from 35.83 to 20.25 tokens per second between the shortest and longest prompts. Qwen's fell from 16.58 to 12.62.6

That is a useful reason to test the context you intend to use. A short chat prompt and a document-sized prompt can give different impressions of the same model on the same computer. These measurements do not isolate the contributions of attention, cache traffic, architecture or runtime choices, so the curves alone cannot identify a single cause.

They also say nothing about which model gives the more useful answer. We fixed the output length to measure timing, not to grade the responses.

Would a larger chunk help?

The falling prefill advantage raised a question: were our 2 048-token chunks holding the M6 back? Chunk size controls how much new input is processed at once. The earlier context remains available; a 2K chunk does not impose a 2K context limit. Here, K means 1 024 tokens.

We followed up with a separate prefill experiment on both machines, keeping the prompt at 32 768 tokens and using the original timing boundaries. Here are the M6 results, in tokens per second:7

Chunk sizeMiniCPMQwen
2K369.0328.3
4K373.3315.5
8K372.7282.9
16K367.4Buffer limit
32K359.0Buffer limit

MiniCPM was fastest with a 4K chunk, only 1.2% above this experiment's 2K baseline. Qwen was fastest at 2K and slowed as chunks grew. Its two largest chunk sizes failed on both machines: the runtime requested individual buffers larger than Metal allowed. The full 32K prompt still worked with smaller chunks.

Across the completed primary comparisons, M6 remained about 1.96–2.12× as fast as M5. A diagnostic that skipped intermediate output-head evaluation, the step that turns model state into token scores, did not uncover a larger relative advantage either. Larger chunks therefore did not recover the short-prompt speedup in this panel. The experiment supports keeping 2K as a reasonable baseline, while leaving the underlying cause of the context-dependent ratio open. The original measurements and figures above remain unchanged.

How we measured

For the original context sweep, we used a timing harness around stock MLX and MLX-LM, with controlled model packages and fixed token inputs. The arrival environment used MLX 0.32.2 and mlx-lm 0.31.3. The harness evaluates the full prompt in chunks before sampling the first output token, then records subsequent token events.8

SettingValue
ModelsMiniCPM5-2B and Qwen3.5-4B, affine Q4, group size 64
Prompt lengths256, 512, 2 048, 8 192, 32 768 tokens
Generated length128 tokens per request
Batch size1
KV-cache policy16-bit
Prompt chunk size2 048 tokens
Sessions3 per machine
Measured repetitions4 per case and session; 2 at 32K
Warmup1 per case and session, excluded from results

Each machine completed 108 measured requests and 30 warmups. Each plotted rate is the equal-weight mean of the three session means. The uncertainty bands come from 10 000 bootstrap resamples of those session means; the ratio calculation resamples the two machines independently.8

The input tokens are controlled fixtures, while the timing comes from real executions. Keeping input length repeatable helps us compare performance. It does not turn the exercise into a test of document understanding or answer quality.

We also checked the path from a recorded event to a plotted number. For Qwen at 2K on the M6, we read the first and last token timestamps in all 12 measured requests and recalculated each decode rate as 127 / (last_token_time - first_token_time), with time in seconds. Averaging within sessions and then across them reproduced the plotted 16.24 tokens per second. This checks the arithmetic behind that point; it is not an independent rerun of the hardware experiment.9

The evidence archive (ZIP) contains the measurement tables, timing-method excerpts and a Python script to check the plotted means and ratios offline.

What this first look leaves open

This is a small-model, single-request test. It does not establish the performance of larger models, several concurrent users, reused prefixes or a complete coding agent. It also does not explain the hardware mechanism behind the improvements.

We are separately investigating FP8 execution through Metal, Apple's GPU API. An initial numerical check shows that the requested software route can execute correctly on this system. It does not establish native FP8 hardware execution or a performance benefit. Those questions need their own measurements and execution evidence.10

Where a mini fits in an agentic workflow

A coding or terminal agent turns those two waits into a loop. Source files, instructions and tool results become input; reasoning, code and tool calls become output. As the work grows, so can the context. Our 128-token responses sample one part of that loop. They do not tell us how long a substantial code change would take, or whether the model would get it right.

There is another variable: how much must the model read again? An inference engine can reuse cached state for an unchanged prefix. LM Studio's own tests show why that matters: on the same M3 Max, mlx-engine 1.8.5 completed a repeated image-prompt request about 3.5 times as fast as version 1.7.0, after improving cache reuse. That is a software-version comparison on a different workload and machine from ours, but a useful reminder that software can change the wait between turns. Our fresh-cache tests leave that question open.11

For demanding local coding, memory capacity and bandwidth deserve separate attention. The model, context cache and working buffers need room alongside the editor and other tools. Higher bandwidth can help generation when moving data is the bottleneck, but the model and runtime still matter.12 JetBrains' current Junie Local setup, for example, specifies at least 64 GB of RAM for its selected Qwen3.6-27B model, beyond our mini's 32 GB. That is a requirement of that setup, not a minimum for every local coding task.13 A larger codebase does not all have to fit in the prompt; what the agent selects and retains matters too.

An always-on mini can also have a more focused job: helping sort incoming mail, drafting replies, searching documents or coordinating tools. Hosting that agent does not require every model to run on the same computer. OpenClaw, for example, connects messaging services to agents using local or cloud models, and supports local embeddings for memory search.1415 MLX-Audio provides local speech recognition and synthesis on Apple silicon.16 These make a hybrid setup possible: search, speech-to-text and text-to-speech on the mini, with a cloud model handling more demanding coding or reasoning. Any text sent to that API still leaves the machine.

That makes the mini interesting as both a place to experiment with small models and a host for selected parts of a larger system. We have measured the two text models in this article; we have not benchmarked those agent or audio workflows.

Conclusion: measure the wait you care about

Our first M6 mini results show improvements in both phases of local inference against the M5 Air we tested. Prompt processing gained more in the point estimates, while answer generation improved by a smaller amount. Longer prompts changed both the absolute speed and the comparison.

For an agent, the useful next measurement is a completed task: reading the relevant files, making a change, checking it and recovering from mistakes. Measure long-context input and sustained output together, with the cache policy and model visible. Include whether the result is correct. Faster tokens only help if they move the work forward.

Our results give us a starting point for smaller local models and focused tasks. Larger models, longer retained context or several agents may call for more memory and higher bandwidth. A hybrid setup offers another route, keeping useful pieces on the mini while a cloud model handles the work that exceeds its local capacity.

In our 2K Qwen example, the prompt-processing wait fell from nearly eleven seconds to under five. The words that followed arrived faster too, but by a different margin. The next question is how much of that saved time reaches someone waiting for a working change, a useful reply or a finished task.

Footnotes

  1. Apple, Mac mini with M6 and M5 Pro. Announced 25 August 2026; customer availability began 22 September. ↩

  2. Precisit, frozen M6/M5 comparison data, F13. ↩

  3. Mean durations recalculated from the measured Qwen 2K request receipts for the M6 run and M5 control. These are means of durations, not reciprocals of mean throughput. ↩

  4. Device records for the M6 Mac mini and M5 MacBook Air. Both report 32 GiB RAM and low-power mode off. macOS builds are 26A428 and 26A5388g, respectively; the mini had no attached display, while the Air reported one. ↩

  5. Model-package verification receipt. Supported Q4 matrices preserve the source carriers; the Qwen package includes 24 Conv1d compatibility adaptations that materialise represented values. These controlled packages should not be assumed equivalent to any other Q4 download. ↩ ↩2

  6. M6 absolute throughput and duration data: prefill, F05 and decode, F06. Article figures are regenerated from these files and F13, retaining the recorded intervals. ↩

  7. Precisit, matched 32K prefill chunk experiment and follow-up protocol, 23 September 2026. Same model packages, inputs and 16-bit KV policy on both machines; three sessions, two measured repetitions per primary cell and one per diagnostic cell per session. Diagnostics cover the 2K and 32K endpoints. Completed cells preserved the final top-1 token against each machine's 2K reference; maximum absolute final-logit error was about 0.000113. Qwen's failed configurations requested single 32 GiB and 64 GiB buffers against a per-buffer limit of about 20.1 GB. This is a separate prefill-only experiment, not a replacement for the original context sweep or an evaluation of answer quality. ↩

  8. Timing and analysis implementation, experiment settings, and recorded runtime versions. Runtime model loading is outside the timed requests. The full-prompt timing policy is full_prompt_last_logits_v2. ↩ ↩2

  9. Independent recomputation of the Qwen 2K decode cell. The mean is 16.24401238410306 tokens per second, matching F06 exactly; this was also recomputed during article review. ↩

  10. Summary of the arrival FP8 numerical check records native_fp8_attribution: unknown and measured_benefit: false. ↩

  11. LM Studio, Improving LM Studio's MLX Engine for Agentic Workflows, 5 June 2026. Its repeated-image-prompt comparison uses Qwen3.6-27B-MLX-4bit on an M3 Max with 36 GB RAM. The second request takes 23.79 s with mlx-engine 1.7.0 and 6.88 s with 1.8.5, which restores most of the cached prompt. This is an application-engine comparison, not a comparison with our M6 measurements. ↩

  12. Tom's Hardware, Apple silicon LLM inference tests on the M4 Max, 30 July 2026. Its measured generation advantage over other unified-memory systems varies across three model architectures despite the same hardware bandwidth advantage. Different hardware, models and runtime make this context for interpreting specifications, not a validation of our speedups. ↩

  13. JetBrains, Junie Local prerequisites, consulted 23 September 2026. The documented setup uses Qwen3.6-27B-4bit with a speculative-decoding draft model and requires M5 or newer Apple silicon and at least 64 GB RAM. ↩

  14. OpenClaw, What is OpenClaw? and where data lives. The agent gateway can run locally while requests go to a remote model provider. ↩

  15. OpenClaw, memory search overview. Embedding-provider options include local GGUF, Ollama and LM Studio as well as remote APIs. ↩

  16. MLX-Audio, the project's documentation for speech-to-text and text-to-speech on Apple silicon. Capabilities are documented here; performance on our mini was not measured in this experiment. ↩

Keep readingSmall weights, fast arithmeticFewer experts, faster answers on a Mac? ← All articles
100%