What we learn along the way

Notes on engineering, research and building a company.

← All articles

Small weights, fast arithmetic

A compact model file does not determine the arithmetic a GPU performs. Measurements on M5, M5 Max and M6 show why that distinction matters for MLX prefill.

Chapters

Making a language model smaller can turn “this does not fit” into something you can run on your own computer. Then comes another wait: you give it a long prompt, and the first word takes its time.

The file size tells us how much space the weights occupy. It does not tell us what numbers the GPU multiplies. A four-bit weight can leave memory in a compact package and arrive at the multiplication as a 16-bit floating-point value.

That distinction led us to compare Apple's matrix-operation paths on an M5 Air, an M5 Max and an M6. We wanted to understand which arithmetic could accelerate prefill, the processing of the input before generation begins, and whether that advantage survived inside a model. Decode, generating the subsequent tokens, needed its own test.

The useful idea is that storage and compute formats can be chosen separately. The difficult part is paying for the conversion without losing the speed gain or damaging the model's predictions.

The representation changes on the way

Our earlier packing experiment asked whether a model could still predict well with the values preserved by a compact weight format. Once those values are worth computing with, there is another choice: how should the runtime present them to the multiplication?

The quantized prefill kernels we inspected in MLX 0.32.3 unpack weight codes, apply their scales and biases, and place reconstructed floating-point values in a small working tile. On the inspected NAX (Neural Accelerator) path, those values feed Apple's MPP (Metal Performance Primitives) matrix operation. The packed checkpoint remains compact; a tile is expanded when it is needed.1

We built an eight-bit working representation in an MLX fork. Its int8 path folds the packed weight's group scale and bias into an integer operand, using one weight scale per output row. Activations are quantized in a separate pass. A second path stages MX-format weights into FP8 (eight-bit floating-point) operands. Both expand small working tiles while keeping the weights compact in device memory.2

This matters because the useful arithmetic changes with the GPU generation. M5 introduced a Neural Accelerator in each GPU core.3 Our M5 probes show an advantage for int8 over fp16 (16-bit floating point).4 On M6, plain FP8 reaches more than twice the same-script fp16 throughput.5 The opportunity is to turn that extra arithmetic capacity into shorter prompt processing while keeping the model compact. Which working format gets us there?

An illustrative four-bit weight is followed from packed storage through a temporary tile to multiplication. A 64 by 64 tile contains 2 KiB of codes. Stock MLX reconstructs an 8 KiB payload of 16-bit values; the fork's int8 path prepares a 4 KiB payload and separately quantizes activations. For the example weight 0.042, fp16 represents 0.0419922 and scaled int8 represents about 0.0418. The FP8 prototype would represent about 0.0430; the implemented fork uses FP8 for MX weights instead. Stock and int8 operands enter the matrix operation from registers, while FP8 operands are read from memory. Accumulators stay in registers, and the staged paths rescale before storing.

Figure A. The same stored code can reach the multiplication in different forms. Our fork implements the int8 route for affine weights; the FP8 numbers illustrate prototype arithmetic on those same weights. The implemented FP8 route takes MX weights. Matrix dimensions and traced values are illustrative. Tile sizes show value payloads; the inspected stock tile allocates 9,216 bytes with padding.2

This is also more than replacing a type name. In the tested MPP interface, FP8 format operands cannot be cooperative input tensors held in registers. They must be tensors over device or threadgroup memory. A cooperative destination for the accumulated result is still possible. That constraint shapes the implementation.6

First, measure the arithmetic paths

We tested operand combinations through mpp::tensor_ops::matmul2d. A small matrix kernel lets us ask a narrower question than a complete model: how fast does this operation run with this representation, shape and tile?

For the integer paths, the best successful cells at the same matrix shape gave these ratios against the probe's own fp16 multiplication:4

OperandsAirMaxM6
fp16 × fp161.00×1.00×1.00×
int8 × int82.03×1.87×1.98×
int8 × int4b1.76×1.53×1.73×

Table 1. Integer-path throughput relative to this probe's fp16 × fp16 baseline.

The shape here is M×K×N=2048×4096×4096M \times K \times N = 2048 \times 4096 \times 4096, for an M×KM \times K matrix multiplied by a K×NK \times N matrix. These are short probe runs, taking the best tested tile for each operand pair. They are not model speedups.

The denominator matters. Our fp16 probe is not stock MLX's best dense kernel. A result near two against that probe cannot be read as “twice as fast as MLX”. Matrix size matters too. On the M5 Max, the int8 × int8 gain over the fp16 probe was 1.87× with 2048 input rows and 1.94× with 4096 rows.4

Nor are all four-bit operands interchangeable. The mixed int8-by-int4b probe outperformed the fp16 probe too. That is a possible execution path, not evidence that an arbitrary four-bit checkpoint can be passed to it unchanged. Its group scales, biases and layout still need an implementation.

We also had to correct the measuring instrument. An early floating-point probe used multiply-accumulate with output buffers recycled by MLX. Old results could be added into later calls, producing apparent mixed-format anomalies. The corrected sweep uses multiply mode for independent products. The old failed cells remain in the record; they do not enter the comparisons below.7

A corrected sweep includes fp16 and FP4 (four-bit floating point) in the same script on all three machines. At the same matrix shape, it gives these ratios against its own fp16 lane:5

OperandsAirMaxM6
fp16 × fp161.00×1.00×1.00×
Plain FP81.08×1.01×2.08×
Transposed FP81.11×1.01×2.17×
MXFP80.81×0.85×1.13×
Plain FP40.93×0.89×1.08×
fp16 × FP40.94×1.04×1.09×

Table 2. Floating-point-path throughput relative to the corrected sweep's fp16 × fp16 baseline.

The plain FP8 and FP4 rows multiply two operands of the named format: E4M3 uses four exponent bits and three mantissa bits; E2M1 uses two and one. All these lanes accumulate into fp32 (32-bit floating point). On M6, plain FP8 reaches 2.08–2.17× the same-script fp16 rate. Four-bit floating point stays much closer to fp16. Fewer bits alone do not select a faster arithmetic path. These are best-tile burst probes, separate from the integer table above and from a complete model.

Bring the gain back to the model

Our W8A8 (eight-bit weights and eight-bit activations) model experiment uses integer operands, an integer accumulator, and a scaled bf16 (bfloat16) output. We tested it against stock four-bit Qwen3-8B, using the same input tokens in paired runs. “W8A8” describes the converted operations, rather than asserting that every model tensor is eight-bit.8

Three side-by-side panels, M5 Air, M5 Max and M6, plot W8A8 prefill throughput relative to stock four-bit against prompt length, at 512, 2048, 4096 and 8192 tokens. Circles show three quartet observations in run order; diamonds and bars show the re-analysed log-mean and its indicative 95 percent bootstrap interval. At 2048 tokens the means are 2.002 on Air, 1.844 on Max and 1.420 on M6, and a mark shows the predeclared twofold gate, not met on any machine. Gains narrow at longer prompts. A table under each panel lists every mean, interval and original median.

Figure B. Qwen3-8B, MLX 0.32.3, cache-only prefill with 2048-token chunks. Circles show three counterbalanced timing quartets. Diamonds show the exponentiated mean of their log-ratios; bars show a 95% percentile bootstrap interval for that same estimator. With only three quartets, the intervals are indicative. The gold mark is the predeclared twofold gate at 2048 tokens; the tables repeat every cell beside its original median. These are sustained model runs, including a heat-soaked Air.9

The advantage survived model-level evaluation on all three machines and at every tested prompt length. The re-analysis puts the Air's mean just above two at 2048 tokens, but its original median was 1.998× and the predeclared twofold headline gate failed. A different summary statistic does not overturn that recorded decision.8

There is another boundary around these timings. The harness evaluates the model's key/value cache after each chunk; it does not force evaluation of the returned logits, the scores used to choose the next token. This measures cache-building prefill, not the complete time from sending a request to seeing the first word.9

The weight footprint also changes. In the measured M6 model inventory, W8A8 holds about 8.79 GiB of weights, against 4.29 GiB for stock four-bit. This experiment keeps the converted model weights in eight-bit form. The staged path shown above instead expands a working tile from compact weights.10

Faster is only one of the gates

Perplexity measures how well a model predicts the next tokens in a text. Lower is better. For Qwen3-8B, we measured its change relative to bf16 on four fixed text samples labelled literature, science, math and code. The table shows the largest increase across those samples; “bf16 tail” means retaining the final four decoder blocks in bf16.11

RecipeWorst increase
bf16 baseline0.00%
Stock four-bit+9.81%
W8A8+3.57%
Plain FP8+1.65%
FP8, bf16 tail+0.71%
Stock MXFP8+0.75%

Table 3. Worst perplexity increase relative to the bf16 baseline on the four text samples.

W8A8 improved on stock four-bit in all four samples. It nevertheless failed the stricter bf16 degradation bound on the prose samples; the code sample improved by 0.73%. Both comparisons belong in the result. A better incumbent comparison does not make the stricter failure disappear.12

Plain FP8 uses E4M3, an eight-bit floating-point format. Here, weights have per-output-channel scales and activations have per-row scales; “plain” means that the operands do not carry MPP's block-scale plane. It does not mean that we ignore scaling.13

The primary FP8 recipe missed its own bound of at most a 1% perplexity increase on every sample. Keeping the final four decoder blocks in bf16 brought the worst increase down to 0.71%. A larger evaluation over the same source texts gave a worst pooled increase of 0.56%, with a reported 95% interval of 0.42–0.70%. That strengthens the observation on these texts, while leaving other models, languages and tasks untested.14

On M6, the tuned retained-block recipe ran at 1.283× stock-four-bit prefill throughput at 2048 tokens in the re-analysis, with an indicative interval of 1.273–1.290 under the same three-quartet method. It ran at 0.909× W8A8. The tempting “within 2% of W8A8” comparison belongs to the faster recipe that does not retain those blocks, and cannot be combined with the retained-block quality result.13

Comparing against bf16 makes the gain look larger. The untuned retained-block recipe measured 1.723× bf16 in the re-analysis at that length, but the baseline had a slowdown: evaluating the feed-forward gate and up projections together took longer than evaluating them separately. Leaving their input lazy removed the effect in the diagnostic. We observed the slowdown; its complete scheduling explanation remains an inference. The direct stock-four-bit comparison is the more useful headline here.15

A scale can change the route

FP8 and MXFP8 both contain eight-bit floating-point values, but they present scaling differently. The tested MXFP8 path attaches a plane of UE8M0 scales (unsigned, eight exponent bits, no mantissa bits), shared across blocks of values. Those scales encode powers of two. The matrix operation receives the values and their block scales together.16

That seemed a promising way to use a compact representation. The measured throughput told a more specific story.

Horizontal bars show the corrected v4 burst sweep per GPU core on M5 Air, M5 Max and M6 for six lanes: fp16, plain FP8, transposed plain FP8, MXFP8, plain FP4 and fp16 by FP4. A line marks each machine's fp16 rate. On M6, transposed plain FP8 reaches 2.50 TOPS per core and MXFP8 1.30, beside 1.15 for fp16. The layout-matched MX-to-plain ratios are 0.724 on M5 Air, 0.846 on M5 Max and 0.520 on M6.

Figure C. The scaled right operand requires a transposed layout, so the headline ratios use the transposed plain control. TOPS means trillions of operations per second. Rates are divided by the 10, 40 and 12 GPU cores confirmed in the device inventories. This normalizes core count, not clock or cooling. The grey rows repeat the fp16 and FP4 lanes from Table 2 above. These are best-tile burst measurements without paired confidence intervals.16

The M6 result is the exciting part: plain FP8 reaches 2.40 TOPS per GPU core, or 2.50 with the transposed layout, against 1.15 for fp16. Those are the 2.08× and 2.17× ratios in Table 2. On the two M5 machines, the same FP8 paths stay much closer to fp16. For this tested M6 path, eight-bit floating point offers a substantial arithmetic advantage.16

That gives FP8 staging a concrete purpose. A quantized model can keep compact weights in memory, then expand a working tile into FP8 to reach that faster arithmetic. Expanding the same values into fp16 would miss the advantage seen in this probe. The gain creates room to pay for unpacking and activation quantization; whether it survives those costs is a model-level question. Our fork's MX-weight path puts this storage-to-compute choice into practice.2

On the tested M6 path, adding the scale plane did not share the plain operands' throughput increase. The scaled rates remained much closer across machines per core, and M6's scaled rate was actually the highest of the three. Its smaller ratio reflects a much faster plain-FP8 denominator, rather than an absolute collapse in scaled throughput.

The baseline explains the apparently different MXFP8 numbers. On the Air, Table 2's 0.81× compares MXFP8 with fp16; Figure C's 0.724× compares it with transposed plain FP8.

Even the layout choice changes the quoted comparison: M6's MXFP8 rate is 0.542× the untransposed plain control, or 0.520× the layout-matched transposed control. A ratio needs the name of the thing underneath it.16

This describes the software paths we measured. It does not prove which physical arithmetic unit executed them or that a future implementation must behave the same way. It also does not erase MXFP8's useful quality result in Table 3.

How much room is there for unpacking?

The storage question now returns. How much time can a kernel afford to spend expanding each packed tile into eight-bit operands?

Choosing a compute format after loading packed weights is an established technique. On NVIDIA GPUs, Marlin converts four-bit weight codes into fp16 register fragments for the matrix operation. QServe takes a different route: it expands four-bit weights to int8 and pairs them with eight-bit activations for INT8 tensor-core computation.17 Our question here is which conversion earns its cost on the tested Apple paths, where the eight-bit probes showed a throughput advantage that plain FP4 did not.

To isolate that budget, we compared stock four-bit matrix multiplication with an int8 probe whose operands were already prepared. It estimates the room available for unpacking and staging under those particular tile choices.18

Two matrix shapes show no-staging int8 throughput divided by stock four-bit throughput. At 2048 by 4096 by 4096 the ratios are 2.172 on Air, 1.170 on Max and 1.535 on M6. At 8192 by 16384 by 4096 they are 2.123, 2.081 and 1.301. Open circles show the other tested tiles, a few of them near or below parity, and each row lists the stock and int8 rates in TOPS. These are headroom measurements, not staged-kernel speedups.

Figure D. Both shapes and all three machines are retained. Filled marks are the best tested tile and open marks the other two. The int8 probe excludes staging costs. No confidence interval is available for these cells; the ratios are measured headroom, not theoretical ceilings or achieved staged-kernel gains.

The Max's smaller-shape result is especially useful to keep: this tile grid leaves much less margin there than at the larger shape. A dispatch rule based only on the largest result could choose poorly. These margins belong to the probe implementations and tile choices; a kernel with different tiling can exceed them.

We also separated weight rounding from activation quantization. A weight-only int8 diagnostic found no meaningful perplexity cost on the three tested Qwen3-8B weight artifacts; activation quantization remained the quality constraint. That finding concerns model predictions on the tested texts. The worked example above still shows that an individual weight can change when rounded into the working format.19

Our fork's staged int8 path for affine weights and staged FP8 path for MX weights have passed their declared model gates. Those evaluations test compact-weight staging, with their own quality rules and timing boundaries. They are distinct from the expanded-weight W8A8 experiment in Figure B and the prepared-operand headroom in Figure D.20

For Qwen3-8B, the staged paths reached the following results. Each row compares against the stock path for the same weight artifact. This study's quality rule requires the upper 95% confidence limit on perplexity degradation to stay within +1% in every tested domain. Its reference is each artifact's stock output, whereas Table 3 compares against bf16.

Table 4. Implemented staging in the fork. Prefill is the geometric mean across 512, 2048, 4096 and 8192 tokens. The final column is the worst domain's upper confidence limit, not its point estimate.

Path / weightsHostPrefillPPL upper
Stock baselineEach1.000×0.00%
int8, 4-bitM61.389×+0.920%
int8, 4-bitM5 Air1.461×+0.920%
FP8, MXFP8M61.372×+0.708%
FP8, MXFP4M61.421×+0.776%

Every path in Table 4 cleared its declared prefill floor and quality bound. The two int8 rows share the same quality evaluation. Decode throughput stayed within about 1% of stock, with bit-identical outputs in the measured decode steps. At 8192 tokens, staging added 6.25 MiB of transient scratch for int8, 36.08 MiB for MXFP8 and 4.08 MiB for MXFP4. No staging allocation remained active after prefill.21

An earlier FP8 configuration missed its speed floors: 1.21× and 1.25× against required gains of 1.30× and 1.33×. Work on output stores and dispatch recovered that margin. The measured version sends complete row tiles through an unguarded kernel and handles the remaining rows separately, so a prompt need not end on a tile boundary. The failed configuration remains part of the result: fast arithmetic only helps if preparing and storing its tiles is cheap enough.21

Choose for the phase you are accelerating

Generation makes a different demand on the implementation. In our M6 decode test, stock four-bit produced 33.56 tokens per second, while W8A8 produced 12.61. Plain FP8 in that expanded-weight form was similarly slow. The eight-bit prefill win did not make it the better decoder.10

Compact-weight staging gave us a more useful combination in a serving test. We ran Qwen3.8-27B in MXFP4, group size 32 through oMLX on a 32 GB M6, with FP8 staging enabled for prefill only. Of 497 linear layers, 496 met the staged path's conditions; the output head stayed stock. The switch applied to quantized products with at least 64 rows, while one-row decode steps kept the stock path. Activations used a power-of-two scale per row, and MX weights were expanded a tile at a time. There was no full-model eight-bit weight copy.

Cold prompts of roughly 1950, 7360 and 15,020 tokens ran 1.49–1.52× faster, measured through the HTTP interface. At the shortest length, throughput rose from 306 to 461 tokens per second. Adding roughly 1010 or 2780 tokens to a cached 8K context gave 1.31× and 1.41×; a short continuation took about 0.9 seconds in either mode. Decode remained around 10.2–10.5 tokens per second. The expanded-weight 8B decode regression did not appear in this configuration.22

Quality still needs its own measurement. A small Qwen3.8-27B evaluation, with two segments per domain, found perplexity increases of +1.21% on literature, +0.58% on science, +0.27% on math and +0.18% on code, relative to that model's stock MXFP4 path. No quality gate was applied to this serving study. The 8B quality result cannot be carried over to a different model.23

Keeping separate representations for prefill and decode is one possible response, but its memory cost must be counted. Staging a small tile from compact weights offers a different tradeoff. Each needs its own quality and complete-request measurements.

For a new format, the practical sequence is to check the represented values, identify the arithmetic path, then measure model speed and quality under the same declared boundaries. Test generation separately. The checkpoint's bit width answers how the numbers are stored; the working tile helps decide how quickly they can be used.

A small Christmas wish list

Working this close to Apple Silicon leaves us with a few wishes. Here is what we would be delighted to try in future Apple hardware and APIs. The labels reflect our researcher's assessment of potential value, rather than measured speedups or an Apple roadmap: highest value, high value and useful.

  • Small floats with a bigger compute gain. We would love FP4 and FP6 paths with an arithmetic advantage like the one these probes found for eight-bit operands. The tested FP4 path stayed near fp16 throughput; there is an appealing possibility beyond saving bytes. [Highest value]
  • Transparent compression along the memory path. If weights and the key/value cache could travel in fewer bytes and be decompressed cheaply as they are read, generation would have a different memory budget to work with. That would be a lovely extension of the compact-storage idea. [Highest value]
  • A whole tile in one description. Imagine describing packed weights, activations and optional scales together, and having the matrix operation handle their preparation, tiling and edge stores. More of a kernel could express the calculation itself. [Highest value]
  • Scales that come along for the ride. MXFP8's quality result makes us curious about a path that combines block scaling with plain FP8's throughput. Keeping the scales close to the values could become an even more useful option. [High value]
  • A shorter trip from packed weights to multiplication. Hardware that helps unpack, scale and arrange a tile as it loads would be exciting to explore. Dedicated data movement could give us another way to reduce preparation instructions and synchronization. [High value]
  • More breathing room for generation. The decode results put compact weights back in the picture. More bandwidth or useful cache capacity would give us new ways to explore that tradeoff. [High value]
  • Partial tiles with less fuss. Prompts need not end neatly on a tile boundary. Matrix operations that handle the remaining rows efficiently would make arbitrary prompt lengths easier to support. [Useful]
  • A format map in the box. A concise guide for each generation, covering operand formats, throughput, scales and layout constraints, would help us choose the next experiment. Good documentation would be a lovely stocking filler. [Useful]

Try it yourself

The implementation is available in our public MLX fork. It is experimental fork code; the examples here pin a commit so that a moving branch does not change the experiment. Start in a fresh virtual environment with native Apple Silicon Python 3.12, the version used for these smoke tests. The fork requires Python 3.10 or newer; the system's python3 may be older. For the FP8 path, use an M6, macOS 27 and Xcode 27 with its Metal tools. The same build block was also tested on an M5 Air running macOS 27.

git clone https://github.com/precisit/mlx
cd mlx
git checkout d03fed7b6965864b61c7c10fc48277e38c1a272f
python3.12 -m venv .venv
source .venv/bin/activate
MACOSX_DEPLOYMENT_TARGET=27.0 \
CMAKE_ARGS="-DCMAKE_OSX_DEPLOYMENT_TARGET=27.0 -DMLX_METAL_JIT=OFF" \
  python -m pip wheel . --no-deps -w wheels
python -m pip install --no-deps wheels/mlx-*.whl

For eligible matrix products, MLX_QMM_FP8=1 enables staging from MXFP4 or MXFP8 weights on M6. MLX_QMM_INT8=1 enables the affine-weight path on M5 or later. Leave the relevant switch unset or set it to 0 for the stock comparison. For an int8-only build on macOS 26.2 or later, replace both deployment-target values of 27.0 in the block with 26.2. The build instructions at the pinned commit cover the general compiler and Python requirements.

Download the evidence package (ZIP, 59 KiB) for the selected measurements, method notes, run identities and small comparison script. Paths in the footnotes are relative to its extracted top-level folder. Some original run identities were not recorded; identities.md lists those unknowns explicitly. The package contains selected evidence rather than the full research repositories. Its scripts are MIT-licensed; text and measurement data are licensed under CC BY 4.0. See LICENSE-NOTE.md for attribution.

The ZIP's SHA-256 is 60cdd6302263ae585f86ff75d8e1859b8e57ad933805343372bfb37a78562ea2. Place it in the cloned mlx directory, then verify, unpack and run:

shasum -a 256 apple-matrix-formats-evidence-20261008-r02.zip
unzip apple-matrix-formats-evidence-20261008-r02.zip
python apple-matrix-formats-evidence-20261008-r02/examples/tryit_check.py

The script compares each switch at 0 and 1 on the same seeded matrix inputs, with shape 2048 × 4096 × 4096. The supplied smoke-test outputs show a changed result for int8 on M5 Air and for both paths on M6. FP8 on M5 Air and both switches on the M1 Max fallback machine produced identical results. These output comparisons indicate a changed execution path; the script does not trace the dispatched kernel. Its short timings measure one matrix operation, not the model gains in Table 4 or model quality.24

The switch alone does not reproduce the model recipes above. Those also control which layers are staged, including keeping the FP8 output head stock. A useful comparison holds the model, inputs and layer policy fixed, confirms which path ran, and measures prefill, decode and quality separately.

Footnotes

  1. Source inspection of MLX 0.32.3, commit 64ea011: the dequantization routine reconstructs values using their scales and biases; the block loader calls it to fill its destination; the inspected NAX prefill kernel loads those tiles into its matrix operations. This description concerns these prefill kernels, not every quantized operation or later MLX release. ↩

  2. Evidence package, mechanism/README.md explains the operand contract; mechanism/visual-evidence.json records the source identities in sources and the MX census in mxfp4_census. The fork int8 kernel folds affine weights and rescales its accumulator before storing. The figure distinguishes the affine FP8 prototype from implemented MX staging; the inspected 35-layer mxfp4 census found no changed weight values, rather than proving exactness for arbitrary scale ranges. The stock tile's 8 KiB payload occupies a padded 64 × 72 × 2-byte allocation in the inspected loader. These are source-level details, separate from the model measurements in Figure B. ↩ ↩2 ↩3

  3. Apple's M5 announcement describes a Neural Accelerator in each GPU core. Its M5/A19 developer talk explains MPP and TensorOps, including four- and eight-bit integer tensor support in macOS 26.4. The int8 and M6 FP8 ratios here are our probe results, with their own baselines; they are not Apple's performance claims or proof of physical-unit attribution for every operand format. ↩

  4. Evidence package, campaign/evidence.json, integer_lane_candidates: select the 2048 × 4096 × 4096 entries for M5 Air, M5 Max and M6; selected retains each chosen cell and derived_vs_f16 its ratio. Ratios divide gops_best for the fastest successful operand-pair cell by the corresponding fastest f16*f16->f32 cell at the same shape. The mixed row is the receipt's i8*i4b->i32 lane. The researcher admitted these integer cells in technical review; the legacy float cells are excluded from the figures. ↩ ↩2 ↩3

  5. Evidence package, campaign/followup-evidence.json, mx_sweep_v4[*].lanes and vs_f16, with the implementation in campaign/mx_sweep.py. The same script uses mode::multiply, five tile configurations and the fastest of five timed calls after warm-up. All plotted lanes pass its constant checks. Ratios divide each gops_best by the same receipt's f16*f16->f32 rate. These post-pin additions supersede the earlier floating-point lane comparison; the original snapshots remain in the audit. They are probe ratios, not comparisons against stock MLX's best dense kernel. ↩ ↩2

  6. Evidence package, mechanism/README.md summarizes the memory-input constraint; mechanism/visual-evidence.json, sources, identifies the prototype interface check (probes/p06_fp8_tiles/README.md) at its pinned source hash. The implemented FP8 kernel reads activations from device memory and weights from threadgroup memory, with a cooperative fp32 destination. This is an API constraint, not a claim about physical register circuitry. ↩

  7. Evidence package, campaign/mx_sweep.py and campaign/notes.md (“Probe correction”). Independent products require a fresh result; deliberate accumulation into an initialized destination is a different use of multiply-accumulate. ↩

  8. Evidence package, campaign/evidence.json, w8a8_prefill[*].summary.per_N and summary.g2_headline_2x_at_2048. At 2048 tokens, the original median points and reported intervals are 1.998 [1.989, 2.018], 1.836 [1.833, 1.863] and 1.415 [1.408, 1.437]. All three record g2_headline_2x_at_2048: false. ↩ ↩2

  9. Evidence package, campaign/notes.md, campaign/wp6-intervals/PREREGISTRATION.md and campaign/wp6-intervals/intervals.json (receipts, then recipe and per_N). The WP6 and FP8 harnesses use ABBA/BAAB/ABBA, three quartets per length, with both models resident and matched tokens. The predeclared re-analysis uses exp(mean(quartet log-ratios)), with a percentile bootstrap over three quartets (10,000 resamples, seed 1). The input logs are rounded to four decimals in the original receipts. These indicative intervals coincide with the observed quartet ranges in these cells; they do not support significance or equivalence claims. The original median points, mixed-estimator intervals and gate decisions are retained for audit. Cache states, rather than logits, are forced. Lazy evaluation may also omit final-layer work needed only to produce the logits. The Air's thermal correction, summarized in campaign/notes.md, describes heat soak across FP8 gates; paired within-gate ratios must not be combined as absolute burst rates. Within-run variation does not measure reproducibility across machines or days. ↩ ↩2

  10. Evidence package, campaign/evidence.json, decode.summary, including resident weight sizes. This short test retains forward and reverse arm order; it is not a long-running generation service measurement. It evaluates these implementations, not the best achievable decoder for every eight-bit representation. ↩ ↩2

  11. Evidence package, campaign/evidence.json, quality[0].summary.deltas (original) and quality[1].summary.deltas (tuned retained-block variant). Each sample supplies 16,384 tokens, of which 16,383 are scored, using step 512. Values are 100 * (PPL_arm / PPL_bf16 - 1). These project corpus labels do not imply broad domain benchmarks. Plain FP8 converts linear layers except the output head; embeddings and normalization remain bf16. MLX 0.32.3; MLX_ENABLE_TF32=0 in the FP8 receipts. ↩

  12. Evidence package, campaign/notes.md (“W8A8 quality documentation defect”) and campaign/evidence.json, quality[0].summary.deltas. The original record reports incumbent parity on all four samples and failure of the bf16 prose bound. Its absolute-value notation conflicts with its recorded code pass; the researcher confirmed in review that the operative rule was one-sided degradation, at most +0.1%. That documentation defect is retained in the evidence audit. The later FP8 +1% and MXFP8 +3% rules are separate decisions. ↩

  13. Evidence package, campaign/evidence.json, fp8_model_gates (identify entries by source or summary.ratio); campaign/wp6-intervals/intervals.json, receipts.fp8tL4_vs_stock4, receipts.fp8tL4_vs_w8a8 and receipts.fp8t_vs_w8a8. At 2048 tokens, the re-analysis gives 1.283 [1.273, 1.290] for retained-block FP8 / stock four-bit, 0.909 [0.908, 0.911] for that recipe / W8A8, and 0.981 [0.979, 0.982] for all-FP8 / W8A8. Intervals are indicative, n = 3. The original retained-block / stock median was 1.287. ↩ ↩2

  14. Evidence package, campaign/evidence.json, quality[2].summary.fp8L4.per_domain: 24 segments, 393,192 scored tokens; paired chunk bootstrap within each domain. This includes the original segments and expands the same four source texts. It is not independent held-out validation; the math-labelled text supplies only three segments. The quoted interval belongs to the worst pooled domain, science. ↩

  15. Evidence package, campaign/evidence.json, fp8_model_gates (the fp8L4/bf16 entry) and bf16_pair_diagnostic.cells[1] at 2048 rows; campaign/wp6-intervals/intervals.json, receipts.fp8L4_vs_bf16.per_N.2048. The original median was 1.721; the re-analysis mean is 1.723 with indicative interval [1.716, 1.733]. The pair diagnostic records 42.64 ms together versus 21.70 ms separately, and 20.83 ms with a lazy normalization input. Distinct materialized inputs also show the slowdown; shared-buffer aliasing is not established as its cause. ↩

  16. Evidence package, campaign/followup-evidence.json, mx_sweep_v4, mx_layout_comparison and device_inventory. GPU core counts are 10, 40 and 12, confirmed by the 2026-10-06 inventories. These are later inventory snapshots, not proof that every earlier run used the recorded OS/SDK build. Layout-matched ratios divide MX gops_best by transposed plain FP8: 0.724, 0.846, 0.520. Against untransposed plain they are 0.748, 0.842, 0.542. Neither comparison has a paired confidence interval. The scaled and plain paths also use different relaxed-precision settings in the script; this is a comparison of the tested paths, not an isolated hardware cost of scale bytes. ↩ ↩2 ↩3 ↩4

  17. Marlin's conversion and scale routines prepare fp16 fragments; its inner loop loads packed weights into registers, converts them and calls the matrix operation. The QServe paper (sections 4.1 and 5.2) describes progressive quantization and INT4-to-INT8 conversion for its W4A8 kernels. These are examples of the storage/compute distinction, with their own quantization schemes and workloads. ↩

  18. Evidence package, campaign/evidence.json, staging_headroom[*].shapes, and campaign/notes.md (“Headroom”). Matrix shapes use M × K × N. The probe is already-prepared int8 multiplication; the stock arm includes quantized multiplication. This deliberate asymmetry estimates headroom, not a matched implemented loader. ↩

  19. Evidence package, mechanism/visual-evidence.json, weight_only_diagnostic: the w_row_noH diagnostic compares each artifact with its stock arm on one segment per domain. The small geomean changes do not establish mathematical exactness or zero error on every domain. The pooled study and kernel gates are separate evaluations. ↩

  20. The public fork at d03fed7 contains the affine int8 path and MX-to-FP8 path, including the row-tile dispatch repair. The evidence package, staged-and-serving/model-evidence.json, models, records the K1 and K4 evaluations. sources identifies the verdict and dated addenda. These results have separate quality gates, artifacts and build identities; they do not replace the original measurement pins. ↩

  21. Evidence package, staged-and-serving/gates-and-methods.md and staged-and-serving/model-evidence.json, models: select by machine, weights and arm; fields prefill, quality, decode_ratio, decode_steps_bit_identical, memory and ragged. The affine artifact uses group size 64; MX artifacts use group size 32. The model harness, identified in sources, pairs the paths in-process, with TF32 disabled and cache states forced rather than logits. At each length the point estimate is the exponential of the median quartet log-ratio, as the methods page states; Table 4 takes the geometric mean of those points across the four lengths. This is not Figure B's re-analysis estimator; no aggregate confidence interval is supplied. Prefill floors are 1.320 (int8 M6), 1.389 (int8 Air), 1.30 (MXFP8) and 1.33 (MXFP4). Quality uses eight segments per domain and the worst upper 95% limit relative to each stock artifact; the shared int8 quality receipt was measured on M6. FP8 decode ratios are 0.998 and 0.993. At non-aligned lengths 8191 and 2000, FP8 ratios are 1.330/1.382 for MXFP8 and 1.374/1.423 for MXFP4. The M5 Max int8 entry gave 1.341×, but had no predeclared floor and is reported without a promotion verdict. Transient scratch is the measured incremental prefill allocation, not total model memory. ↩ ↩2

  22. Evidence package, staged-and-serving/serving.md, identities.md (section G) and staged-and-serving/model-evidence.json, serving.cold, serving.followup and serving.decode. oMLX 0.7.0 at 25aebb3e, mlx-lm 94cdcae, Python 3.12.13, macOS 27; fork build bb8189463 plus the K4 patch, matching the two modified files at public d03fed7. Cold runs used source-code prompts with a randomized first line, zero reported cached tokens, one output token, and two off/on/on/off rounds in one loaded server. Ratios divide staged and stock median throughput (ratio_of_tps_medians in the export; speedup_median in the original receipt), not the separately reported quartet ratios. The 1010/2780-token follow-ups followed an 8K cached context; the server memory guard reduced the base requests' prefill chunks in both arms. Decode observations used 200 generated tokens per run. The serving benchmark left TF32 unset; the quality evaluation disabled it. These are serving observations, not the frozen model gates in Table 4. The layer wrapper, not the global flag alone, excludes the output head. ↩

  23. Evidence package, staged-and-serving/model-evidence.json, serving.quality.per_domain: two segments and 32,766 scored tokens per domain. Literature's +1.21% point estimate has a reported interval of [+0.94%, +1.48%], which crosses +1%. This small, ungated reading does not establish general quality equivalence. ↩

  24. Evidence package, examples/tryit_check.py, examples/tryit_check_m5-air.txt, examples/tryit_check_m6.txt, examples/tryit_check_hub.txt and examples/host-identities.txt. The researcher built the pinned fork with Python 3.12.13, target 27.0 on M5 Air/M6 and target 26.2 on the M1 Max. Each arm has two warm-up calls and five timed calls; the script tests int8 with affine four-bit weights and FP8 with MXFP8 weights. The example is a single-operation smoke test, with output changes used as a route indication. It does not implement the serving wrapper, verify model quality or establish kernel dispatch by tracing. The checks were run by the researcher; no new GPU measurement was made during editorial review. ↩

Keep readingDoes your weight format deserve a kernel?M6 Mac mini, first look: the two speeds of local AI ← All articles
100%