M6 Mac mini first look: public evidence for the Precisit article

This focused export supports the article's C01 context sweep and C05 chunk-size
follow-up. It is an audit dataset, not the full internal research repository or
a complete hardware reproduction kit. No model weights are included.

Files:
  prefill.json, decode.json, comparison.json: unchanged F05/F06/F13 figure data.
  measurements.csv: C01 per-request timings, including 108 measured requests
    and 30 excluded warmups per machine. No hostnames or generated token text.
  environment.json: the two systems, relevant runtime wheel identities and
    shared timing settings. These are two complete systems, not isolated chips.
  model-packages.json: exact package identities and compatibility qualifications.
    Do not assume arbitrary community Q4 packages reproduce these results.
  independent-cell-check.json: the twelve-request Qwen 2K M6 arithmetic check.
  chunk-comparison.json: C05 rates, session means and capacity outcomes. The
    primary and diagnostic timing policies are explicitly distinguished.
  method-excerpts.txt: exact timing/statistical functions from the frozen source.
  fp8-scope.json: numerical-check scope only, not a speed or native-hardware claim.
  verify.py: standard-library audit of hashes, timing arithmetic and all twenty
    C01 plotted throughput means/ratios. Run: python3 verify.py
  provenance.json: original source revisions/hashes and public file hashes.

Method:
  Two controlled Q4/group-64 models: MiniCPM5-2B and Qwen3.5-4B. Fixed token
  fixtures of length 256, 512, 2048, 8192 or 32768. Batch one, 16-bit KV policy,
  fresh cache per request, stock MLX 0.32.2 / mlx-lm 0.31.3. Model loading is
  outside timing. Prefill evaluates every 2048-token chunk's last logits and
  cache. Decode covers the 127 intervals between 128 generated tokens.
  Three sessions, four measured repetitions per cell (two at 32K), plus one
  excluded warmup per cell/session. Means weight sessions equally. Descriptive
  95% intervals use 10000 session-bootstrap draws, independently by device for
  ratios. See method-excerpts.txt for seeds and exact implementation. These
  intervals do not estimate variation across all units of either Mac model.

  C05 retains the full 32768-token prompt and sweeps chunks of 2048, 4096, 8192,
  16384 and 32768. Primary cells use the original timing policy; diagnostics
  omit intermediate output-head evaluation at the endpoints. This follow-up
  does not replace C01. Neither experiment measures answer quality or complete
  agent task latency. Larger chunks did not recover the short-context speedup.

The source inventory names frozen internal paths for provenance; it does not
require or promise access to that repository. The files in this download are
the reader-accessible evidence. The wider research project remains private.
