A game-playing AI in 1.6 MB
Training with ternary (just three weight values) brought a new model down to 1.6 MB. We compared Base243 and T34 packing, with training and conversion telling different stories.
Chapters
Before an AI can make its first move in a browser, it has to arrive. A quick decision is useful once the model is running, but the player still has to download the numbers that make that decision possible.
In our first Connect Four experiment, we trained a small model to read the board, score the legal columns and choose a move in one forward pass, without searching future positions. Its int8 model file was 7.8 MB. In the follow-up, we made the runtime smaller and faster. That left the weights as the obvious next question: how much of the download could we remove before the opponent forgot how to play?
A new model with the same architecture and training data fits in 1.59 MB and retains a similar level of play in our tests. Most of its weights are ternary: −1, 0 and +1, multiplied by a shared scale. Training with that restriction worked. Applying it to the finished model was much less successful.1
With a longer final training stage, the T34 model scored 93.0% against a four-move search bot and 91.0% against a six-move bot, counting a draw as half a win. On held-out positions, 98.9% of its moves preserved the game's exact outcome, compared with 98.7% for the original fp32 model. Repeating the training helped us put those results in perspective.
The difference is useful beyond the game: a storage format can be something a model learns to live with, rather than something imposed once learning is over.
Why use a game as a measuring stick?
Connect Four gives us an unusually clear answer to whether a decision was good. The game is solved. An exact solver can tell us whether a position is a win, draw or loss with perfect play, and whether a move preserves that outcome.
That made it possible to ask a precise question about the smaller model. Instead of judging whether a sentence sounded convincing, we could inspect 17 325 held-out board positions and then let the models play complete games. The training labels came from an exact solver too, building on Benjamin Rall's MIT-licensed connect-four-ai. The model learned from those labels; it did not discover the game by itself.
If the only goal is a strong Connect Four opponent, that solver remains a good choice. Its 1.26 MB WebAssembly build is smaller even than our ternary model file. The experiment is about learning compact decisions, with a game whose answers we can check. Other tasks may offer examples of good decisions without an exact solver available at runtime.
Three values, packed two ways
A weight normally stores a number with many possible values. Ternary weights reduce that choice to three. Within each group of 128 weights, a shared scale turns the stored −1, 0 and +1 into the values used by the model. A scale of 0.08, for example, gives −0.08, 0 and +0.08.
The same three-value idea appears in BitNet b1.58.2 We tested two ways of arranging and packing those choices; this was not a comparison with BitNet's training recipe.
Base243, our name for the base-3 packing used in llama.cpp's TQ1_0 format, permits any combination of the three values. Five ternary values have possible combinations, so they fit in a byte's 256 states. Our 128-weight groups use 26 bytes for the codes, including padding, and two bytes for an fp16 scale: 1.75 bits per ternary weight, including its share of the scale. The five values use 243 of a byte's 256 possible values, leaving 13 unused. That is unused encoding capacity, rather than a whole spare bit.
T34, our shorthand for Sherry's 3:4 format, adds a restriction: in each group of four weights, exactly one is zero and the other three are either −1 or +1. There are four places for the zero and eight possible sign patterns. That gives states, which fit in five bits. With the shared scale, a 128-weight group occupies 22 bytes: 1.375 bits per ternary weight.3
Neither file is entirely ternary. The byte embedding and scoring head's matrices remain int8; positions, biases and normalization parameters use fp16. These smaller parts still take space, as do scales and the ONNX graph. The 1.59 MB headline counts the complete T34 model file, not just its packed codes.
Give the model the constraint while it learns
For both formats, we tried three routes to a small file:
- Convert after training. Start with the finished v2 model and fit its weights to the ternary format, without further learning.
- Fine-tune with ternary weights. Start with v2, but let it adapt while using the restricted weights in its forward passes.
- Train with ternary weights from the start. Start from random weights and use the original data and training schedule, with the same restriction.
Before running that six-model grid, we fixed a modest acceptance test: a file under 2 MB, better play than our small first attempt, and agreement between the exported file and the model we evaluated. All six passed. That was a floor, though. Keeping the stronger v2 model's level of play was the more interesting question.
Training used MLX on Apple silicon. It kept floating-point weights for the optimizer, but replaced them with their quantized values for each forward pass. A straight-through gradient let learning update those underlying weights despite the discrete rounding step. The exported file then used the same quantization rule.4
The fine-tuning runs took 12 000 steps and also learned from v2's distribution of move scores. The original from-scratch runs followed the two-stage schedule, 6 000 then 12 000 steps, using the solver labels without that extra v2 teacher. These are new trained models, not alternative packaging of an unchanged set of weights.
A small file is not enough
Each game comparison used 200 games from an empty board, alternating colours. On both sides, a move was chosen randomly 5% of the time. A win counted one point and a draw half a point. The percentages below are shares of the available points, not necessarily win rates.
| Model | File | bpw | Four-move bot | Six-move bot |
|---|---|---|---|---|
| v2, original fp32 | 29.7 MB | 32.20 | 91.5% | 89.3% |
| v2, original int8 export | 7.8 MB | 8.47 | 90.5% | 87.8% |
| Base243, converted only | 1.93 MB | 2.09 | 54.8% | 51.3% |
| T34, converted only | 1.59 MB | 1.73 | 12.8% | 10.8% |
| Base243, fine-tuned | 1.93 MB | 2.09 | 91.5% | 88.0% |
| T34, fine-tuned | 1.59 MB | 1.73 | 88.5% | 89.5% |
| Base243, trained from scratch | 1.93 MB | 2.09 | 88.5% | 87.8% |
| T34, original from-scratch run | 1.59 MB | 1.73 | 94.5% | 87.0% |
Here bpw means effective bits per model parameter: the complete file size in bytes × 8 ÷ 7 382 528 parameters. It includes scales, the non-ternary parts and graph overhead, so it is higher than the packed-weight figures above.
All four models trained with ternary weights reached the strong-player thresholds used for the original model. Neither convert-only model did. On held-out positions, the trained models preserved the exact game outcome on 98.5-98.7% of moves, compared with 98.7% for fp32 v2. Conversion alone reached 92.0% with Base243 and 86.0% with T34.
Two more game seeds put all four trained ternary models between 86% and 95% against the four-move bot. Those are additional game runs, not independently trained models. Replaying a trained model and training another one answer different questions.
Does it survive another training run?
We trained the from-scratch T34 recipe again with a different random seed. Its held-out result was close: 98.4% outcome-preserving moves, against the first run's 98.5%. Its score against the four-move bot was lower, though: 88.3% rather than 94.5%. The first result was a favourable run, not a number we should expect every training run to reproduce.
We also tested a longer final training stage, planned before that run began. It used the same first stage, then 24 000 steps instead of 12 000 in the second stage. The file remained 1.59 MB.
| T34 training run | Four-move bot | Six-move bot | Outcome preserved |
|---|---|---|---|
| Original schedule, first seed | 94.5% | 87.0% | 98.5% |
| Original schedule, second seed | 88.3% | 89.3% | 98.4% |
| Longer final stage | 93.0% | 91.0% | 98.9% |
More training improved the held-out result and the score against the deeper bot, without adding bytes to the model. The longer-trained model scored 92.5-93.5% against the four-move bot across three game seeds, and 88.5-92.3% against the six-move bot. This is the file behind the article's headline result.5
It still does not establish that T34 is a generally stronger player than v2. Against the perfect player, the longer-trained model scored 42.8%, compared with v2's 47.5%, under the same noisy game rules. The primary four-move scores have overlapping reported 95% intervals: 88.6-95.8% for the longer-trained T34 and 86.8-94.6% for fp32 v2. Only the original T34 from-scratch recipe has been repeated with a second training seed; the longer schedule and the other training recipes each have one run.
The game panel also revisits familiar paths. Against the deterministic four-move bot, only 95-137 of the 200 games were distinct for the trained ternary models, despite the random moves. Those scores describe a narrower set of encounters than the 17 325 held-out positions. The intervals capture variation within the game test, not the uncertainty from retraining or choosing a different opponent.
The closest weights did not make the best moves
The convert-only result left another question. Was learning essential, or had we simply chosen a poor way to map the original weights to three values?
We explored 60 training-free variants, changing the group size, the fitting rule or which matrices became ternary. The most revealing adjustment was the scale, in a follow-up investigation after the first sweep results.
The starting fit minimized squared error between the original and ternary weights. For T34, it zeroed the smallest weight in each group of four and set the shared scale from the remaining magnitudes. That is a sensible answer to a numerical approximation problem. It was a poor answer to preserving the model's decisions.
| Conversion change | Four-move bot | Outcome preserved |
|---|---|---|
| T34, least-squares scale | 12.8% | 86.0% |
| T34, alternative scale, same zeros | 80.5% | 92.2% |
| Base243, least-squares scale | 54.8% | 92.0% |
| Base243, that scale multiplied by 1.2 | 80.5% | 96.5% |
The alternative T34 scale came from a Gaussian Lloyd-Max rule, using each group's root-mean-square weight magnitude. It changed how large the surviving weights were, without changing which ones were zero. The Base243 adjustment was simpler still: make the same scale 20% larger.6
A better approximation of the original weights was not necessarily a better approximation of what the model did. The lesson is to measure the decisions you care about, not only the error introduced into the stored numbers.
The attention matrices that form queries, keys and values were particularly sensitive to conversion. Within the 2 MB budget, the best training-free variant reached 96.5% outcome-preserving moves, against 98.4-98.9% across the trained models. Keeping all attention matrices at int8 brought Base243 to 98.1% without training, but its tensors alone occupied about 3.7 MB. That recovered much of the playing quality while giving up much of the size reduction.
Packed on disk, packed on the GPU
The ternary files run in the same onepass-webgpu runtime as the speed experiment, with a small format plugin. The GPU reads the packed weights and decodes them inside the matrix-multiplication kernels, rather than expanding every matrix into a full floating-point copy first.
For T34, that means extracting a five-bit state, locating its zero, reading the three signs and applying the group's scale. The calculations still use floating-point arithmetic. A compact weight representation does not mean that the browser has become a ternary computer.
We checked the original Base243 and T34 from-scratch exports, and the longer-trained T34, through ONNX Runtime's reference decode. We then compared WebGPU with each file's reference on all 17 325 held-out positions. All three files selected the same column as their own reference on every tested position. This checks that the browser runs the model we evaluated; it does not mean those models always choose v2's move.7
On an idle M5 Pro with Chrome 154 and Metal, the warm timing comparison was:
| Model file | Download | Per move | GPU weights |
|---|---|---|---|
| v2 fp32 | 29.7 MB | 1.3 ms | 29.5 MB |
| v2 int8 | 7.8 MB | 0.9 ms | 7.7 MB |
| T34, original from-scratch run | 1.59 MB | 1.1 ms | 1.96 MB |
| Base243, trained from scratch | 1.93 MB | 1.3 ms | 2.30 MB |
Each row is the median of three run medians, with 500 timed positions after 20 warm-ups per run. These are warm decisions, not page-load or first-move times, and the memory column counts weight buffers rather than all browser or GPU memory.8
The timed T34 file was the original training run. The longer-trained file uses the same format, matrix shapes and kernels, but the table is not a new timing measurement of that file.
The gain here is size, not speed. T34 is about one fifth of the int8 file, but it is slightly slower in this runtime. Its 1.59 MB model is accompanied by roughly 37 KB of runtime code and a 6.5 KB format plugin, before network compression and separate page assets. The smaller file stays small during execution, while extra decoding still has a cost.
The T34 kernel has not yet been tuned for its format. We kept the runtime and kernel structure, changing how the weights are read. Each decoded value is multiplied by its group's scale before an ordinary floating-point multiply-add. With 64 positions scored per call, T34 took 0.38 ms per position against fp32's 0.32 ms, about 18% longer in the same timing panel. Those are batch-throughput figures, not the turn times in the table above.
There are concrete optimizations to try: apply the shared scale once per group, or reuse each decoded weight across more positions. An integer path, like those explored for BitNet inference, is another possibility, but would also quantize the activations and need fresh accuracy checks.9 The measurements describe our current implementation, not a speed limit inherent to ternary weights.
What fits in a smaller opponent?
We started with a practical question: could a browser download much less without turning a capable opponent into an easy one?
For this 7.4-million-parameter architecture, training data and game, the answer was yes. The longer-trained 1.59 MB T34 model kept a similar level of play to our earlier model, preserving the game outcome on 98.9% of held-out positions. The shortcut of converting finished weights lost much more. Adjusting their scale recovered some of that loss; teaching with the restriction in place recovered more.
That suggests a useful order of work for another small decision model: choose a representation you can deliver, expose the model to it during learning, and check the exported file on decisions that matter. Repeat the training too: the difference between our two initial T34 runs was a useful reminder that a good checkpoint and a repeatable recipe are different things. A familiar numerical error metric is only part of the answer.
This experiment covers one model size and one task, with browser measurements on Apple silicon in Chrome. Whether another task or a larger model adapts as readily is a question for its own experiment. Connect Four's solver made this one unusually easy to grade.
A small download can be a design goal from the first training step. Small ternary models could put useful decisions into more games, apps and browser tools, without sending every choice to a cloud model. Connect Four is a starting point. The next interesting question is what else we can teach to fit.
Try the ternary models in your browser and see how much of an opponent fits in a 1.6 MB model, and how fast its inference runs locally. Can you beat it?
Footnotes
-
The six-model comparison, file sizes and game results are in the grid summary, with the concept gate fixed before the first run and raw records available for inspection. The original v2 recipe and evaluation are available in one-pass-specialists; its int8 export and fp32 model are separate comparison rows here. ↩
-
Ma et al., The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits, 2024, describes BitNet b1.58's ternary weights. We have not run a matched quality comparison with its training recipe. ↩
-
Huang et al., Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification, 2026, describes the 3:4 format we call T34. Base-3 packing of five ternary values per byte was implemented by compilade in llama.cpp PR #8151 for BitNet b1.58 and TriLM models in 2024. The codec implementation is available alongside the tests. Our group size and scale accounting are specified above; using these packing designs does not imply using the papers' complete training methods. ↩
-
Training uses MLX. For the underlying techniques, see Bengio, Léonard and Courville, Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation, 2013, on straight-through estimation, and Hinton, Vinyals and Dean, Distilling the Knowledge in a Neural Network, 2015, on learning from a teacher's outputs. See our MLX trainer and the scripts for the second seed and longer stage. ↩
-
The raw evaluation records include the second training seed and longer final stage. Percentages are rounded directly from raw game scores: 0.8825 becomes 88.3%, 0.8925 becomes 89.3%, 0.9225 becomes 92.3% and 0.4275 becomes 42.8%. The longer-trained export contains 1 593 727 bytes; its SHA-256 in the verification record matches the T34 file on Hugging Face. ↩
-
The scale-fitting method builds on Lloyd, Least squares quantization in PCM, 1982, and Max, Quantizing for minimum distortion, 1960. Our scale investigation was a follow-up to the first sweep, not part of the original six-model grid. See the sweep plan and dated addendum and sweep results. ↩
-
The reference decode is stored as an ONNX model-local function, which standard ONNX Runtime can execute. Full-position WebGPU parity was recorded on an M1 Max, separately from the M5 Pro timing run. The original Base243 and T34 exports and the longer-trained T34 each matched their own decoded references on 17 325 of 17 325 choices. See the WebGPU parity record and export verification code. ↩
-
See the public speed protocol and our preceding speed article. The frozen 28 September timing record uses runtime revision
517cb46, Chrome 154 headless with a Metal adapter, an idle Apple M5 Pro and three runs. The int8 result here is 0.9 ms; the earlier article's separate run rounded to 1.0 ms. ↩ -
The batch-64 comparison comes from the same timing record as the per-move table. See the ternary kernels and optimization notes. For related integer-kernel work, see Wang et al., 1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs, 2024. That work concerns CPUs; it does not measure the proposed changes to our WebGPU implementation. ↩