Parameter Golf Architecture Explorer
How do tiny language models fit more intelligence into 16 MB?
Parameter Golf submissions squeeze an entire language model into 16 MB, but the winning ideas are buried in code and scoreboards. This interactive explorer maps the architecture, explains each component, and compares what the strongest submissions changed.

What is the Parameter Golf Architecture Explorer?
Parameter Golf is a model-building challenge with a severe constraint: fit the language model into a 16 MB artifact while preserving as much performance as possible. The explorer turns several of its strongest submissions into one readable architecture map.
ELI5: it is an exploded diagram of a tiny language model. Click any part to learn what it does, why it matters, and how different competitors changed it.
Explore the Architecture
The main diagram follows the model from input tokens through embeddings, transformer blocks, skip connections, and the output layer. Training and compression techniques sit alongside the model itself, including Muon, quantization-aware training, stochastic weight averaging, and sliding-window evaluation.
Each node includes:
- A plain-language explanation
- The important parameters
- Alternatives worth exploring
- Links for studying the idea in more depth
Compare Submissions
The explorer includes a naive baseline and five stronger approaches. Pick any two to highlight exactly where their architectures diverge, then inspect the change at each affected node. That makes it easier to see how small improvements compound: lower-bit quantization creates room for wider layers, local-context tricks reduce pressure on attention, and evaluation choices can improve the score without changing training.
Notes That Stay Local
Every component has a notes tab for observations and follow-up tasks. Notes stay in the browser through local storage, so the explorer doubles as a lightweight study guide without requiring an account or backend.
