Week 1: From Matmul to Text
This week, we will start with basic array and matrix operations and use them to turn Qwen3 model parameters into a model that generates text. We will implement the neural network layers used by Qwen3 with MLX’s array APIs.
We will use Qwen/Qwen3-0.6B-MLX-4bit. Week 1 dequantizes model parameters into bfloat16, so start with the 0.6B model
before trying larger Qwen3 models.
What We Will Cover
- Attention, multi-head attention, grouped-query attention, and multi-query attention
- Positional encodings and RoPE
- Using
mx.fast.rms_normfor Qwen3’s per-head Q/K normalization, then implementing RMSNorm ourselves - Implementing the MLP, combining the attention components, and building the complete Transformer model
- Loading Qwen3 model parameters and generating text
What We Will Not Cover
To make the journey as interesting as possible, we will skip a few things for now:
- Quantization and dequantization internals. These will be covered in Week 2. For now, we use a provided helper to dequantize the Qwen3 weights before passing them to our layer implementations.
- Low-level implementations of operations such as softmax, exponentiation, and logarithms. These operations are simple enough that using the MLX versions does not detract from the learning objectives.
- Tokenization. We use the
mlx_lmtokenizer rather than implementing one from scratch. - Decoding model-weight files. We use
mlx_lmto load the model, then transfer its weights into our layer implementations.
Basic Matrix APIs
MLX’s Python API is designed to be familiar to NumPy users. If you are new to array programming, start with NumPy: the absolute basics for beginners.
You can also refer to the MLX Operations API for more details.
Qwen3 Models
You can run Qwen3 with MLX or vLLM. The readings below provide context for what we will build. By the end of the week, you will be able to use Qwen3 as a causal language model to generate text.
Reference implementations of Qwen3 are available in Hugging Face Transformers, vLLM, and mlx-lm. Use them to explore the model’s internals and compare them with this week’s implementation.
📚 Readings
- Qwen3: Think Deeper, Act Faster
- Hugging Face Transformers — Qwen3
- vLLM Qwen3
- mlx-lm Qwen3
- Qwen3 Technical Report
Your feedback is greatly appreciated. Join our Discord community.
Found an issue? Open an issue or pull request at github.com/skyzh/tiny-llm.
tiny-llm-book © 2025 by Alex Chi Z is licensed under CC BY-NC-SA 4.0.