Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Week 1: From Matmul to Text

This week, we will start with basic array and matrix operations and use them to turn Qwen3 model parameters into a model that generates text. We will implement the neural network layers used by Qwen3 with MLX’s array APIs.

We will use Qwen/Qwen3-0.6B-MLX-4bit. Week 1 dequantizes model parameters into bfloat16, so start with the 0.6B model before trying larger Qwen3 models.

What We Will Cover

  • Attention, multi-head attention, grouped-query attention, and multi-query attention
  • Positional encodings and RoPE
  • Using mx.fast.rms_norm for Qwen3’s per-head Q/K normalization, then implementing RMSNorm ourselves
  • Implementing the MLP, combining the attention components, and building the complete Transformer model
  • Loading Qwen3 model parameters and generating text

What We Will Not Cover

To make the journey as interesting as possible, we will skip a few things for now:

  • Quantization and dequantization internals. These will be covered in Week 2. For now, we use a provided helper to dequantize the Qwen3 weights before passing them to our layer implementations.
  • Low-level implementations of operations such as softmax, exponentiation, and logarithms. These operations are simple enough that using the MLX versions does not detract from the learning objectives.
  • Tokenization. We use the mlx_lm tokenizer rather than implementing one from scratch.
  • Decoding model-weight files. We use mlx_lm to load the model, then transfer its weights into our layer implementations.

Basic Matrix APIs

MLX’s Python API is designed to be familiar to NumPy users. If you are new to array programming, start with NumPy: the absolute basics for beginners.

You can also refer to the MLX Operations API for more details.

Qwen3 Models

You can run Qwen3 with MLX or vLLM. The readings below provide context for what we will build. By the end of the week, you will be able to use Qwen3 as a causal language model to generate text.

Reference implementations of Qwen3 are available in Hugging Face Transformers, vLLM, and mlx-lm. Use them to explore the model’s internals and compare them with this week’s implementation.

📚 Readings

Your feedback is greatly appreciated. Join our Discord community.
Found an issue? Open an issue or pull request at github.com/skyzh/tiny-llm.
tiny-llm-book © 2025 by Alex Chi Z is licensed under CC BY-NC-SA 4.0.