Glossary Index
- Scaled Dot Product Attention
- Multi Head Attention
- Linear
- Rotary Positional Encoding
- Grouped Query Attention
- Qwen3 Attention Module
- RMSNorm
- SiLU
- SwiGLU
- MLP
- Embedding
- Qwen3 Transformer Block
- Week 1 Qwen3 Model
- dequantize_linear
- KV Cache and Request-Bounded Capacity
- Benchmarking, Profiling, and Decode Roofline
- Packed W4 Quantization
- Fused RMSNorm, RoPE, and SwiGLU
- Tiled Dense Prefill Attention and Selected Model
- Historical: SIMD-Matrix Prefill
- Historical: Bounded Decode Attention
- Historical: Split-K Prefill
- Flash Attention
- Paged Attention
Your feedback is greatly appreciated. Join our Discord community.
Found an issue? Open an issue or pull request at github.com/skyzh/tiny-llm.
tiny-llm-book © 2025 by Alex Chi Z is licensed under CC BY-NC-SA 4.0.