Keyboard shortcuts

Press ← or β†’ to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

🚧 Week 2: A Faster Single Request

Week 1 leaves you with a readable Qwen3 model that regenerates from the full prefix. The current Week 2 route contains Days 1–5: reuse previous keys and values, bound their dense storage, keep projection weights packed, then reuse W4 and activation tiles during matrix-shaped prefill. Day 4 integrates RMSNorm, RoPE, and SwiGLU as three separate model checkpoints. Day 5 adds tiled dense prefill attention and runs the final selected model. Its nine cumulative checkpoints are kv-cache, capacity-cache, quantized-matvec, simd-matmul, rmsnorm, rope, swiglu, tiled-prefill, and selected.

Begin with Day 1: Cache and Measure. Its first feedback loop builds the extension required for test collection, runs the focused KV test, and sends only the new token during decode. Its second loop bounds storage without exposing unused capacity to attention. After both checkpoints, compare the same request across Week 1, kv-cache, and capacity-cache; keep the serving-only comparison separate from the all-logit algorithm comparison.

Continue with Day 2: Keep W4 Packed. Start from a working capacity-cache model. Implement selected-row embedding dequantization, a checked quantized operator, a readable Metal control and decode matvec, then wire packed projections into the cached model. The Day 2 checkpoint is complete when the model actually calls the packed-weight path.

Continue to Day 3: SIMD Matrix Prefill. Keep Day 2’s quantized-matvec checkpoint as the pre-edit control. Build a cooperative matrix kernel for the larger activation shapes, verify partial output tiles against the readable control, then compare the two checkpoints under the same cached-model workload.

Continue to Day 4: Fused Model Primitives. Keep simd-matmul as the pre-edit model control. Implement register-cached RMSNorm, RoPE over the model’s head layout, and fused SwiGLU in order. Compare each operator with its readable equation, run its cumulative checkpoint, then measure all four model variants under one matched 0.6B workload.

Finish with Day 5: Tiled Dense Prefill Attention. Keep swiglu as the pre-edit product control. Build the BF16/D128 tiled path, preserve causal and additive masks, GQA, and the short-query fallback, then run tiled-prefill and the completed selected model on the same cached 0.6B request. The selected checkpoint names the final cumulative feature set; it adds no second attention kernel.

Day 1 route

StepWhat you ownFeedback
PrepareBuild the Week 2 extension after Week 1 Day 7pdm run build-ext
Cache the prefixImplement dense K/V reuse in the model and generation looppdm run test --week 2 --day 1 -- -k 'not capacity'
Bound the cacheAllocate from the request limit, expose only the logical prefix, and preserve reset/rewind/overflow behaviorpdm run test --week 2 --day 1 after the capacity work
MeasureKeep the workload and prefill-logit mode matchedDay 1 measurement loop

Day 2 route

StepWhat you ownFeedback
PrepareKeep the Day 1 capacity-cache control; build learner and reference extensions for the native Day 2 checkspdm run build-ext and pdm run build-ext-ref
Keep W4 packedDequantize selected embedding rows, validate the wrapper, and implement the Metal matrix control and SIMD matvecFocused Day 2 tests
IntegrateRoute the cached model’s projections and output head through the packed operatorComplete Day 2 gate and a live quantized-matvec run
MeasureCompare capacity, packed W4, and MLX with one model and one workloadDay 2 measurement loop

Day 3 route

StepWhat you ownFeedback
PrepareComplete Day 2, build both native extensions, and save a quantized-matvec prefill controlDay 3 baseline
Build the tileLoad BF16 activation and reconstructed W4 fragments cooperatively; accumulate in FP32 and guard partial outputsFocused partial-tile test
IntegrateKeep the decode-shaped matvec, dispatch larger matrices to SIMD, and wire simd-matmul through the cached modelComplete Day 3 test and live model checkpoint
MeasureCompare old and new prefill paths with the same cached 0.6B model and workloadDay 3 product loop

Day 4 route

StepWhat you ownFeedback
PrepareComplete Day 3 and save a simd-matmul 0.6B product controlDay 4 baseline
NormalizeAdd register-cached RMSNorm with a wider-row fallback and integrate it into every Week 2 normFocused Task 1 check and live rmsnorm checkpoint
RotateReuse each RoPE angle across head pairs and accept per-batch offsetsFocused Task 2 check and live rope checkpoint
ActivateFuse SiLU and the gate/up product, then verify the cumulative modelFocused Task 3 check and complete Day 4 gate at swiglu
MeasureCompare the Day 3 control and three Day 4 checkpoints at one cached 0.6B workloadDay 4 product loop

Day 5 route

StepWhat you ownFeedback
PrepareKeep the Day 4 swiglu model and save its matched 0.6B product controlDay 5 baseline
TileImplement BF16/D128 grouped attention with BQ32/BK16 tiles and online softmaxFocused causal GQA and partial-tail check
PreserveApply causal/additive masks, return zero for fully masked rows, and keep readable short-query attentionMask and fallback checks
IntegrateRun the cumulative tiled-prefill model, then the named final selected modelComplete Day 5 gate and live model commands
MeasureCompare Day 4 control, tiled prefill, selected, and MLX on one cached 0.6B requestDay 5 product loop

The Day 1 starter supplies model loading, test entrypoints, and benchmark and attribution helpers; you own the cache state and serving loop. Day 2 adds the packed-weight container and native operator boundaries, but you implement the embedding, kernels, and model wiring. Day 3 adds the SIMD matrix path behind that packed operator. The reference solution and full MLX model are separate controls; they do not fill your learner TODOs. Day 4 adds three native primitives to the same cached, packed model. Day 5 adds the tiled attention operator and model selector. A cache counter shows which bytes moved, while a synchronized complete-request comparison shows whether a mechanism helped the chosen workload.

Historical lessons and Week 3

The five active days end at selected. Week 3 reuses the model and dense-cache interfaces while adding paging, batching, and serving policy; its learner work remains separate from the single-request Week 2 route.

Earlier seven-day lessons whose URLs are not reused for active Days 1–5 remain at their old addresses as historical Week 2 material. The former Day 4 address now serves the active fused-primitives lesson. The historical pages preserve benchmark method, W4 derivation, Apple M1–M4 bandwidth and roofline calculations, fusion/SIMD mechanisms, optional capture, and the old bounded-decode and Split-K experiments. Their operator-attribution diagram and decision diagram are also historical evidence, not diagrams of the current checkout. Those pages describe a different checkpoint order and may show commands unavailable in this partial branch. Use Days 1–5 above for the current learner workflow. The performance evidence ledger is likewise historical context, not a performance claim for this checkout.

Your feedback is greatly appreciated. Join our Discord community.
Found an issue? Open an issue or pull request at github.com/skyzh/tiny-llm.
tiny-llm-book Β© 2025 by Alex Chi Z is licensed under CC BY-NC-SA 4.0.