Learning track
The Transformer
Assemble the architecture behind modern LLMs, beginning with self-attention and its query, key, and value operations.
The Transformer
Assemble the architecture behind modern LLMs, beginning with self-attention and its query, key, and value operations.
- 05.01"Attention Is All You Need" in context→
- 05.02Self-attention from first principles: Q, K, V→
- 05.03Scaled dot-product attention & the √d_k→
- 05.04Causal masking & why order matters→
- 05.05Multi-head attention→
- 05.06Positional encoding I: sinusoidal & learned→
- 05.07Positional encoding II: RoPE→
- 05.08Positional encoding III: ALiBi & relative bias→
- 05.09The feed-forward block & where knowledge lives→
- 05.10Residuals, pre-norm vs post-norm→
- 05.11The classic block — and Qwen's two block types→
- 05.12Transformer families, from encoder-only to hybrid→
- 05.13Build a GPT from scratch: the uniform baseline→
- 05.14Reading real weights: opening Qwen3.8-27B's checkpoint→
- 05.15Gated DeltaNet: the mixer in 48 of 64 layers→
- 05.16Hybrid layouts: three DeltaNet, one attention→
- 05.17How images join the sequence: the vision encoder→