Learning track
Inference & Efficiency
Understand how trained models generate text and how decoding choices trade off diversity, coherence, latency, and cost.
Inference & Efficiency
Understand how trained models generate text and how decoding choices trade off diversity, coherence, latency, and cost.
- 08.01What actually happens when you hit send→
- 08.02The KV cache→
- 08.03Prefill vs decode: two different machines→
- 08.04Sampling: temperature, top-k, top-p, and min-p→
- 08.05Beam search, speculative decoding & Medusa→
- 08.06MQA, GQA & MLA→
- 08.07FlashAttention & IO-awareness→
- 08.08PagedAttention & continuous batching→
- 08.09Quantization I: int8, int4, the basics→
- 08.10Quantization II: GPTQ, AWQ, GGUF, QAT→
- 08.11Serving stacks: vLLM, SGLang, TensorRT-LLM→
- 08.12Running models locally→
- 08.13Choosing and judging a community quant→