Learning track

Post-training & Alignment

Turn pretrained predictors into useful assistants through instruction tuning, preference learning, reinforcement learning, and distillation.

07

Post-training & Alignment

Turn pretrained predictors into useful assistants through instruction tuning, preference learning, reinforcement learning, and distillation.

  1. 07.01Why base models aren't assistants
  2. 07.02Supervised fine-tuning & instruction data
  3. 07.03LoRA, QLoRA & PEFT
  4. 07.04Reward models & human preference data
  5. 07.05RLHF with PPO
  6. 07.06DPO: skipping the reward model
  7. 07.07ORPO, KTO, SimPO & the alignment zoo
  8. 07.08GRPO & RL on verifiable rewards
  9. 07.09Reasoning models: test-time compute & long CoT
  10. 07.10Constitutional AI & RLAIF
  11. 07.11Distillation: making small models punch up