Learning track
Building & Serving Stacks
Master the software that actually runs Qwen3.8-27B — PyTorch, Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, MLX, and Modular MAX — and learn which one each situation calls for.
Building & Serving Stacks
Master the software that actually runs Qwen3.8-27B — PyTorch, Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, MLX, and Modular MAX — and learn which one each situation calls for.
- 09.01PyTorch: the reference runtime→
- 09.02The Hugging Face Transformers ecosystem→
- 09.03torch.compile and CUDA graphs→
- 09.04vLLM→
- 09.05SGLang→
- 09.06TensorRT-LLM→
- 09.07llama.cpp and GGUF→
- 09.08Ollama and local runners→
- 09.09MLX→
- 09.10Modular MAX and Mojo→
- 09.11Serving in the cloud→
- 09.12Choosing your stack→