Stanford CS25: Transformers United V6 I The Ultra-Scale Talk: Scaling Training to Thousands of GPUs

Stanford Online · 61:48

Modern LLM training (trillion-parameter models, ~15T tokens, long context) does not fit on one GPU and cannot be scaled by data parallelism alone. This talk walks through five orthogonal parallelisms—DP/ZeRO/FSDP, ten...

Read the full summary on tuber

Redirecting...