LLM Model Parallelism
Ethereum & AI Engineering Group · 47:04
Peter Robinson's talk surveys how large language models are split across GPUs and even across data centers, and argues that the hottest current research problem — cutting the communication cost of pipeline parallelism...