GPU Communications for Python - Benjamin Glick, Michael Yh Wang

PyCon US · 29:40

Ben and Michael walk through NVIDIA’s native-Python CUDA stack for multi-GPU work: host/device memory and streams first, then NCCL (collectives such as all-reduce, plus send/recv) and NVSHMEM (PGAS symmetric memory an...

Read the full summary on tuber

Redirecting...