Optimizing the Full Stack for Generative Image and Video Models
MIT OpenCourseWare · 30:02
This talk argues that serving modern text-to-image and text-to-video diffusion (and flow) models is a full-stack problem, not a single-model speedup: the iterative diffusion network is compute-bound and memory-heavy,...