Optimizing the Full Stack for Generative Image and Video Models

MIT OpenCourseWare · 30:02

This talk argues that serving modern text-to-image and text-to-video diffusion (and flow) models is a full-stack problem, not a single-model speedup: the iterative diffusion network is compute-bound and memory-heavy,...

Read the full summary on tuber

Redirecting...