8: Deep Learning for Natural Language – Transformers, Self-Supervised Learning
MIT OpenCourseWare · 76:46
This lecture completes the 2017 transformer encoder: it recasts last class’s parameter-free self-attention as GPU-friendly matrix math, then adds three “industrial” pieces—learnable Q/K/V matrices, residual connection...