8: Deep Learning for Natural Language – Transformers, Self-Supervised Learning

MIT OpenCourseWare · 76:46

This lecture completes the 2017 transformer encoder: it recasts last class’s parameter-free self-attention as GPU-friendly matrix math, then adds three “industrial” pieces—learnable Q/K/V matrices, residual connection...

Read the full summary on tuber

Redirecting...