Mixture of Experts (MoE), Visually Explained
Jia-Bin Huang · 31:46
Mixture of Experts (MoE) scales a transformer’s parameter count by splitting each large feedforward network into many smaller experts and routing every token to only a sparse top-k subset, so training and inference co...