Pretrained LLMs Are Surrounded by Task Experts feat Yulu Gan from MIT
Deep Learning with Yacine · 72:18
This paper argues that pre-trained neural networks don't converge to one optimal solution, but land in a broad flat "basin" surrounded by many nearby "task-expert" configurations reachable simply by adding random nois...