Can synthetic data unlock AI recursive self-improvement? — Mark Zuckerberg
Dwarkesh Patel · 4:06
Even after training Llama’s 70B on around 15 trillion tokens, the model was still improving rather than flattening as expected, so Meta stopped further pretraining as a GPU-allocation call—not because data had run out...