Stanford CS25: Transformers United V6 I From Language Models to Native Multimodal Intelligence
Stanford Online · 64:39
Native multimodal language models work by converting images, audio, and video into tokens and training a transformer the same way as a text LLM—but “native” is not solved. Understanding-only products and true omni mod...