Let's build the GPT Tokenizer
Andrej Karpathy · 133:34
Tokenization is a separate, non-neural preprocessing stage that compresses UTF-8 bytes into a finite vocabulary via byte-pair encoding (plus regex split rules and special tokens), and many LLM “model bugs”—spelling, n...