The Engineering Behind LLM Inference: The Memory Wall
PY · 31:13
A frontier language model spends about 30 ms on each output token, but less than 1% of that time is arithmetic: decode is memory-bandwidth bound because every token streams the full weight set from HBM, while the math...