The Engineering Behind LLM Inference: The Memory Wall

PY · 31:13

A frontier language model spends about 30 ms on each output token, but less than 1% of that time is arithmetic: decode is memory-bandwidth bound because every token streams the full weight set from HBM, while the math...

Read the full summary on tuber

Redirecting...