DeepSeek Splits the Transformer in Two. Here’s Why.
Jia-Bin Huang · 16:51
DeepSeek's new model (called V4.1 Flash in the video) cuts the global KV cache to about 890 bytes per token, 3.9× less than V4 Flash, and stays competitive on agent benchmarks. It does this by combining a causal encod...