← All insights

AI's unbelievable 437× memory shrinkage

In 2023, DeepSeek's first model used an astonishing 1,167,360 bytes of memory for a single word, unbelievable. Since then, clever optimisations have slashed this. The recent V4.1 Flash model requires only 2,670 bytes for the same task. This represents a 437-fold reduction in less than three years. This is an important but little-noticed advance, which is part of why models can become more capable and cheaper compared to past versions.

Top: the word unbelievable split into three tokens, each above a stack of 437 blocks, which corresponds to 1,167,360 bytes in 2023. Bottom: the same word as three tokens, each with just one block, which corresponds to 2,670 bytes in 2026.

The example involves one word and three tokens in both the initial DeepSeek model and V4.1 Flash. Each block represents 890 bytes. These are still images from the 30-second video.

Where the bytes go

A model reads text as tokens, breaking longer words into smaller chunks. When entered as a standalone message, unbelievable consists of three tokens in both models. In the first model, they are "un", "bel", and "ievable". In V4.1 Flash, they are "un", "belie", and "vable". For every token, the model remembers "keys" and "values" in each layer. This eliminates the need to re-process the entire conversation history for every newly generated word. This storage area is known as the KV cache. It resides in High Bandwidth Memory (HBM) located directly next to the processor, a resource that is currently scarce and expensive.

The initial DeepSeek model features 95 layers, each with 8 key-value heads of 128 dimensions, stored in 16-bit format. This results in 95 × 8 × 128 × 2 (keys and values) × 2 bytes = 389,120 bytes per token, or 1,167,360 bytes for the three tokens making up the word un-bel-ievable. This was typical for 2023. For comparison, Llama 2 required 524,288 bytes per token for the 7B variant, 819,200 for the 13B, and 327,680 for the 70B.

Fast-forward to today: the paper on DeepSeek's V4.1 Flash specifies a global KV cache of 890 bytes per token. This represents a 437-fold reduction compared to the first model. Storing the cache in 4-bit floating-point format instead of 16-bit accounts for a factor of four at most. The rest comes from what DeepSeek calls Compressed Sparse Attention 2, which compresses the long-term memory and shares it across layers. A small sliding window holds the most recent tokens and does not grow with the conversation.

What this means for you

More cost-effective artificial intelligence. Model weights represent a fixed cost factor on every chip. The cache grows with each conversation. It determines how many conversations can take place simultaneously. A cache that is 437 times smaller frees up space for far more conversations. This does not automatically translate to 437 times as many users, since weights and computational overhead also play a role. But expect the smaller cache to reduce memory pressure and lower deployment costs.

Fast, long chats. According to the paper, the computational cost per generated token remains nearly constant even as the context grows. It increases by only about a quarter when scaling from 4,000 to one million tokens. And V4.1 Flash supports contexts of up to one million tokens. A 100,000-word novel corresponds to approximately 130,000 tokens.

Why openness matters

Refreshingly, DeepSeek provides a lot of insight into the inner workings of its models. That allows a different kind of exploration than benchmark comparisons, which reveal less and less. Many frontier labs are very tight-lipped and leave their models a black box. When OpenAI released GPT-6.1 Sol only a week after GPT-6 Sol, the rushed-looking follow-up led to a lot of speculation. So little is published that we do not even know whether they are the same model with different quantisation or settings, or completely different models.

Hopefully, DeepSeek will continue publishing details to share with the public some of the more interesting advances in AI.