topic: ai-technology
author: Crashtech Editorial
date: Oct 7, 2026 · read: 2 min
updated: October 8, 2026
---
The KV Cache: How GQA and DeepSeek MLA Reduce Inference Memory
KV-cache size depends on layers, heads, precision and context. GQA shares key-value heads; MLA caches a learned latent plus positional data.
On this page
Conceptual architecture; arrows show relationships, not measured latency, energy or safety guarantees.
Calculate the cache before quoting a percentage
For conventional cached attention, a useful simplified formula is:
bytes per sequence = 2 × layers × KV_heads × head_dimension
× tokens × bytes_per_element
The first factor counts keys and values. The formula excludes model weights, buffers, allocator overhead and implementation-specific padding or replication.
As an illustrative configuration, use 80 layers, eight KV heads, a head dimension of 128 and two bytes per element. That gives 327,680 bytes per token. At 131,072 tokens, the cache is 40 GiB for one sequence before the excluded overheads. With 64 KV heads instead, the same calculation gives 320 GiB. These are arithmetic examples, not measurements of a named production model.
GQA shares key-value heads
Grouped-query attention allows several query heads to share key and value heads. Relative to MHA with otherwise identical dimensions, fewer KV heads mean a smaller cache. The GQA paper studies uptraining and reports quality close to its multi-head baselines with better inference efficiency. It does not establish that GQA universally fails at long-context recall. [3]
The ratio matters: 128 KV heads reduced to eight is a sixteen-fold change, not eight-fold. State the dimensions rather than mixing head counts from unrelated models.
MLA changes the model architecture
DeepSeek-V2 introduces joint low-rank compression of keys and values. A learned latent representation is cached alongside a separate component for rotary positional information. Decoupling the positional part enables inference computations that would otherwise obstruct useful matrix absorption. [1]
The V2 report’s 93.3% KV-cache reduction is a result relative to its specified comparison, not a percentage obtained from any arbitrary combination of 512 latent dimensions and 128 heads. DeepSeek-V3 also uses MLA, but model-specific dimensions and serving kernels must be checked. [2]
Smaller cache is not free capacity
MLA does not eliminate context-dependent cache growth. Nor does it make total model storage disappear: the number of active MoE parameters per token is not the number of weights that must be available to the runtime.
Measure peak memory, prefill time and decode throughput at realistic concurrency. Compare output quality on the intended tasks. Architectural compression, cache quantization, paging and offloading solve related but different problems; none alone establishes a universal batch-size or cost multiplier.
Frequently asked questions
Is a KV cache always 80% of GPU memory?
No. The share depends on model weights, cache precision, context length, concurrency and serving buffers. A long context can be expensive without a universal percentage.
Does MLA losslessly compress an existing transformer cache?
No. MLA is a learned model architecture. DeepSeek’s reported reduction is relative to its stated baseline; it is not a generic lossless compressor for any pretrained model.
/* Comments */
Comments are offline right now — we reconnect automatically, nothing is lost.