topic: ai-technology
author: Crashtech Editorial
date: Oct 7, 2026 · read: 2 min
updated: October 8, 2026
---
Mamba, RWKV and Hybrid Models: Different Ways to Carry Long Context
Selective state spaces and recurrent models change sequence-processing costs. Constant recurrent state does not provide unlimited exact recall.
On this page
Conceptual architecture; arrows show relationships, not measured latency, energy or safety guarantees.
Separate prefill, decoding and memory
For full self-attention, processing N positions involves pairwise query-key interactions with quadratic sequence-length compute. A naive implementation also materializes an N-by-N matrix. FlashAttention avoids storing that full matrix in high-bandwidth memory; it does not remove the quadratic arithmetic of full attention. [4]
During autoregressive decoding, conventional cached attention stores keys and values from the history. That cache grows linearly with sequence length, and a new query attends over the retained history. Calling every attention memory cost quadratic confuses distinct phases and implementations.
State spaces compress the past
Mamba introduces input-dependent selection within a state-space sequence model. Its hardware-aware selective scan enables efficient processing, while recurrent decoding updates a state rather than appending a transformer-style KV cache. [1]
Because selection depends on the input, Mamba’s scan should not be described as simply the fixed convolution or FFT formulation of a time-invariant SSM. Mamba-2 develops related ideas through structured state space duality. [2]
RWKV is another recurrent model family, with its own architecture and training formulation. It should not be treated as a second name for Mamba. [5]
Fixed state has a tradeoff
A fixed-size state is an efficiency property, not an unlimited transcript. The model must encode useful history into that state. Recovering an arbitrary identifier far back in a stream may differ from recognizing an overall pattern.
Evaluate exact retrieval, copying, repeated entities and distractors separately from language-modeling loss. Neither attention nor recurrence guarantees perfect long-context recall merely from its architectural label.
Hybrids retain selected attention layers
Jamba is a published example combining transformer and Mamba layers. That makes it a concrete hybrid reference; Gemma 2 should not be cited as a Mamba hybrid. [3]
Fewer attention layers can reduce KV-cache requirements relative to a comparable all-attention design, but the percentage depends on layer counts, head structure and cache precision. Retaining attention layers also means the complete hybrid is not automatically linear-time for every context length.
For a deployment decision, compare prefill time, per-token latency, cache size and task quality at the required context lengths. Constant recurrent state is valuable when the application tolerates its information tradeoffs. It does not make the attention design obsolete for every workload.
Frequently asked questions
Does FlashAttention make full attention linear-time?
No. It reduces memory traffic and avoids storing the full attention matrix, while exact full attention still has quadratic sequence-length compute during prefill.
Does Mamba offer infinite usable context?
No. Its recurrent state can remain fixed in size, but finite state does not preserve arbitrary detail from an unlimited history. Quality must be evaluated at the intended sequence length.
/* Comments */
Comments are offline right now — we reconnect automatically, nothing is lost.