Does FlashAttention make full attention linear-time?
No. It reduces memory traffic and avoids storing the full attention matrix, while exact full attention still has quadratic sequence-length compute during prefill.
Answered in
Mamba, RWKV and Hybrid Models: Different Ways to Carry Long ContextSelective state spaces and recurrent models change sequence-processing costs. Constant recurrent state does not provide unlimited exact recall.
Read the full analysisOther questions this article answers
More how ai actually works questions
- What are the three scenarios for the AI economy?
- What happens in the 25% Bull Case?
- What would trigger the 15% Bear Case crash?
- What does the 60% base case look like?
- Who are the winners in the 60% base case?
- Why is BitNet called 1.58-bit?
- Can any existing model be converted to BitNet?
- Does GraphRAG eliminate hallucinations?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the How AI Actually Works beat.