What is read amplification and how do Bloom filters fix it?
Read amplification occurs when a single lookup must search many SSTables across multiple levels before finding the key (or confirming its absence). A Bloom filter is a probabilistic data structure that can definitively say a key is NOT in an SSTable with no disk I/O; it eliminates false positives but allows false negatives.
Answered in
LSM-Trees vs B-Trees: Why Cassandra Chose Sequential WritesB-trees seek random disk positions. LSM-trees buffer in memory and flush sequentially, converting random I/O into sequential writes for millions of ops/sec.
Read the full analysisOther questions this article answers
More system design questions
- Why doesn't Google just run Dijkstra faster?
- What is a shortcut edge and when is it precomputed?
- How much space do shortcut edges take compared to the original graph?
- Can Contraction Hierarchies handle dynamic graphs like traffic or road closure?
- Why contract low-degree nodes first instead of high-degree ones?
- What is a CRDT and why does it matter for real-time collaboration?
- How do CRDTs handle concurrent edits without a central server referee?
- Why did Figma move from operational transforms to CRDTs?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the System Design beat.