Why is vectorized execution mentioned alongside columnar storage?
Columnar data is naturally amenable to vectorization. Modern CPUs operate on vectors of values efficiently (SIMD instructions). With columns in contiguous memory, a CPU can process 8–64 values per instruction cycle. Row storage scatters related data, preventing vectorization. That's why columnar databases are fast—they align data layout with CPU capabilities, not despite the layout.
Answered in
Columnar Storage: Why Column Stores Beat Row Stores for AnalyticsColumnar storage reads only needed columns, skipping the rest. Dictionary encoding shrinks data 50–100×. Analytics queries go from minutes to milliseconds.
Read the full analysisOther questions this article answers
More system design questions
- Why does an index's internal data structure matter if it all ends up 'faster than a scan'?
- Why is disk I/O the thing index structures are actually optimizing for?
- Why can't a hash index handle range queries?
- What makes a bitmap index different from a B-tree, and when is it better?
- Why do B-trees stay balanced automatically as data is inserted?
- Why does a database need an index at all — why can't it just scan the table?
- When is a hash index better than a B-tree index?
- What is a composite index and why does column order matter?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the System Design beat.