How does Git detect file changes without scanning every file on disk?
Git builds a Merkle tree: files are hashed into blob objects, directories into tree objects, and trees into a root tree. If a tree's hash matches the previous commit, everything inside is provably unchanged. Git skips entire subtrees, descending only where hashes differ. This reduces checking millions of files to comparing a handful of hashes.
Answered in
Merkle Trees: How Git Detects Changes in MillisecondsGit hashes files into nested cryptographic trees to skip unchanged directories in one comparison, finding changes across millions of files faster than scanning.
Read the full analysisOther questions this article answers
More system design questions
- Why doesn't Google just run Dijkstra faster?
- What is a shortcut edge and when is it precomputed?
- How much space do shortcut edges take compared to the original graph?
- Can Contraction Hierarchies handle dynamic graphs like traffic or road closure?
- Why contract low-degree nodes first instead of high-degree ones?
- What is a CRDT and why does it matter for real-time collaboration?
- How do CRDTs handle concurrent edits without a central server referee?
- Why did Figma move from operational transforms to CRDTs?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the System Design beat.