Deriving Git from simple file system ideas
Summary
Tom Preston-Werner develops Git's basic model step by step from files, snapshots, hash values, pointers and shared history. Git uses content-addressed objects. Commits form a directed graph through parent references.
Ideas
- Complete project copies initially form simple manual snapshots.
- Hash values give immutable content stable names.
- Tree objects link file names with stored content.
- Commits add parents, author and description.
- Branches name the movable tips of a sequence of commits.
- Distributed repositories exchange missing objects instead of central states.
Insights
- Complex tools become understandable when their data model is reconstructed step by step.
- Immutable content simplifies replication, verification and caching.
- Names and content need different lifecycles.
- Distribution works more easily when objects are independently addressable.
Facts
- Conceptually, branches consist of references to commits.
Recommendations
- Draw the object relationships of a small repository by hand.
- Use git cat-file to examine stored objects.
- Separate immutable objects from movable references.
References
Links to the original source and the Web Archive open in a new tab.