The biggest and weirdest Linux kernel commits
Summary
Gary Bernhardt examines extreme objects in the Linux Git history and explains how imports, tools and data formats create surprising outliers. The Linux kernel has a very large public Git history. Git addresses objects via cryptographic hash values.
Ideas
- Git stores commit, tree and content as separate object types.
- Automated imports create different patterns from normal development work.
- Large changes can arise from metadata or generated files.
- Historical outliers reveal earlier tools and ways of working.
Insights
- Repository statistics need context before size or number of authors gain meaning.
- Version history contains traces of organisational migrations alongside code changes.
- Extreme cases are good entry points for understanding a data model.
Facts
- The article analyses data directly from the repository.
Recommendations
- Examine conspicuous commits with git show and git cat-file.
- Separate generated, imported and hand-edited content in analyses.
References
Links to the original source and the Web Archive open in a new tab.