XetHub brings block-based storage to the Hub
Summary
XetHub joins Hugging Face to version large model and data set files more efficiently. Content-based deduplication is meant to transfer changed blocks instead of complete files. Xet works with block-based, content-addressed storage.
Ideas
- ML versioning needs deduplication below the level of huge monolithic files.
- Content addressing can share storage and transfer across many similar model versions.
Insights
- ML storage has to move large, similar artefacts efficiently without giving up reproducibility.
Facts
- The technology is being integrated into the infrastructure of the Hugging Face Hub.
Critique
- A new storage layer increases protocol and vendor dependency as long as alternative tools are missing.
Recommendations
- Measure upload, checkout, cache requirements and recovery with your real model revisions.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.