Uploads and downloads from chunk to block
Summary
The article explains the conversion of large Hub transfers to block-based Xet storage. Deduplication and parallel transfer are meant to handle changes to large files efficiently. The new architecture splits files into smaller content-addressed blocks.
Ideas
- Block boundaries should recognise identical content even after local changes to a file.
- Transfer performance comes from the interplay of deduplication, parallelism and cache.
Insights
- ML storage has to move large, similar artefacts efficiently without giving up reproducibility.
Facts
- Blocks that already exist do not have to be transferred again.
Critique
- Deduplication costs indexing work and helps little with completely new or encrypted data.
Recommendations
- Test the behaviour on aborts, cache loss and many small changes to large files.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.