bk99.de entertain the web since 1997

Uploads and downloads from chunk to block

Summary

The article explains the conversion of large Hub transfers to block-based Xet storage. Deduplication and parallel transfer are meant to handle changes to large files efficiently. The new architecture splits files into smaller content-addressed blocks.

Ideas

  • Block boundaries should recognise identical content even after local changes to a file.
  • Transfer performance comes from the interplay of deduplication, parallelism and cache.

Insights

  • ML storage has to move large, similar artefacts efficiently without giving up reproducibility.

Facts

  • Blocks that already exist do not have to be transferred again.

Critique

  • Deduplication costs indexing work and helps little with completely new or encrypted data.

Recommendations

  • Test the behaviour on aborts, cache loss and many small changes to large files.

References

Read the original article on Hugging Face

Search the Web Archive