bk99.de entertain the web since 1997

Sentence Transformers in the Hub

Summary

Sentence Transformer models get a shared home with metadata and a directly usable loading function. This makes semantic search and similarity comparison easier to reproduce. Sentence Transformers produce fixed vectors for sentences and texts.

Ideas

  • Embeddings only become comparable when model, pooling and preprocessing are versioned together.
  • Model cards link technical artefacts with their intended and unsuitable uses.

Insights

  • The best representation is the one whose quality and operating costs suit the specific search task.

Facts

  • Models can be loaded directly by their Hub identifier.

Critique

  • A popular embedding model can perform considerably worse on technical language or other languages.

Recommendations

  • Document the distance measure, normalisation and model revision used next to every stored index.

References

Read the original article on Hugging Face

Search the Web Archive