Sentence Transformers in the Hub
Summary
Sentence Transformer models get a shared home with metadata and a directly usable loading function. This makes semantic search and similarity comparison easier to reproduce. Sentence Transformers produce fixed vectors for sentences and texts.
Ideas
- Embeddings only become comparable when model, pooling and preprocessing are versioned together.
- Model cards link technical artefacts with their intended and unsuitable uses.
Insights
- The best representation is the one whose quality and operating costs suit the specific search task.
Facts
- Models can be loaded directly by their Hub identifier.
Critique
- A popular embedding model can perform considerably worse on technical language or other languages.
Recommendations
- Document the distance measure, normalisation and model revision used next to every stored index.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.