bk99.de entertain the web since 1997

Training a language model from scratch

Summary

Using Esperanto as an example, Hugging Face shows how the tokenizer, data set and transformer are built up from scratch together. The experiment makes visible which decisions finished models usually hide. The example model has six transformer layers and around 84 million parameters.

Ideas

  • A tokenizer of your own prevents a target language from being processed with unsuitable word boundaries.
  • Training from scratch separates architectural knowledge from merely reusing finished weights.

Insights

  • A base model of your own is only worthwhile when data or language measurably overwhelm existing models.

Facts

  • The trained encoder is then fine-tuned for part-of-speech tagging.

Critique

  • The small Esperanto example does not yet prove superiority over strong multilingual base models.

Recommendations

  • First check whether an existing multilingual model serves the purpose before you train a new one.

References

Read the original article on Hugging Face

Search the Web Archive