bk99.de entertain the web since 1997

Understanding transformers as encoder-decoders

Summary

The article takes apart encoder-decoder transformers and explains how they differ from pure encoder or decoder models. Cross-attention in particular connects the representation of the input with the step-by-step output. The encoder processes the input sequence bidirectionally.

Ideas

  • Encoder and decoder solve different parts of a sequence transformation.
  • Cross-attention makes the encoded input specifically usable at every output step.

Insights

  • Architecture decisions distribute computing costs and information paths; they are not mere implementation details.

Facts

  • The decoder masks future output tokens and takes the encoder states into account.

Critique

  • The conceptual explanation says little about the training costs and failure patterns of specific models.

Recommendations

  • Choose the architecture by task and data flow rather than by a model’s mere popularity.

References

Read the original article on Hugging Face

Search the Web Archive