Understanding transformers as encoder-decoders
Summary
The article takes apart encoder-decoder transformers and explains how they differ from pure encoder or decoder models. Cross-attention in particular connects the representation of the input with the step-by-step output. The encoder processes the input sequence bidirectionally.
Ideas
- Encoder and decoder solve different parts of a sequence transformation.
- Cross-attention makes the encoded input specifically usable at every output step.
Insights
- Architecture decisions distribute computing costs and information paths; they are not mere implementation details.
Facts
- The decoder masks future output tokens and takes the encoder states into account.
Critique
- The conceptual explanation says little about the training costs and failure patterns of specific models.
Recommendations
- Choose the architecture by task and data flow rather than by a model’s mere popularity.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.