bk99.de entertain the web since 1997

Transformers for long documents

Summary

The literature review classifies methods that avoid the quadratic cost of full attention. Local windows, sparsity and compressed representations trade a global view for scalability. Full self-attention grows quadratically with the sequence length.

Ideas

  • Long contexts need a defined information route between distant tokens.
  • Every attention approximation determines which relationships a model can learn easily.

Insights

  • Architecture decisions distribute computing costs and information paths; they are not mere implementation details.

Facts

  • Longformer and BigBird use sparse attention patterns for longer inputs.

Critique

  • Maximum context length alone measures neither usable memory nor answer quality.

Recommendations

  • For long documents, check whether the information you are looking for can reach the chosen attention structure.

References

Read the original article on Hugging Face

Search the Web Archive