Reformer for long sequences
Summary
Reformer reduces the memory and computing requirements of classic transformers through locality-sensitive hashing and reversible layers. This makes much longer inputs possible on limited hardware. Reformer replaces full attention with locality-sensitive hashing.
Ideas
- Approximation can push quadratic attention down to manageable costs.
- Reversible layers save activation memory by reconstructing intermediate states during the backward pass.
Insights
- Architecture decisions distribute computing costs and information paths; they are not mere implementation details.
Facts
- The architecture combines LSH attention with reversible residual layers.
Critique
- Efficiency gains depend heavily on sequence length, implementation and hardware utilisation.
Recommendations
- Measure quality loss and runtime with your actual sequence length rather than only with model size.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.