Faster output through self-speculative decoding
Summary
LayerSkip uses early layers of the same model as a fast draft and checks the proposals with the complete model. This makes a separate draft model unnecessary. Self-speculative decoding uses the same model for drafting and verification.
Ideas
- A model can use its own intermediate representations as cheap proposals.
- The speed-up depends on how often the complete model accepts early tokens.
Facts
- Training with layer dropout and early exit improves the usefulness of early layers.
Critique
- The method requires suitably trained models and gains little at a low acceptance rate.
Recommendations
- Measure the acceptance rate and output identity for different tasks and sequence lengths.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.