bk99.de entertain the web since 1997

Faster output through self-speculative decoding

Summary

LayerSkip uses early layers of the same model as a fast draft and checks the proposals with the complete model. This makes a separate draft model unnecessary. Self-speculative decoding uses the same model for drafting and verification.

Ideas

  • A model can use its own intermediate representations as cheap proposals.
  • The speed-up depends on how often the complete model accepts early tokens.

Facts

  • Training with layer dropout and early exit improves the usefulness of early layers.

Critique

  • The method requires suitably trained models and gains little at a low acceptance rate.

Recommendations

  • Measure the acceptance rate and output identity for different tasks and sequence lengths.

References

Read the original article on Hugging Face

Search the Web Archive