Longer generation through a quantised KV cache
Summary
The article quantises the key-value cache, which takes up a growing share of GPU memory with long contexts. This leaves more memory free for longer sequences or larger batches. The KV cache grows with sequence length, number of layers and batch size.
Ideas
- In autoregressive generation, the state cache can become more important than the model weights.
- Cache quantisation trades accuracy of memory for context capacity.
Facts
- Transformers supports quantised cache implementations for selected backends.
Critique
- Additional quantisation steps can slow down short requests and introduce subtle errors.
Recommendations
- Test long contexts for quality loss and measure from which length quantisation actually helps.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.