bk99.de entertain the web since 1997

Longer generation through a quantised KV cache

Summary

The article quantises the key-value cache, which takes up a growing share of GPU memory with long contexts. This leaves more memory free for longer sequences or larger batches. The KV cache grows with sequence length, number of layers and batch size.

Ideas

  • In autoregressive generation, the state cache can become more important than the model weights.
  • Cache quantisation trades accuracy of memory for context capacity.

Facts

  • Transformers supports quantised cache implementations for selected backends.

Critique

  • Additional quantisation steps can slow down short requests and introduce subtle errors.

Recommendations

  • Test long contexts for quality loss and measure from which length quantisation actually helps.

References

Read the original article on Hugging Face

Search the Web Archive