bk99.de entertain the web since 1997

Fine-tuning language models to 1.58 bits

Summary

The article applies ternary weights to existing language models and shows an accessible fine-tuning path. Extreme quantisation promises small models but requires careful quality measurement. 1.58 bits correspond to the information content of three possible weight values.

Ideas

  • Ternary weights restrict parameters to minus one, zero and plus one.
  • Extreme quantisation works better as a training method than as blind rounding after the fact.

Insights

  • Lower precision is a system decision between quality, memory, kernels and target hardware.

Facts

  • The approach uses BitNet-like linear layers and fine-tuning.

Critique

  • A compact representation does not automatically bring energy or time savings without specialised kernels.

Recommendations

  • Compare task quality and real kernel speed with 4-bit and 8-bit baselines.

References

Read the original article on Hugging Face

Search the Web Archive