bk99.de entertain the web since 1997

Serving thirty LoRA models with one server

Summary

Text Generation Inference loads many LoRA adapters on top of a shared base model and serves them together. This avoids complete model copies and improves utilisation. The demonstration serves 30 LoRA models through one TGI instance.

Ideas

  • Adapter multi-tenancy shares expensive base weights between specialised services.
  • Schedulers have to batch different adapters together without unfair latency spikes.

Facts

  • All adapters use the same loaded base model.

Critique

  • The number of possible adapters depends on size, traffic pattern and the shared base model.

Recommendations

  • Set memory limits per adapter and test fairness and tail latency under mixed load.

References

Read the original article on Hugging Face

Search the Web Archive