Serving thirty LoRA models with one server
Summary
Text Generation Inference loads many LoRA adapters on top of a shared base model and serves them together. This avoids complete model copies and improves utilisation. The demonstration serves 30 LoRA models through one TGI instance.
Ideas
- Adapter multi-tenancy shares expensive base weights between specialised services.
- Schedulers have to batch different adapters together without unfair latency spikes.
Facts
- All adapters use the same loaded base model.
Critique
- The number of possible adapters depends on size, traffic pattern and the shared base model.
Recommendations
- Set memory limits per adapter and test fairness and tail latency under mixed load.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.