Training LLaMA with RLHF
Summary
StackLLaMA combines supervised fine-tuning, a reward model and PPO into a complete RLHF pipeline. The guide shows how PEFT and quantisation lower the hardware requirements. The pipeline first uses instruction data for supervised fine-tuning.
Ideas
- RLHF consists of several interdependent models and data phases.
- A faulty reward model can encourage strategies that look convincing but are unwanted.
Insights
- Scalable training needs reproducible transitions between code, data, devices and checkpoints.
Facts
- PPO then optimises answers against a learned reward model.
Critique
- Preference data reflects the values and blind spots of the people annotating it.
Recommendations
- Evaluate each training stage separately and look specifically for reward hacking.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.