bk99.de entertain the web since 1997

Training LLaMA with RLHF

Summary

StackLLaMA combines supervised fine-tuning, a reward model and PPO into a complete RLHF pipeline. The guide shows how PEFT and quantisation lower the hardware requirements. The pipeline first uses instruction data for supervised fine-tuning.

Ideas

  • RLHF consists of several interdependent models and data phases.
  • A faulty reward model can encourage strategies that look convincing but are unwanted.

Insights

  • Scalable training needs reproducible transitions between code, data, devices and checkpoints.

Facts

  • PPO then optimises answers against a learned reward model.

Critique

  • Preference data reflects the values and blind spots of the people annotating it.

Recommendations

  • Evaluate each training stage separately and look specifically for reward hacking.

References

Read the original article on Hugging Face

Search the Web Archive