bk99.de entertain the web since 1997

Training CodeParrot from scratch

Summary

CodeParrot documents how a GPT-2-like code generator was built, from data collection to training. The article also shows how licence filtering and duplicates shape the data. The model was trained on publicly available Python code from GitHub.

Ideas

  • For code models, data curation is at least as important as the architecture.
  • Duplicates can distort evaluation and overweight frequently copied code.

Insights

  • For code generators, data hygiene and execution checks determine the security gain more than model size.

Facts

  • Among other things, the pipeline filters files by size and licence notices.

Critique

  • Repository metadata does not prove that every piece of code it contains can be used legally or safely.

Recommendations

  • Check origin, licences, secrets and overlaps with test data before every code model training.

References

Read the original article on Hugging Face

Search the Web Archive