How GitHub sped up its application
Summary
GitHub describes the overhaul of its early architecture through measurement, caching, background work and targeted database optimisation. At the time the application was based mainly on Ruby on Rails. Memcached played an important role for frequently read data.
Ideas
- Profilers map slow pages to concrete methods and queries.
- Caches avoid repeatedly computing frequently read repository data.
- Background jobs remove expensive work from interactive requests.
- Database indexes considerably speed up known access patterns.
- Denormalised values trade extra write effort for faster read paths.
- Incremental changes keep improvements measurable and reversible.
Insights
- Performance work starts with latency distributions, not architectural fashions.
- Caching only pays off with understood validity and clear invalidation.
- Asynchrony improves response times but shifts failures into other systems.
- Growing products need observability before they need more infrastructure.
Facts
- Background processes took over non-interactive tasks.
Recommendations
- Profile slow endpoints before any structural change.
- Define cache lifetime and invalidation together.
- Monitor background jobs for age, errors and retries.
References
Links to the original source and the Web Archive open in a new tab.