What Google learned from a hundred thousand disk failures
Summary
A Google study uses large-scale field data to show which SMART values predict disk failures and which common assumptions do not hold. Google examined more than a hundred thousand hard drives over nine months. Drives with their first scan errors failed considerably more often in the short term.
Ideas
- Real operating data separates lab assumptions from actual failure patterns.
- Individual SMART signals correlate strongly but capture only part of later defects.
- Scan errors clearly increase the subsequent risk of failure.
- Low average temperatures do not guarantee a longer lifetime.
- High utilisation explains failures less reliably than is often assumed.
- Prediction models must explicitly account for false-negative drives.
Insights
- Telemetry is better suited to risk assessment than to reliable individual prediction.
- Large samples can disprove persistent operational myths.
- An unremarkable health value never replaces a tested recovery strategy.
- Maintenance decisions need probabilities, costs and impact together.
Facts
- Many failed drives showed no strong SMART warning signs beforehand.
- Temperature and usage correlated less with failures than expected.
Recommendations
- Treat SMART warnings as a replacement signal, not as a complete failure forecast.
- Test restores regularly and monitor replication independently of drive health.
- Decide replacement thresholds based on your own fleet data and the cost of damage.
References
- BBC: Hard disc test surprises Google
- Google Research: Failure Trends in a Large Disk Drive Population
Links to the original source and the Web Archive open in a new tab.