bk99.de entertain the web since 1997

What Google learned from a hundred thousand disk failures

Summary

A Google study uses large-scale field data to show which SMART values predict disk failures and which common assumptions do not hold. Google examined more than a hundred thousand hard drives over nine months. Drives with their first scan errors failed considerably more often in the short term.

Ideas

  • Real operating data separates lab assumptions from actual failure patterns.
  • Individual SMART signals correlate strongly but capture only part of later defects.
  • Scan errors clearly increase the subsequent risk of failure.
  • Low average temperatures do not guarantee a longer lifetime.
  • High utilisation explains failures less reliably than is often assumed.
  • Prediction models must explicitly account for false-negative drives.

Insights

  • Telemetry is better suited to risk assessment than to reliable individual prediction.
  • Large samples can disprove persistent operational myths.
  • An unremarkable health value never replaces a tested recovery strategy.
  • Maintenance decisions need probabilities, costs and impact together.

Facts

  • Many failed drives showed no strong SMART warning signs beforehand.
  • Temperature and usage correlated less with failures than expected.

Recommendations

  • Treat SMART warnings as a replacement signal, not as a complete failure forecast.
  • Test restores regularly and monitor replication independently of drive health.
  • Decide replacement thresholds based on your own fleet data and the cost of damage.

References

Read the original article

Search the Web Archive