Matching subjective server outages against monitoring data
Summary
Manuel Schmitt describes a customer who reports frequent outages, although monitoring shows continuous availability and no fault had been reported before. The operator’s monitoring showed the server as permanently reachable. The customer had not reported any individual faults beforehand.
Ideas
- Central monitoring and user perception often observe different network paths.
- Late, collective reports make it hard to assign problems to specific points in time.
- A first complaint can already claim a series of suspected outages in its wording.
Insights
- Contradictory observations are a reason for joint measurement rather than hasty blame.
- Good incident reporting needs time, destination, source, error pattern and affected path.
Facts
- Manuel suspected the cause outside his own infrastructure but waited for feedback.
Recommendations
- When faults occur, ask for exact timestamps, source networks and error messages.
- Complement internal monitoring with external measuring points and end-to-end tests.
References
Read the original article on Hostblogger
Links to the original source and the Web Archive open in a new tab.