Linux crisis tools
Summary
Brendan Gregg recommends a prepared toolbox so that performance outages can be measured immediately instead of improvised. Gregg's minimum list includes procps, util-linux, sysstat, iproute2 and tracing packages. Once loaded, BCC and bpftrace tools produce comparable BPF bytecode.
Ideas
- Diagnostic tools belong in every server image before an outage.
- procps and sysstat provide quick CPU, memory, process and device measurements.
- iproute2 and tcpdump cover network state and packet traffic.
- perf, BCC and bpftrace open up scheduler, kernel and application paths.
- Installing tools afterwards can fail because of load, DNS, firewalls or immutable systems.
Insights
- Diagnostic readiness is a property of the system, not a spontaneous ability of the operator.
- Small image costs buy particularly valuable minutes during outages.
- Missing measurement tools encourage risky reboots and unproven explanations.
Facts
- Special accelerators also need diagnostic tools of their own.
Recommendations
- Install and test crisis tools when creating the server image.
- Document a first measurement procedure for CPU, memory, disks and network.
References
Links to the original source and the Web Archive open in a new tab.