Five write-ups from live hosting and virtualization environments: what broke, how we traced it, and what changed afterwards. Written the way we hand them over internally, with the root cause and the numbers intact.
Two shared IPs blacklisted by Spamhaus and Barracuda after a compromised contact form turned one account into a spam relay. Contained, delisted and rate-limited within 18 hours.
18 h To confirmed blacklist removal
Domains stopped resolving for a high-value downstream customer. The suspected cause was our own recent migration. The real cause was a healthy, symptom-free server nobody had connected to the authority chain.
0 Rollbacks required
60+ production VMs moved to a new SAN backend across a 5-day window, with no maintenance window and zero customer-reported downtime.
60+ VMs migrated live
A 15-second WordPress site that plugin-disabling could not fix. The bottleneck was 4.2 million orphaned rows in wp_options. Page load went from 12 seconds to 1.8.
1.8 s Page load, down from 12 s
Random I/O stalls across VMs on one KVM node, with no pattern by workload. The cause was one failing disk behind a single Ceph OSD. No data lost.
40 VMs returned to stable I/O
Most of the work above started as a symptom nobody could reproduce. If that sounds familiar, we should talk.