Incidents we were called in to solve

Five write-ups from live hosting and virtualization environments: what broke, how we traced it, and what changed afterwards. Written the way we hand them over internally, with the root cause and the numbers intact.

Mail Reputation Containment

Two shared IPs blacklisted by Spamhaus and Barracuda after a compromised contact form turned one account into a spam relay. Contained, delisted and rate-limited within 18 hours.

18 h To confirmed blacklist removal

Hidden DNS Dependency Outage

Domains stopped resolving for a high-value downstream customer. The suspected cause was our own recent migration. The real cause was a healthy, symptom-free server nobody had connected to the authority chain.

0 Rollbacks required

Live SAN Migration

60+ production VMs moved to a new SAN backend across a 5-day window, with no maintenance window and zero customer-reported downtime.

60+ VMs migrated live

WordPress Performance Recovery

A 15-second WordPress site that plugin-disabling could not fix. The bottleneck was 4.2 million orphaned rows in wp_options. Page load went from 12 seconds to 1.8.

1.8 s Page load, down from 12 s

Hypervisor I/O Stalls

Random I/O stalls across VMs on one KVM node, with no pattern by workload. The cause was one failing disk behind a single Ceph OSD. No data lost.

40 VMs returned to stable I/O

Have something that does not add up?

Most of the work above started as a symptom nobody could reproduce. If that sounds familiar, we should talk.