Category: Incident Response
-
The Server That Never Talked: ARP TTL, FDB Aging, and a Broadcast Storm
An ARP TTL FDB aging mismatch turned a data center into a light show. Around 2005, every top-of-rack switch in the data center was blinking like a light show at a disco rave. I traced the cause back to a server doing exactly what its owner designed it to do. The failure wasn’t a misconfiguration,…
-
When the Interior Doesn’t Speak PIM: SPBM Multicast Root Cause
SPBM Multicast and a Replacement Plan That Assumed the Wrong Protocol A hospital system spanning several states runs its live IPTV distribution out of a single metro region’s headend: DirecTV and BlonderTongue encoder headends feeding 60+ active multicast streams. When the core node serving that region was rip-and-replaced, removing legacy Avaya-era SPBM hardware out, standing…
-
1500 != 1500: MTU, OSPF ExStart, and a 14-Byte Blind Spot
What OSPF is actually doing when it stalls in EXSTART, why MTU is the non-obvious suspect, and what to check first when you hit it.
-
When the Fix Becomes the Failure: ECMP, Zone Protection, and a 64KB Ceiling
Three ECMP firewall cutovers went cleanly. The fourth did not — and the cause turned out to be Palo Alto MSS clamping, hidden inside a zone protection profile that had passed through three clean rollouts undetected. It was the highest profile pair in the sequence, sitting at the data center boundary. The Setup The Baylor…
-
BGP Prefix Leak, RPKI, and the Cold Email That Confirmed It
A BGP route leak at an Internet2 customer site propagated more-specific prefixes into the global table, causing every major CDN to black-hole return traffic to an entire campus network. Diagnosing it required a cold email to a Google network engineer at 8:40am. This is the full story: the triage, the root cause, and the architectural…
