An SE evaluating me for a role told me something about my own resume I’d never noticed. It carries more “held a technical position against management pressure, then was proven right” stories than any resume he’d reviewed, he said, and he called it an unclaimed sales instinct. I wrote about that on LinkedIn. This post is about the two halves of it I didn’t have room for there — both from the same fabric overhaul, the same year, and neither of them a comfortable story to tell alone. While both are about technical leadership, one is about where I held a position, and the other is where I owned a mistake.
The argument I wouldn’t stop making
2024, a campus fabric overhaul on Extreme’s SPBM. The design called for edge ports — the 802.1q connections where hosts actually attach — to terminate on 7520s, with the 7720 core switches running NNI-only, no host-facing ports at all. That split is vendor best practice, and it isn’t arbitrary: it’s the same reason a chassis has line cards and a backplane instead of one board doing both jobs. The backplane moves traffic between cards. It doesn’t try to also be a line card.
My manager wanted to collapse that: terminate edge ports directly on the core, skip a layer, save a hop. I didn’t have the authority to say no. What I had was the argument, made the same way, for as many whiteboard sessions as it took:
- Failure domain collapse. An edge problem — a bad port, a flapping host, a misbehaving switch — now shares a failure domain with the traffic transiting the entire distribution layer. What used to be contained becomes structural.
- Debuggability. A layered design gives you a deterministic place to reason from: if it’s not in the edge, it’s in the core, and vice versa. Collapse the layers and you’ve destroyed the mental model that makes troubleshooting fast.
- Change control scope inflation. A routine access-port change — someone’s new desk phone, a reprovisioned port — becomes a core-change event, with all the review and blast-radius anxiety that implies, for something that used to be nobody’s Tuesday afternoon.
I kept coming back to the same line: collapsing access onto core doesn’t just add risk, it merges two failure domains that were deliberately kept apart. When the edge breaks, you want it to break quietly. On a core switch, nothing breaks quietly.
The design held. NNI-only core went to production the way I’d argued it, on the strength of repetition and the argument itself — no title backed it up.
The same backplane framing did a second job that year. The team was picking up VRF-based segmentation on the same fabric, and the isolated-FIB concept wasn’t landing — a routing table that’s logically separate but physically shares the backbone is not an intuitive thing to hold in your head the first time. I used the same picture: the backbone doesn’t participate in IP routing any more than a chassis backplane does; it moves frames between switches and leaves the routing decisions to the boxes at the edges of that picture. I stopped the analogy there rather than pushing it further into “VRF is a slot boundary” territory, because VRF’s isolation is logical, not physical, and overselling it would have taught the team the wrong mental model to troubleshoot from. It took repeated sessions, not one breakthrough, but the team got to the point where they could reason about the boundary themselves instead of just deferring to me on it.
The mistake I made on the same fabric
The same year, the same rollout, I was on the other side of an incident I caused myself.
I was allocating IP space for the new routed interconnects between the fabric and existing infrastructure. Two addresses I assigned to a pair of routed /31s turned out to belong to a pair of session border controllers — and our homegrown IPAM didn’t clearly show that. I didn’t check somewhere I should have, and I allocated over them. The SBCs went down. Voice was out for most of a day, on a rollout with enough visibility that a day-long voice outage was not a quiet problem.
A multi-discipline war room stood up — server ops, net ops, infosec — because of how long it ran and how visible it was. Everyone knew the SBCs were down. Nobody had yet connected that to the addresses I’d just reallocated, including me.
The server ops manager asked an obvious-sounding question: have we checked the firewall logs? My mental model said no — in the new routed-fabric architecture, the Palo Alto shouldn’t have been anywhere in the path for those addresses. I checked anyway, mostly to close it off.
It wasn’t closed off. The firewall logs showed a legacy policy still tied to those IPs — a NAT configuration I can no longer reconstruct the exact purpose of, but whose presence was unambiguous. Those addresses had an entire prior life in the old architecture, one nobody had documented or cleaned up, and it had survived intact into the new design. The conflict wasn’t a single mistake; it was my new allocation running headlong into an old one nobody remembered was there, on infrastructure that had quietly outlived the reason it existed.
Root cause: a documentation gap in IPAM, compounded by firewall artifacts that stayed invisible until someone checked the one place I’d already decided didn’t apply. SBCs restored, IPAM updated, and I carried the lesson — checking the place you’ve already ruled out is sometimes exactly where the answer is — into the rest of the allocation work on that rollout.
This one’s mine. I made the call, on my own project, and it cost a day of voice service on a system a lot of people were watching. I’m not interested in a version of this where I quietly slide credit for the fix off myself, either: a colleague’s question is the only reason the log entry got read at all.
Technical leadership under pressure: both ways
I’ve written about SPBM before, watching multicast state disappear on somebody else’s fabric I’d been on for four weeks. This time the fabric was one I’d designed, and the boundary I was defending — and, separately, the boundary I got wrong — were both mine to answer for.
Holding a position against a manager and admitting an allocation error in front of a war room don’t look like the same skill from the outside. One reads as backbone; the other reads as an apology. But both come down to the same willingness: to say the true thing about the system in front of you, whether that costs you an argument with your manager or costs you comfort in a room full of people watching you find your own mistake.
I don’t know yet whether that’s the sales instinct the SE saw in my resume, or just what happens to your relationship with being right when you’ve been wrong enough times to know the difference matters more than the outcome. Which one does your own record actually show more of — the times you held, or the times you owned it?

Leave a Reply
You must be logged in to post a comment.