Sorry to hear that, losing a member mid upgrade with no window to fall back on is about as bad as it gets. Before you kick off the next attempt, get a full backup and image/boot set inventory on the new secondary, and confirm it's running the exact same base build as production before it ever joins the VSX pair, so the upgrade isn't doing a bigger jump than you think.
Ask TAC for the RCA on the dead unit before you touch the second one: if it's a DPU firmware issue specific to that image path, they may tell you to stage the secondary at 10.16 standalone first and rejoin it, instead of letting the automated sequence do it live. Also pull a fresh health check on the primary DPU before you start, so you've got a known good baseline if something goes sideways again.
------------------------------
Dustin Burns
Lead Mobility Engineer @Worldcom Exchange, Inc.
ACCX 1271| ACMX 509| ACSP | ACDA | MVP Guru 2022-2023
If my post was useful accept solution and/or give kudos
------------------------------
Original Message:
Sent: Aug 03, 2026 11:18 AM
From: MR-90c854
Subject: CX10000 AOS-CX 10.16.1051 with VSX/EVPN-VXLAN – Production Experience?
Update to my initial post:
To clarify the context and the urgency behind this thread: The reason I asked is because our first upgrade attempt failed severely during a live production run.
Since our environment is a 24/7 mission-critical infrastructure, we do not use maintenance windows. We have to rely entirely on the VSX cluster's built-in redundancy to perform live upgrades during full production.
During the automated VSX upgrade sequence from 10.13.1130 to 10.16.1051, the secondary CX10000 went completely dark a few minutes after the process started. It became entirely unresponsive, lost all console/management access, and never booted back up.
This left our entire production fabric running on a single, non-redundant primary switch while we scrambled to open an emergency TAC case. Ultimately, the secondary unit was declared dead and required an RMA/hardware replacement.
We now have the replacement hardware in place and are facing our next upgrade attempt. Because we must do this in production again, we are highly alert.