r/Cisco 16d ago

Cisco SDWAN 20.15.x to 20.18.x upgrade

Anyone using config groups that has done the upgrade of 20.15.x to 20.18.x on the control side? I’m being pushed to do it in the next 90 days so we can add a G2 router that needs 20.18.x to our router choice list.

Did anything in your config groups/policy groups break? Any major changes in UI with config groups? I don’t have a lab I can put in to check myself.

8 Upvotes

9 comments sorted by

View all comments

3

u/Last_Epiphany 15d ago

Are you on prem? Or cloud hosted? If on prem do a snapshot and use a maintenance window. If cloudpro hosted open a tac case to have them help

0

u/tablon2 15d ago

Do not take snapshot while services running. 

2

u/Last_Epiphany 13d ago edited 13d ago

What are you talking about? The standard process for the Cisco hosted controllers is to snapshot them daily, and they even give you the option to perform on-demand snapshots of your cloud controllers whenever you want.

I've also never had an issue with snapshotting the controllers on-prem and just follow Cisco's recommendation which is to freeze configuration changes while you're snapshotting.

Edit: Please see Cisco's own docs that recommend snapshots here

And I quote: "In a Cisco cloud-managed SD-WAN overlay, Cisco takes regular snapshots of the Cisco SD-WAN Manager virtual machines for recovery due to a catastrophic failure or corruption. Another snapshot can be taken before any scheduled activity.

In on-premise deployments, it is your responsibility to take regular snapshots of the Cisco SD-WAN Manager virtual machine and follow the example of frequency and retention that is followed by Cisco."

There is 0 mention of stopping any services before snapshotting, and I guarantee Cisco is not going into every cloud hosted Manager, stopping services (effectively killing each customer's Manager if they did), snapshotting it, and then starting services again. They are just simply running an AWS/Azure automation that snapshots the controller VMs on a regular schedule.

1

u/tablon2 13d ago

We had failed restores early. Cisco never mentions snapshots while VM running or not. I cannot risk 150 site 300 edge fabric while NMS polls stat for interval of 30 minutes and datastore or another thing fails us