Due to an ongoing issue with our VXLAN where the IPsec protocol sometimes desynchronizes causing nodes to loose communication one way, we saw all servers having issues communicating properly with other internal services.
Caused some app-servers to be unable to serve admin, viewer, api and popups
We have an underlying issue with our provider where very small networking glitches can cause the VXLAN IPsec protocol to get out of sync, which causes communications issues between nodes in our cluster.
While we can manually reset this and recover, which we did in this instance, we are still working with our vendors for a permanent fix.
To mitigate this, we are adding even more infrastructure and capacity to prevent more downtime.