I have a customer who is having issues with his VMware ESX 4 servers reaching the storage VLAN(VLAN100) using two ProCurve Switches(2810). Originally, everything worked w/o any issues until we migrated his new hardware into his server rack. After the move, vCenter will send my customer several emails about losing connectivity to multiple datastores on the SAN. Everything stays up and running, users don't seem to notice it, but we continue to get spammed with error messages.
I would like to verify the configuration and make sure that we've got everything configured as it should be.
Both switches have ports 1-12 untagged for VLAN100 and NOT configured as DEFAULT_VLAN ports. Ports 13-22 are configured as untagged for the DEFAULT_VLAN and not configured as access ports for VLAN100. Pretty basic but jumbo frames is enabled for VLAN100(and the SAN/ESX Servers).
My worry is that there is something wrong with our redundant trunk links between the switches. We have ports 23-24 configured as Trk1 on both switches. I set Trk1 to be untagged for DEFAULT_VLAN and tagged for VLAN100. When I test everything from my laptop and/or other devices, everything appears to be functional. However, randomly we are getting messages about losing the path to our datastores. The message indicates that all of the servers/datastores are losing connectivity at random times during the day.
Right now, I have everything on one of the 2810Gs and it hasn't generated an error for a couple days. If move any of my storage links to the second switch then after a few hours, the errors will start appearing again.
Summary:
sw1
ports 1-12 DEFAULT_VLAN: No
ports 1-12 VLAN100: Untagged
ports 13-22 DEFAULT_VLAN: Untagged
ports 13-22 VLAN100: No
ports 23-24 configured as Trk1
Trk1 DEFAULT_VLAN: Untagged
Trk1 VLAN100: Tagged
I haven't configured anything else on the switches other than jumbo frames on VLAN100 for both switches. We had a slight power issue on the second switch which we replaced to eliminate that issue. We also had a bad network cable from one of the SAN ports and replaced that as well. I can't see anything else that could cause this issue other than my switch config so any help would be greatly appreciated.