Wireless Access

 View Only
  • 1.  Standby AAC Not assigned on all APs on new cluster

    Posted Apr 24, 2023 11:54 AM
    Edited by GG-62df62 Apr 24, 2023 04:36 PM

    Hello,

    Please ignore the subject (senior moment from me), but the issues described below are still relevant

    new cluster AOS 8.10.0.6   (existing cluster 8.10.0.5)
    1 x 9240 in active new 'cluster' and the same for the new standby cluster
    Active and standby conductors ArubaMM-HW-5K
    All physical appliances

    We have a new cluster (we have nearly reached the AP limit on our existing cluster so this is to take some of the load), we have transferred some APs (700-odd) APs from the existing cluster to this new cluster. Things seem to be working but I noticed an issue today, in the logs of APs (of the APs that I have checked) that are on the new cluster we are seeing lots of this type of event:

    Mon Apr 24 13:04:49 2023 System Up
    Mon Apr 24 13:04:48 2023 System Device rebooted while marked as down: Change in number of reboots detected (got 2, expected 1)
    Mon Apr 24 13:04:48 2023 System Remote LAN IP changed from 172.x.x.x to 172.x.x.x.
    Mon Apr 24 13:04:01 2023 System Tunnel IP changed from 172.x.x.x to 172.x.x.x.
    Mon Apr 24 13:04:01 2023 System AP's associated controller is changed from x-A2 to x-C1
    Mon Apr 24 12:50:08 2023 System Status changed to 'AP is no longer associated with controller'
    Mon Apr 24 12:50:08 2023 System Down


    So the APs look like they are trying to switch back to the existing (old) cluster - that cluster is on different firmware so the AP then looks to be up/downgrading which makes the whole problem particularly messy. The APs switch back and forwards like this a lot. Things appear to 'work' on the new cluster, at least well enough that we haven't been inundated with complaints, but clearly there is an issue (and we are starting to get some reports through now).

    A lot of the APs were still in the allowlist and gap-db on the older cluster which I guess was exacerbating the issue because the AP would not be blocked from talking to the old cluster and would downgrade, I have now removed the APs from both the db and allowlist.

    Any ideas why this might be happening, or what we might check? I'm confused as to why they would still try to contact the cluster they were previously on. Could it be that previous cluster members are still listed in the AP's memory somewhere?


    Guy




  • 2.  RE: Standby AAC Not assigned on all APs on new cluster

    Posted Apr 28, 2023 06:36 AM

    How are your APs discovering the controller? DNS, DHCP, Static config?



    ------------------------------
    Herman Robers
    ------------------------
    If you have urgent issues, always contact your Aruba partner, distributor, or Aruba TAC Support. Check https://www.arubanetworks.com/support-services/contact-support/ for how to contact Aruba TAC. Any opinions expressed here are solely my own and not necessarily that of Hewlett Packard Enterprise or Aruba Networks.

    In case your problem is solved, please invest the time to post a follow-up with the information on how you solved it. Others can benefit from that.
    ------------------------------



  • 3.  RE: Standby AAC Not assigned on all APs on new cluster

    Posted Apr 28, 2023 07:32 AM

    Initially we provision APs on provisioning VLANs where they pick up the cluster address via DHCP option 43. These APs would have been provisioned in that way when they were first installed in the locations involved (ie all the APs in locations that have now been shifted to the new cluster). Then a few weeks ago, to move them across to the new cluster we used a new provisioning profile to set the Conductor IP (which is an aname pointing at the new clusters).




  • 4.  RE: Standby AAC Not assigned on all APs on new cluster

    Posted Apr 28, 2023 08:00 AM

    If the APs are flipping between the old and new cluster (and controller versions), there should be a reason for that. If you do static controller assignment, you should point the AP to the controller (one, round-robin DNS, VRRP), but not to the Mobility Conductor:
    If you use DNS/DHCP/ADP, make sure that these only point to the new controllers. It could be that if you configured static to the mobility conductor, that one does not answer and the AP may fallback to another discovery and find the old controllers.



    ------------------------------
    Herman Robers
    ------------------------
    If you have urgent issues, always contact your Aruba partner, distributor, or Aruba TAC Support. Check https://www.arubanetworks.com/support-services/contact-support/ for how to contact Aruba TAC. Any opinions expressed here are solely my own and not necessarily that of Hewlett Packard Enterprise or Aruba Networks.

    In case your problem is solved, please invest the time to post a follow-up with the information on how you solved it. Others can benefit from that.
    ------------------------------



  • 5.  RE: Standby AAC Not assigned on all APs on new cluster

    Posted Apr 28, 2023 08:52 AM

    It's kind of odd because it isn't clear that there is anything service affecting actually happening - despite the messages saying the APs were switching controllers (back to the 'old' system). There would be several of these types of messages a day, and messages about downgrading (which makes sense as the old system is on a minor revision behind the new one). I removed the APs from the whitelist and db of the old system and the messages about switching controller have stopped (as have the downgrade messages) but instead we are seeing a lot of these:

    Fri Apr 28 13:32:15 2023 System Status changed to 'OK'
    Fri Apr 28 13:32:15 2023 System Up
    Fri Apr 28 11:24:59 2023 System Status changed to 'AP is no longer associated with controller'
    Fri Apr 28 11:24:59 2023 System Down
    Fri Apr 28 11:12:04 2023 System Status changed to 'OK'
    Fri Apr 28 11:12:04 2023 System Up
    Fri Apr 28 08:17:12 2023 System Status changed to 'AP is no longer associated with controller'
    Fri Apr 28 08:17:12 2023 System Down
    Fri Apr 28 08:11:43 2023 System Status changed to 'OK'
    Fri Apr 28 08:11:43 2023 System Up
    Fri Apr 28 07:02:51 2023 System Status changed to 'AP is no longer associated with controller'
    Fri Apr 28 07:02:51 2023 System Down

     
    So perhaps previously the attempted switch to the old system was as a result of the Down status's that we continue to see, but we no longer see them trying to connect to the old system because they are no longer in the whitelist/db and so can't talk to it (but under the hood perhaps they still make an attempt?).

    It's strange that the APs think they are Down so often. And what is stranger is that throughout all of this (even when the APs were saying that they were changing controller and downgrading etc) we have had no reports of issues on wireless, and we have two UX sensors in areas that would be affected but there have been zero issues reported by them either! So it _seems_ as if all is fine, but something is clearly going on, at least according to the AP logs!

    I did wonder whether because the new cluster is actually just a cluster of 1 maybe the APs have retained some cluster members from the old system in memory as their in-cluster failover options. But even if that's true it doesn't explain why the new system sees so many Down log messages.

    ......aaaah so that got me thinking about in-cluster failover, and it looks like that is turned on for the new cluster with 0ms as the heartbeat threshold (on our main system this is set to 3072ms). If we have a cluster of 1 (though this may increase at some point in the future) should we turn cluster redundancy off? Or leave it on and change the heartbeat threshold?