API Doc
Search Docs...
⌘ K
OverviewDeploymentManagementOperationReferenceGlossary

Viewing the active-active cluster information

In addition to the cluster information, you can also view the health status of the cluster, which is based on the ACOS monitoring and alert service. The service rules are the same for ACOS clusters without the active-active feature enabled. You can find more detailed information about the availability zones on their specific page in AOC.

Procedure

  1. Click an active-active cluster to access its page, select Active-active status from the tab bar.

  2. Select a component or network from the topology diagram. On the pop-up overview page, you can view the basic information and monitoring details of the availability zones and nodes, network traffic between availability zones, and ping status between the witness node and availability zones.

    Component or network Information Details
    Primary or secondary availability zone Basic information

    Numbers of hosts and virtual machines

    vCPUs and memory allocated to active virtual machines

    Data center

    Numbers of running and suspended virtual machines

    Total storage space

    Used, invalid, and available storage space

    Witness node Basic information

    IP address

    Allocated CPU

    Allocated memory

    Total system storage space

    Used and available system storage space

    Monitoring information vCPU and memory usage of the node within two hours
    Network between availability zones Network traffic

    Network traffic from the primary availability zone to the secondary availability zone

    Network traffic from the secondary availability zone to the primary availability zone

    Network between witness node and availability zones Ping status

    Maximum round-trip latency from the witness node to the primary or secondary availability zone

    Maximum round-trip latency from the primary or secondary availability zone to the witness node

  3. Check the status of the availability zones, the network connectivity between the primary and the secondary availability zones, and the network connectivity between the witness node and the availability zones. In addition, pay attention to and address info and critical alerts to recover the cluster health in time.

    Component or network Health check item Details
    Primary or secondary availability zone Memory resources

    Displays whether the available memory in the other availability zone is sufficient to host the running and suspended virtual machines when the current availability zone fails.

    • Healthy: The available memory is sufficient to host all the running and suspended virtual machines.
    • Notice: The available memory is insufficient to host all the running and suspended virtual machines.
    Total data space

    Displays whether the data space in the other availability zone is sufficient to restore all current data to three replicas when the current availability zone fails.

    • Healthy: The available data space is sufficient to restore all data to three replicas.
    • Notice: The available data space is insufficient to restore all data to three replicas.
    • Critical: The available data space is seriously insufficient to restore all data to the two replicas.
    Host system time

    Displays whether the system time of all hosts in the current availability zone differs from the cluster's system time by more than three seconds.

    • Healthy: No time differences are greater than three seconds.
    • Critical: The system time of some hosts differs from the cluster's system time by more than three seconds, which requires special attention.
    Host power state

    Displays the power state of all hosts in the current availability zone.

    • Healthy: All hosts are powered on.
    • Notice: Some hosts are powered off.
    Host status

    Displays the status of all hosts in the current availability zone.

    • Healthy: All hosts are in the running state.
    • Critical: The status of some hosts is unknown.
    Health status of host storage

    Displays the health status of all hosts' storage services in the current availability zone.

    • Healthy: The storage services of all hosts are available.
    • Critical: The storage services of some hosts are abnormal.
    Status of the zbs-metad, zookeeper, and mongod services on all master nodes

    Displays whether the zbs-metad, zookeeper, and mongod services on all master nodes are running in the current availability zone.

    • Healthy: The zbs-metad, zookeeper, and mongod services on all master nodes are running.
    • Critical: The zbs-metad, zookeeper, or mongod service on some master nodes is not running.
    Witness node Witness node status

    Displays the status of the witness node.

    • Healthy: The witness node is in the running state.
    • Critical: The status of the witness node is unknown.
    System time of witness node

    Displays whether the system time of the witness node differs from the cluster's system time by more than three seconds.

    • Healthy: The time difference is no greater than three seconds.
    • Critical: The system time difference is greater than three seconds, which requires special attention.
    Status of the witness node's mongod, zookeeper, ntpd, and master-monitor system services

    Displays whether the mongod, zookeeper, ntpd, and master-monitor services on the witness node are running.

    • Healthy: The mongod, zookeeper, ntpd, and master-monitor services on the witness node are running.
    • Critical: The mongod, zookeeper, ntpd, or master-monitor service on the witness node is not running.
    CPU usage of witness node

    Displays the witness node's CPU usage and the lasting time.

    • Healthy: The CPU usage of the witness node does not exceed 80%, or the CPU usage has once exceeded 80% but never lasted for more than five minutes.
    • Notice: The CPU usage of the witness node exceeds 80% and has lasted for more than five minutes.
    • Critical: The CPU usage of the witness node exceeds 90% and has lasted for more than five minutes.
    Memory usage of witness node

    Displays the memory usage of the witness node and the lasting time.

    • Healthy: The memory usage of the witness node does not exceed 80%, or the memory usage has once exceeded 80% but never last for more than five minutes.
    • Notice: The memory usage of the witness node exceeds 80% and has lasted for more than five minutes.
    • Critical: The memory usage of the witness node exceeds 90% and has lasted for more than five minutes.
    System disk usage of witness node

    Displays whether the system disk usage of the witness node exceeds 90%.

    • Healthy: The system disk usage of the witness node does not exceed 90%.
    • Notice: The system disk usage of the witness node exceeds 90%.
    Network connectivity between availability zones Maximum round-trip latency between availability zones

    Displays the average maximum round trip delay between hosts in the current availability zone and the other zone in the past five minutes.

    • Healthy: In the past five minutes, the average maximum round trip delay between hosts in the current availability zone and the other zone has not exceeded 10 milliseconds.
    • Info: In the past five minutes, the average maximum round trip delay between hosts in the current availability zone and the other zone has exceeded 10 milliseconds.
    • Notice: In the past five minutes, the average maximum round trip delay between hosts in the current availability zone and the other zone has exceeded 200 milliseconds.
    • Critical: In the past five minutes, the average maximum round trip delay between hosts in the current availability zone and the other zone has exceeded two seconds.
    Network connection between availability zones

    Displays whether the storage network connection between the primary availability zone and the secondary availability zone is normal.

    • Healthy: The storage network connection between the primary availability zone and the secondary availability zone is normal.
    • Critical: The storage network connection between the primary availability zone and the secondary availability zone is abnormal.
    Network between witness node and availability zones Maximum round-trip latency between the witness node and availability zones

    Displays the average maximum round trip delay from the witness node to any host in the primary availability zone and the secondary availability zone in the past five minutes.

    • Healthy: In the past five minutes, the average maximum round trip delay from the witness node to any host in the primary availability zone and the secondary availability zone has not exceeded 10 milliseconds.
    • Critical: In the past five minutes, the average maximum round trip delay from the witness node to any host in the primary availability zone and the secondary availability zone has exceeded two seconds.
    • Notice: In the past five minutes, the average of the maximum round trip delay from the witness node to any host in the availability zones has exceeded 200 milliseconds.
    • Critical: In the past five minutes, the average maximum round trip delay from the witness node to any host in the availability zones has exceeded two seconds.
    Network connection between the witness node and availability zones

    Displays whether the storage network connection is normal between the witness node and the primary availability zone, and between the witness node and the secondary availability zone.

    • Healthy: The storage network connection is normal between the witness node and the primary availability zone, and between the witness node and the secondary availability zone.
    • Critical: The storage network connection is abnormal between the witness node and the primary availability zone, and between the witness node and the secondary availability zone.