| Service name | Single-node failure impact | Cluster-wide failure impact |
|---|---|---|
| zbs-metad |
|
Scenario: Fewer than half of the meta nodes alive
|
| zbs-chunkd |
|
Scenario: Failure on all nodes
|
| zbs-taskd | The failed node cannot execute long-running tasks. |
Scenario: Failure on all nodes
|
| vipservice | The failed node cannot host the VIP service. |
Scenario: Failure on all nodes
|
| zbs-iscsi-redirectord |
|
Scenario: Failure on all nodes
|
| zbs-aurora-monitord | Because the HA feature for virtual machines in vhost mode on failed nodes becomes unavailable, the virtual machines cannot perform read or write operations upon a storage failure. |
Scenario: Failure on all nodes
|
| zbs-inspectord | The failed node cannot perform data inspection. |
Scenario: Failure on all nodes The cluster loses its data inspection capability. |
| zbs-aurorad | Because the HA feature for virtual machines in vhost mode on failed nodes becomes unavailable, the virtual machines cannot perform read or write operations upon a storage failure. | Scenario: Failure of zbs-chunkd and zbs-aurorad on all nodes The cluster becomes unavailable and I/Os are interrupted. |
| timemachine | If timemachine fails on the Meta Leader node, the cluster cannot provide the scheduled snapshot service. Failure on other nodes has no impact. |
Scenario: Failure on all nodes
|
| zbs-watchdogd | If the zbs-watchdogd service fails on the ZooKeeper Leader node, when the node's operating system becomes abnormal, the ZooKeeper service may not be able to re-elect a leader quickly. | - |
| Service name | Single-node failure impact | Cluster-wide failure impact |
|---|---|---|
| job-center-worker |
|
- |
| job-center-scheduler | No impact |
Scenario: Failure on all nodes
|
| elf-vm-monitor | Virtual machines on the failed node cannot trigger HA. |
Scenario: Normal functioning on fewer than two nodes
|
| elf-vm-scheduler |
|
- |
| master-monitor | If the network connection is lost between the primary availability zone and the secondary availability zone, the REST API may become unavailable. | - |
| elf-dhcp |
|
Scenario: Failure on the Leader node The DHCP configuration function cannot be provided for VM networks. |
| elf-exporter | During the failure, the monitoring data on the failed node cannot be retrieved. | Scenario: Failure on all nodes Monitoring fails on the entire cluster. |
| vnc-proxy | The failed node cannot provide the VNC Web console service. | - |
| vmtools-agent |
|
Scenario: Failure on all nodes
|
| elf-fs | Failure has no impact on virtual machine import and export of the failed node or the entire cluster. The failed node does not respond to virtual machine import and export requests while CloudTower sends requests to other healthy nodes. | Scenario: Failure on all nodes OVF import and export of the cluster become unavailable. |
| Service name | Single-node failure impact | Cluster-wide failure impact |
|---|---|---|
| zbs-rest-server |
|
- |
| zbs-deploy-server | You cannot add hosts. | - |
| cluster-upgrader | - | Scenario: Failure on all nodes
|
| tuna-rest-server | The hardware information API on the failed node becomes inaccessible. | - |
| usbredir-manager | Virtual machines cannot access the USB devices on the failed node. | - |
| ntpm | You cannot modify the NTP configuration on the failed node. |
Scenario: Failure on all nodes
|
| Service name | Single-node failure impact | Cluster-wide failure impact |
|---|---|---|
| octopus | The octopus API on this node becomes unavailable, and monitoring data cannot be queried through this node. |
Scenario: Failure on all nodes
|
| oscar | The oscar service on this node becomes unavailable. |
Scenario: Failure on all nodes All |
| siren |
|
Scenario: Failure on the Leader node
|
| aquarium | The aquarium service becomes unavailable on this node, and virtual machines on this node cannot access the aquarium service. |
Scenario: Failure on all nodes
|
| harbor | All V3 APIs on this node become unavailable. |
Scenario: Failure on all nodes
|
| crab | Most APIs related to the node become unavailable. |
Scenario: Failure on all nodes
|
| dolphin |
|
Scenario: Failure on the Leader node
|
| seal | No impact |
Scenario: Failure on all nodes
|
| fluent-bit | The audit logs for the current node are not collected. |
Scenario: Failure on all nodes
|
| snmpd | The failed node cannot retrieve data through SNMP. | Scenario: Failure on all nodes All nodes cannot retrieve data through SNMP. |
| network-monitor | Network monitoring data on the failed node cannot be retrieved. |
Scenario: Failure on all nodes
|
| disk-healthd |
|
- |
| sd-offline | If a disk blocks the HBA card, this disk may not be automatically taken offline. | - |
| consul-exporter | During the failure, the failed node cannot retrieve consul-server status monitoring data. | Scenario: Failure on all nodes consul-server status monitoring fails completely. |
| tuna-exporter | During the failure, the monitoring data on the failed node cannot be retrieved. | Scenario: Failure on all nodes Monitoring fails on the entire cluster. |
| svcresctld |
|
|
| netreactor |
|
Scenario: Failure on the Leader node The cluster cannot provide network fail-slow detection and isolation services for this node. |
| l2ping@storage | Scenario: Failure on all nodes The node cannot perform network fail-slow detection and isolation for its storage NIC. |
All nodes cannot perform network fail-slow detection and isolation for their storage NICs. |
| l2ping@access | The node cannot perform network fail-slow detection and isolation for its access NIC. | Scenario: Failure on all nodes All nodes cannot perform network fail-slow detection and isolation for their access NICs. |
| vmagent | Monitoring metric data cannot be collected. | - |
| vmagent-prod (observability) | Monitoring metric data cannot be collected. | Scenario: Failure on all nodes Monitoring metrics cannot be collected for all nodes. |
| vector (observability) | The node cannot receive monitoring metrics, collect logs, or send monitoring metrics and logs to the observability virtual machine. | Scenario: Failure on all nodes All nodes cannot receive monitoring metrics, collect logs, or send the data to the observability virtual machine. |
| net-health-check | The node cannot check the health of the network port, which may affect the results of other node checks. | - |
| Service name | Single-node failure impact | Cluster-wide failure impact |
|---|---|---|
| mongod | No impact |
Scenario: Failure on half or more of the nodes
|
| nginx | The API and Web console of the failed node become unavailable. | - |
| chronyd |
|
- |
| zookeeper | No impact |
Scenario: Failure on more than half of the nodes
|
| envoy |
|
- |
| envoy-xds | The API on the failed node cannot function properly. |
Scenario: Failure on all nodes
|
| consul | The API on the failed node cannot function properly. |
Scenario: Failure on all nodes
|
| consul-server | No impact |
Scenario: Failure on half or more of the nodes
|
| libvirtd | The virtual machine lifecycle on the failed node cannot be controlled. | - |
| prometheus |
|
Scenario: Meta Leader fails and no new Meta Leader is elected.
|
| containerd | Containers on the failed node cannot be managed. | - |
| everoute-agent |
|
- |
| everoute-collector | The failed node loses its traffic information. | Scenario: Failure on all nodes The cluster loses its traffic information. |
| rpcbind | On the failed node, services using rpcbind cannot receive I/Os. |
Scenario: Failure on all nodes In the cluster, services usingrpcbind cannot receive I/Os. |
| network-firewall | The firewall on the failed node cannot be modified, which may result in desynchronization between the node and cluster settings. | - |