API Doc
Search Docs...
⌘ K
OverviewDeploymentManagementOperationReferenceGlossary
    ACOS 6.3.0
  • Arcfra Cloud Operating System>
  • Internel service notes

Service failure impact

Block storage services

Service name Single-node failure impact Cluster-wide failure impact
zbs-metad
  • Failure on the Leader node: The storage service becomes briefly inaccessible (approximately 40 seconds).
  • Failure on Follower nodes: No impact
Scenario: Fewer than half of the meta nodes alive
  • The cluster storage becomes unavailable.
  • All virtual machine I/Os time out.
  • Operations related to ACOS virtual machine services and storage services become unavailable.
zbs-chunkd
  • Data on the failed node cannot be retrieved.
  • Data recovery is initiated.
  • Virtual machine I/Os experience a temporary lag.
  • The disk storage service on the local hypervisor is taken over by the remote service, which may cause a temporary lag.

Scenario: Failure on all nodes

  • Storage becomes unavailable and I/Os are interrupted.
zbs-taskd The failed node cannot execute long-running tasks.

Scenario: Failure on all nodes

  • The cluster loses its long-running task scheduling and execution capabilities.
vipservice The failed node cannot host the VIP service.

Scenario: Failure on all nodes

  • The VIP service becomes unavailable.
zbs-iscsi-redirectord
  • The failed node cannot establish new iSCSI connections.
  • Existing iSCSI connections on the failed node remain unaffected.

Scenario: Failure on all nodes

  • The cluster cannot establish new iSCSI connections.
  • Existing iSCSI connections in the cluster remain unaffected.
zbs-aurora-monitord Because the HA feature for virtual machines in vhost mode on failed nodes becomes unavailable, the virtual machines cannot perform read or write operations upon a storage failure.

Scenario: Failure on all nodes

  • Virtual machines in vhost mode cannot start.
  • Started virtual machines cannot perform read or write operations.
zbs-inspectord The failed node cannot perform data inspection.

Scenario: Failure on all nodes

The cluster loses its data inspection capability.
zbs-aurorad Because the HA feature for virtual machines in vhost mode on failed nodes becomes unavailable, the virtual machines cannot perform read or write operations upon a storage failure.

Scenario: Failure of zbs-chunkd and zbs-aurorad on all nodes

The cluster becomes unavailable and I/Os are interrupted.
timemachine If timemachine fails on the Meta Leader node, the cluster cannot provide the scheduled snapshot service. Failure on other nodes has no impact.

Scenario: Failure on all nodes

  • The scheduled snapshot function fails.

zbs-watchdogd If the zbs-watchdogd service fails on the ZooKeeper Leader node, when the node's operating system becomes abnormal, the ZooKeeper service may not be able to re-elect a leader quickly. -

Virtual machine services

Service name Single-node failure impact Cluster-wide failure impact
job-center-worker
  • You cannot manage virtual machines on the failed node.

  • You cannot modify the failed node, such as changing its management IP address, altering its VDS, or inserting and removing disks.

  • System resources and service information of the failed node cannot be updated to the database.

  • The status of virtual machines on the failed node cannot be updated to the database.

-
job-center-scheduler No impact

Scenario: Failure on all nodes

  • Virtual machine management becomes unavailable.

  • The monitoring and alerting function fails.

elf-vm-monitor Virtual machines on the failed node cannot trigger HA.

Scenario: Normal functioning on fewer than two nodes

  • Virtual machine HA in the cluster fails.

elf-vm-scheduler
  • If the service fails on the Meta Leader node, automatic scheduling of virtual machines fails.

  • If the service fails on other nodes, the failure has no impact.

-
master-monitor If the network connection is lost between the primary availability zone and the secondary availability zone, the REST API may become unavailable. -
elf-dhcp
  • If the service fails on the Meta Leader node, the VM network DHCP configuration function fails.

  • If the service fails on other nodes, the failure has no impact.

Scenario: Failure on the Leader node

The DHCP configuration function cannot be provided for VM networks.
elf-exporter During the failure, the monitoring data on the failed node cannot be retrieved.

Scenario: Failure on all nodes

Monitoring fails on the entire cluster.
vnc-proxy The failed node cannot provide the VNC Web console service. -
vmtools-agent
  • The host cannot report the static information and performance data of its virtual machines in real time.

  • You cannot configure network information such as IP and DNS for the virtual machines on this node.

Scenario: Failure on all nodes

  • Static data and performance data of all virtual machines in the cluster cannot be retrieved in real time through Arcfra VMTools.

  • You cannot configure network information such as IP and DNS for all virtual machines in the cluster.

elf-fs Failure has no impact on virtual machine import and export of the failed node or the entire cluster. The failed node does not respond to virtual machine import and export requests while CloudTower sends requests to other healthy nodes.

Scenario: Failure on all nodes

OVF import and export of the cluster become unavailable.

Operations and maintenance services

Service name Single-node failure impact Cluster-wide failure impact
zbs-rest-server
  • The API on the failed node becomes inaccessible.

  • During the failure, the performance data on the failed node cannot be stored.

-
zbs-deploy-server You cannot add hosts. -
cluster-upgrader -

Scenario: Failure on all nodes

  • You cannot perform one-click upgrade.

  • Historical upgrade records cannot be displayed.

tuna-rest-server The hardware information API on the failed node becomes inaccessible. -
usbredir-manager Virtual machines cannot access the USB devices on the failed node. -
ntpm You cannot modify the NTP configuration on the failed node.

Scenario: Failure on all nodes

  • The cluster cannot provide the NTP HA service.

  • The cluster may experience storage system unavailability due to clock desynchronization between nodes.

Monitoring services and microservices

Service name Single-node failure impact Cluster-wide failure impact
octopus The octopus API on this node becomes unavailable, and monitoring data cannot be queried through this node.

Scenario: Failure on all nodes

  • No new monitoring data is generated.

  • No new alerts are generated.

  • Monitoring data cannot be queried.

oscar The oscar service on this node becomes unavailable. Scenario: Failure on all nodes

All oscar services become unavailable.

siren
  • Failure on the Meta Leader node:
    • Alert messages from octopus cannot be received.
    • Alerts that have been resolved cannot be displayed as Resolved.
    • Alerts cannot be marked as Resolved.
  • Failure on other nodes: The siren API on the node becomes unavailable.

Scenario: Failure on the Leader node

  • No new alert notification emails are sent.

  • Alert information cannot be queried.

  • Alerts cannot be marked as Resolved.

aquarium The aquarium service becomes unavailable on this node, and virtual machines on this node cannot access the aquarium service.

Scenario: Failure on all nodes

  • Registration information cannot be provided.

harbor All V3 APIs on this node become unavailable.

Scenario: Failure on all nodes

  • The V3 API cannot be provided.

crab Most APIs related to the node become unavailable.

Scenario: Failure on all nodes

  • Most APIs of the cluster become unavailable.

dolphin
  • Failure on the Meta Leader node: Advanced monitoring cannot be deployed or disabled.

  • Failure on other nodes: The dolphin API on the node becomes unavailable.

Scenario: Failure on the Leader node

  • Advanced monitoring cannot be deployed or disabled.

  • The dolphin API becomes unavailable.

seal No impact

Scenario: Failure on all nodes

  • Event audit becomes unavailable.

fluent-bit The audit logs for the current node are not collected.

Scenario: Failure on all nodes

  • The audit logs for all failed nodes in the cluster are not collected.

snmpd The failed node cannot retrieve data through SNMP.

Scenario: Failure on all nodes

All nodes cannot retrieve data through SNMP.
network-monitor Network monitoring data on the failed node cannot be retrieved.

Scenario: Failure on all nodes

  • Network monitoring data cannot be retrieved.

disk-healthd
  • The disk health status of the failed node cannot be updated.

  • The failed node cannot trigger new disk health alerts.

  • Disk health alerts triggered by the failed node may not reflect the current health status of the disk.

-
sd-offline If a disk blocks the HBA card, this disk may not be automatically taken offline. -
consul-exporter During the failure, the failed node cannot retrieve consul-server status monitoring data.

Scenario: Failure on all nodes

consul-server status monitoring fails completely.
tuna-exporter During the failure, the monitoring data on the failed node cannot be retrieved.

Scenario: Failure on all nodes

Monitoring fails on the entire cluster.
svcresctld
  • Monitoring alerts for the service status and resource usage on the failed node become ineffective.
  • Service status information of the failed node cannot be retrieved.
  • Monitoring alerts for the service status and resource usage in the cluster become ineffective.
  • Service status information of the entire cluster cannot be retrieved.
netreactor
  • Failure on Meta Leader: The cluster cannot provide network fail-slow detection and isolation services for this node.

  • Failure on other nodes: No impact

Scenario: Failure on the Leader node

The cluster cannot provide network fail-slow detection and isolation services for this node.

l2ping@storage

Scenario: Failure on all nodes

The node cannot perform network fail-slow detection and isolation for its storage NIC.
All nodes cannot perform network fail-slow detection and isolation for their storage NICs.
l2ping@access The node cannot perform network fail-slow detection and isolation for its access NIC.

Scenario: Failure on all nodes

All nodes cannot perform network fail-slow detection and isolation for their access NICs.
vmagent Monitoring metric data cannot be collected. -
vmagent-prod (observability) Monitoring metric data cannot be collected.

Scenario: Failure on all nodes

Monitoring metrics cannot be collected for all nodes.
vector (observability) The node cannot receive monitoring metrics, collect logs, or send monitoring metrics and logs to the observability virtual machine.

Scenario: Failure on all nodes

All nodes cannot receive monitoring metrics, collect logs, or send the data to the observability virtual machine.
net-health-check The node cannot check the health of the network port, which may affect the results of other node checks. -

Third-party services

Service name Single-node failure impact Cluster-wide failure impact
mongod No impact

Scenario: Failure on half or more of the nodes

  • APIs of all services do not function.

  • Job Center does not function.

nginx The API and Web console of the failed node become unavailable. -
chronyd
  • Failure affects clock synchronization between the failed node and other nodes, resulting in accumulated clock drift.

  • Clock drift leads to:

    • The latency of some I/Os increases.

    • Misjudgment of virtual machine HA causes unexpected virtual machine shutdowns.

-
zookeeper No impact

Scenario: Failure on more than half of the nodes

  • Storage becomes inaccessible for all nodes in the cluster.

  • Most virtual machine operations become unavailable.

  • All virtual machine I/Os time out.

envoy
  • The API on the failed node becomes unavailable.

  • The failed node cannot perform ZBS asynchronous backup.

-
envoy-xds The API on the failed node cannot function properly.

Scenario: Failure on all nodes

  • All APIs cannot function properly.

consul The API on the failed node cannot function properly.

Scenario: Failure on all nodes

  • All APIs cannot function properly.

consul-server No impact

Scenario: Failure on half or more of the nodes

  • All APIs cannot function properly.

libvirtd The virtual machine lifecycle on the failed node cannot be controlled. -
prometheus
  • Failure on the Meta Leader node: All monitoring services become unavailable.
  • Failure on other nodes: No impact

Scenario: Meta Leader fails and no new Meta Leader is elected.

  • No new monitoring data is generated.
  • No new alerts are generated.
  • Monitoring data cannot be queried.
containerd Containers on the failed node cannot be managed. -
everoute-agent
  • The failed node cannot distribute new network security rules.

  • The failed node cannot retrieve IP addresses of virtual machines.

-
everoute-collector The failed node loses its traffic information.

Scenario: Failure on all nodes

The cluster loses its traffic information.
rpcbind On the failed node, services using rpcbind cannot receive I/Os.

Scenario: Failure on all nodes

In the cluster, services using rpcbind cannot receive I/Os.
network-firewall The firewall on the failed node cannot be modified, which may result in desynchronization between the node and cluster settings. -