API Doc
Search Docs...
⌘ K
OverviewDeploymentManagementOperationReferenceGlossary
    ACOS 6.3.0
  • Acrfra Cloud Operation System cluster>
  • ACOS fault handling>
  • Cluster performance and load

Cluster data recovery

Description

When the data redundancy level in the cluster falls below the expected level, and there is available physical disk capacity on the nodes, the system automatically triggers data recovery to restore the expected redundancy level as soon as possible.

Alert message

The cluster is recovering data.

Impact

During data recovery, the security level of volumes with a replication factor below the expected level is slightly reduced. For newly written data, the system typically enables a temporary replication mechanism to maintain the same data security level as the healthy state. However, for existing data, recovery triggered by replica loss or other anomalies may leave the data at a higher risk level than in a healthy state.

In addition, data access performance is slightly reduced for data in the recovery state.

Cause

Data recovery can be triggered by either of the following causes:

  • Planned maintenance operations

    For example, when shutting down, restarting, or performing maintenance on a node or storage service, immediate recovery of hot data may be triggered regardless of whether maintenance mode is enabled. In addition, when you increase the replication factor, data recovery is also triggered to bring the replication factor up to the desired level.

  • Failure events

    Data recovery for physical disks is almost always triggered by failures. Such data recovery is typically accompanied by other alert messages; refer to the related alerts to further determine the cause. Common causes include:

    • Physical disk failure
    • NIC failure
    • Memory failure
    • Node failure
    • Storage service anomaly

Solution

  • Data recovery triggered by planned maintenance operations

    Typically, no additional action is required. Once the maintenance operation is complete, the system automatically returns to the normal state.

  • Data recovery accompanied by node, physical disk, network, or storage service alerts

    Resolve the related node, physical disk, network, or storage service failures first. Once the failures are resolved and data recovery is complete, no new alerts will appear, and no further action is required.

  • Data recovery not accompanied by other alerts

    The issue is most likely caused by temporary network fluctuations or sudden hardware anomalies. Try identifying any related anomalies by checking the host kernel logs. If necessary, contact an Arcfra after-sales engineer to identify the specific cause.