When the data redundancy level in the cluster falls below the expected level, and there is available physical disk capacity on the nodes, the system automatically triggers data recovery to restore the expected redundancy level as soon as possible.
The cluster is recovering data.
During data recovery, the security level of volumes with a replication factor below the expected level is slightly reduced. For newly written data, the system typically enables a temporary replication mechanism to maintain the same data security level as the healthy state. However, for existing data, recovery triggered by replica loss or other anomalies may leave the data at a higher risk level than in a healthy state.
In addition, data access performance is slightly reduced for data in the recovery state.
Data recovery can be triggered by either of the following causes:
Planned maintenance operations
For example, when shutting down, restarting, or performing maintenance on a node or storage service, immediate recovery of hot data may be triggered regardless of whether maintenance mode is enabled. In addition, when you increase the replication factor, data recovery is also triggered to bring the replication factor up to the desired level.
Failure events
Data recovery for physical disks is almost always triggered by failures. Such data recovery is typically accompanied by other alert messages; refer to the related alerts to further determine the cause. Common causes include:
Data recovery triggered by planned maintenance operations
Typically, no additional action is required. Once the maintenance operation is complete, the system automatically returns to the normal state.
Data recovery accompanied by node, physical disk, network, or storage service alerts
Resolve the related node, physical disk, network, or storage service failures first. Once the failures are resolved and data recovery is complete, no new alerts will appear, and no further action is required.
Data recovery not accompanied by other alerts
The issue is most likely caused by temporary network fluctuations or sudden hardware anomalies. Try identifying any related anomalies by checking the host kernel logs. If necessary, contact an Arcfra after-sales engineer to identify the specific cause.