When ACOS is deployed independently on two physical disks, these two disks serve as system disks and form a hardware RAID 1. The following conditions indicate that a system disk failure has occurred.
A member system disk in the hardware RAID experiences read or write errors or high latency, and is marked as a failed component.
A member physical disk in the hardware RAID fails the S.M.A.R.T. check.
A member physical disk in the hardware RAID shows no signs of read or write timeouts, increased I/O latency, or disk failure, but the system has determined that its remaining lifetime is insufficient and that issues are imminent.
The physical disk { .labels._serial } in hardware RAID on the host { .labels.hostname } failed the S.M.A.R.T. test.
The physical disk { .labels._serial } in hardware RAID on the host { .labels.hostname } has a remaining lifespan below { .threshold }.
The hardware RAID virtual disk { .labels._vd_name } on the host { .labels.hostname } has insufficient redundancy.
The hardware RAID virtual disk { .labels._vd_name } on the host { .labels.hostname } is abnormal: { .labels.status }.
When either of the two system disks fails, the OS automatically switches to the other system disk on the node for read and write operations. Services are not affected throughout this process.
When both system disks fail simultaneously, the operating system partition and metadata partition become completely unavailable, rendering the node inoperable. If the virtual machines on the failed node are configured with HA, they can be migrated to other nodes to continue running normally. Because the node reduces the number of available replicas of some data blocks below the expected replica count, data recovery is triggered.
Refer to the corresponding alert message.
You can check the status of the system disk and its member physical disks via AOC or the BMC system to determine whether a hardware RAID failure has occurred.
AOC
View the status of the Arcfra system disk and its member physical disks on the host overview page and the system disk details page in AOC.
BMC system
Use the server's BMC interface to confirm the storage controller model and the status and information of the RAID and physical disks under it.
You can query ACOS-supported storage controllers using the Hardware Compatibility Checker. The table below lists several common storage controllers used for Arcfra system disks and their corresponding BMC systems. BMC system interfaces vary by vendor; refer to the respective vendor documentation for operation.
| Vendor | BMC system | RAID controller |
|---|---|---|
| Dell PowerEdge | iDRAC | Marvell SATA M.2 (Dell BOSS-S1/S2) |
| Dell PowerEdge | iDRAC | Marvell NVMe M.2 (Dell BOSS-N1) |
| xFusion | iBMC | Broadcom MegaRAID (LSI SAS3XXX) |
When the failure occurs:
In most cases, you only need to replace the faulty physical disk. The RAID controller will automatically handle the relevant configurations.
In some cases, you may need to configure RAID through the server's BMC or BIOS interface.