The alert is triggered when a physical disk sustains physical damage, causing command timeouts, I/O errors, or critical device errors.
Common alert messages for physical disk hardware damage are as follows:
Hardware fault: the physical disk { .labels._dev } on the host { .labels._hostname } encountered an I/O error on the sector { .labels._sector }. The latest error log recorded in the OS is: { .labels.message }
Hardware fault: the physical disk { .labels._dev } on the host { .labels._hostname } encountered a command timeout. The latest error log recorded in the OS is: { .labels.message }
Hardware fault: the physical disk { .labels._dev } on the host { .labels._hostname } encountered a critical device error on the sector { .labels._sector }. The latest error log recorded in the OS is: { .labels.message }
When a physical disk experiences a hardware failure, command timeouts and I/O errors may occur depending on the severity of the damage. The system may fail to recognize the damaged disk, causing it to go offline. Once a disk goes offline, disk quarantine and data recovery are also triggered.
The physical disk has sustained hardware damage.
Physical disks typically reserve redundant space to replace a small number of damaged sectors or data blocks. When the number of errors does not reach the system threshold, the disk firmware automatically activates redundant sectors to continue operation. If only a small number of I/O errors occur and the error count has not reached the threshold for triggering an unhealthy or failing state, the disk is still available and does not require immediate attention.
If the physical disk is confirmed to be severely damaged, unmount and replace the faulty disk. After the replacement, manually resolve the alert.