API Doc
Search Docs...
⌘ K
OverviewDeploymentManagementOperationReferenceGlossary
    ACOS 6.3.0
  • Acrfra Cloud Operation System cluster>
  • ACOS fault handling>
  • Hosts

MegaRAID controller reset failure on a host

Description

The system detected a MegaRAID controller reset failure in the host kernel logs, triggering the alert.

Alert message

Hardware fault: failed to reset the MegaRAID controller of the host {.labels.hostname}. The error log recorded in the OS is: { .labels.message }

Impact

  • There is a risk posed to the host's access to physical disks.

  • If the operating system is not deployed on a separate M.2 disk, the overall operation of the host is also at risk.

  • During the MegaRAID controller anomaly, services on the host may fail as a whole.

Cause

When I/O requests from the host to physical disks remain unresponsive for an extended period, or when the data link returns certain errors, a MegaRAID controller reset may be triggered to clear I/Os that failed to complete normally. Before the reset is performed, physical disks connected to the controller may experience prolonged I/O blocking.

Common causes include:

  • A physical disk is not responding to I/O requests.

  • The MegaRAID firmware or driver fails.

Solution

Check the logs to confirm whether there are physical disk error messages that triggered the MegaRAID controller reset. If any, replace the physical disk with a new one.

If the error messages involve multiple physical disks, there may be a controller firmware or driver anomaly. Contact the relevant hardware vendor to confirm whether to upgrade the firmware or replace the controller.