API Doc
Search Docs...
⌘ K
OverviewDeploymentManagementOperationReferenceGlossary
    ACOS 6.3.0
  • Acrfra Cloud Operation System cluster>
  • ACOS fault handling>
  • Service anomalies

I/O rerouting service anomaly

Description

When ACOS is deployed on the VMware vSphere virtualization platform, an I/O rerouting service is deployed on each ESXi host to monitor the health status of the storage service on that host in real time. If a storage service anomaly is detected, the storage service automatically redirects I/O requests to other healthy nodes in the cluster, preventing storage anomalies from directly triggering virtual machine HA, and reducing the impact of planned maintenance operations (such as upgrades and restarts) or other storage failures on services.

Alert message

  • SCVM with storage IP {.labels._data_ip}: The IO rerouting service on its host stops working.

  • SCVM with storage IP {.labels._data_ip}: IO rerouting service on its host is not directed to the current node as expected.

Impact

  • If the I/O rerouting service on the host where the SCVM resides is not directed to the current node, the I/O performance of the host may degrade.

  • When the I/O rerouting service stops completely, if the SCVM on that host experiences an anomaly or the storage service experiences a prolonged anomaly (the specific duration depends on the HA policy configured in VMware vSphere, typically on the order of minutes), virtual machine HA will be triggered.

Cause

Possible causes include:

  • The I/O rerouting service is not correctly deployed.

  • The storage service on the alerted SCVM is in an abnormal state.

Solution

  • The alert is normal if it appears during a cluster upgrade or storage service maintenance. If the alert disappears on its own after the operation is complete, no additional action is needed.

  • If the alert is not triggered by a planned operation, it is typically accompanied by other storage, network, or SCVM anomaly alerts. Investigate and resolve those issues first and this alert is expected to disappear automatically after the underlying issue is resolved.

  • If troubleshooting does not resolve the issue or the alert persists, contact an Arcfra after-sales engineer for further assistance.