API Doc
Search Docs...
⌘ K
OverviewDeploymentManagementOperationReferenceGlossary
    ACOS 6.3.0
  • Acrfra Cloud Operation System cluster>
  • ACOS fault handling

Troubleshooting principles

When a failure occurs in an ACOS cluster, the primary principle is to minimize the impact on cluster services and restore them as quickly as possible. In addition, you must also follow the requirements below:

  • Confirm the failure symptoms and identify the root cause. Avoid blind operations (such as arbitrary shutdown, restart, or hot-plugging disks) when the cause is unclear, as this may worsen the situation.

  • If this document provides a solution, follow the documented procedure. If the scenario is not covered, contact an Arcfra after-sales engineer for assistance.

  • During fault handling, thoroughly record all raw information. Do not delete data or logs arbitrarily.

  • For hardware failures on servers (except the Arcfra Halo HCI appliance), contact the hardware vendor's after-sales engineer as soon as possible, confirm the solution together with an Arcfra after-sales engineer, and then proceed with the hardware replacement.

  • After the cluster resumes operation, continue monitoring for a period of time to confirm that the failure has been fully resolved.

Collecting environment information

After a failure occurs in the cluster, collect and organize the following information and promptly report it to your Arcfra after-sales engineer.

CategoryDetails to collect
Cluster versionThe software version.
Failure occurrence timeRecord the exact time when the cluster failure occurred.
Failure symptomsProvide a detailed description of the specific manifestations when the failure occurred in the cluster.
Operations before the failureRecall and record the specific operations performed in the cluster before the failure occurred, along with their results.
Operations after the failureRecord the specific operations performed in the cluster after the failure occurred, along with their results.

Collecting cluster logs

To identify the cause of the cluster failure as quickly as possible, collect cluster logs and node information as follows:

  • Log type: cluster OS logs, service logs, and coredump logs

  • Node scope: all nodes in the cluster

  • Service scope: logs for all services

  • Time range: 12 hours before and 12 hours after the failure occurred

Collect and download cluster logs, and promptly submit them to your Arcfra after-sales engineer when reporting the failure.