When a failure occurs in an ACOS cluster, the primary principle is to minimize the impact on cluster services and restore them as quickly as possible. In addition, you must also follow the requirements below:
Confirm the failure symptoms and identify the root cause. Avoid blind operations (such as arbitrary shutdown, restart, or hot-plugging disks) when the cause is unclear, as this may worsen the situation.
If this document provides a solution, follow the documented procedure. If the scenario is not covered, contact an Arcfra after-sales engineer for assistance.
During fault handling, thoroughly record all raw information. Do not delete data or logs arbitrarily.
For hardware failures on servers (except the Arcfra Halo HCI appliance), contact the hardware vendor's after-sales engineer as soon as possible, confirm the solution together with an Arcfra after-sales engineer, and then proceed with the hardware replacement.
After the cluster resumes operation, continue monitoring for a period of time to confirm that the failure has been fully resolved.
After a failure occurs in the cluster, collect and organize the following information and promptly report it to your Arcfra after-sales engineer.
| Category | Details to collect |
|---|---|
| Cluster version | The software version. |
| Failure occurrence time | Record the exact time when the cluster failure occurred. |
| Failure symptoms | Provide a detailed description of the specific manifestations when the failure occurred in the cluster. |
| Operations before the failure | Recall and record the specific operations performed in the cluster before the failure occurred, along with their results. |
| Operations after the failure | Record the specific operations performed in the cluster after the failure occurred, along with their results. |
To identify the cause of the cluster failure as quickly as possible, collect cluster logs and node information as follows:
Log type: cluster OS logs, service logs, and coredump logs
Node scope: all nodes in the cluster
Service scope: logs for all services
Time range: 12 hours before and 12 hours after the failure occurred
Collect and download cluster logs, and promptly submit them to your Arcfra after-sales engineer when reporting the failure.