This section provides you with some operations to locate the fault when a cluster becomes unavailable.
Possible causes are listed in order of likelihood.
If the fault persists after you have ruled out a cause, check other causes.
If the fault persists, contact the customer service to help you locate the fault.
The name of this security group is in the format of cluster_name-cce-control-ID.
For details, see How Do I Modify Cluster Security Group Rules?
Symptom
The resource usage on the master nodes in the cluster reaches 100%.
Possible Cause
When a cluster has a large number of resources created simultaneously, it causes an overload on the API server. This, in turn, overloads the master nodes and leads to OOM issues.
Solution
Increase the cluster management scale. A larger cluster management scale means higher capacity and improved performance of the master nodes. For details, see Changing Cluster Scale.
If a cluster is overloaded, you can submit a service ticket for technical support.