This section shows you how to locate the fault when a cluster becomes unavailable.
Possible causes are listed in order of likelihood.
Check these causes one by one until you find the cause of the fault.
If the fault persists, contact the customer service to help you locate the fault.
The name of this security group is in the format of {cluster_name}-cce-control-{ID}.
For details, see How Do I Modify Cluster Security Group Rules?
Symptom
The resource usage on the master nodes in the cluster reaches 100%.
Possible Cause
When a cluster has a large number of resources created simultaneously, it causes an overload on the API server. This, in turn, overloads the master nodes and leads to OOM issues.
Solution
Increase the cluster management scale. A larger cluster management scale means higher capacity and improved performance of the master nodes. For details, see Changing a Cluster Scale.
If a cluster is overloaded, you can submit a service ticket for technical support.