GKE Standard clusters that no longer send system metrics to Cloud Monitoring
What does ZopNight detect here?
GKE clusters are flagged at high severity when ZopNight records Cloud Monitoring as disabled, typically after `--monitoring=NONE`. Google states that with system metrics off, basic CPU, memory and disk usage are unavailable for the cluster, even though ingesting GKE system metrics under the `kubernetes.io` prefix costs nothing.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1220 |
| Category | compliance |
| Severity | high |
| Metric | none — pure configuration read |
| Threshold | cluster monitoring recorded as disabled |
| Source | ZopNight |
| Permissions used | container.clusters.list · container.clusters.get |
Where it applies
Free metrics that disappear with one flag
When a GKE cluster is created, it collects metrics from system components by default and sends
them to Cloud Monitoring under the kubernetes.io prefix. Google’s
metrics configuration page
notes that Cloud Monitoring does not charge for ingesting these system metrics, and that if you
disable them, basic information such as CPU, memory and disk usage is no longer available for the
cluster.
Two more consequences are easy to miss. Google warns that with Cloud Logging or Cloud Monitoring disabled, GKE support is offered on a best-effort basis. And Autopilot clusters cannot disable system metrics at all, so this finding almost always concerns a Standard cluster.
Checking a cluster’s metric packages
gcloud container clusters describe CLUSTER_NAME --location=LOCATION \ --format="value(monitoringConfig.componentConfig.enableComponents)"SYSTEM_COMPONENTS in the result means system metrics are on. Other packages, such as
APISERVER, or POD and DEPLOYMENT, which need Managed Service for Prometheus, are optional
extras on top.
When ZopNight raises it
ZopNight records a monitoring status for each GKE cluster during inventory and fires only on an explicit disabled value. No utilisation data or time window is involved; the finding is about the collection setting itself.
Situations where it stays silent
A cluster with no recorded monitoring status is skipped. The rule does not assess alerting policies, dashboards or retention. Missing logs are a separate finding: GKE Cluster Logging Disabled.
No bill, only blindness
The saving is $0, and turning system metrics back on does not add ingestion charges for them. What it restores is the evidence needed to right-size nodes and to see an outage coming.
Restoring system metrics
- Enable system metrics:
gcloud container clusters update CLUSTER_NAME --location=LOCATION --monitoring=SYSTEM. - Add
API_SERVERor other packages if you use them, for example--monitoring=SYSTEM,API_SERVER,POD; the Pod and workload packages require Managed Service for Prometheus. - Check that the cluster’s observability metrics show node CPU and memory usage again.
- Review alerting policies built on
kubernetes.iometrics; they had no data while collection was off.