Ein Azure-Dienst, der verwendet wird, um Telemetriedaten aus Azure und lokalen Umgebungen zu sammeln, zu analysieren und auf sie zu reagieren
AKS platform metric alert creation fails for node_memory_working_set_percentage although metric definition exists
We can no longer create Azure Monitor metric alerts for the AKS platform metric node_memory_working_set_percentage.
Metric namespace:
Microsoft.ContainerService/managedClusters
Metric name:
node_memory_working_set_percentage
Affected resource type:
Microsoft.ContainerService/managedClusters
The metric definition exists for this AKS resource. This command returns the metric successfully:
az monitor metrics list-definitions --resource "$AKS_ID" --namespace Microsoft.ContainerService/managedClusters --query "[?name.value=='node_memory_working_set_percentage']" -o jsonc
It shows:
- category: Nodes PREVIEW
- dimensions: node and nodepool
- supported aggregations: Maximum and Average
- primaryAggregationType: Average
- unit: Percent
But creating the metric alert fails:
az monitor metrics alert create \
-g "$ALERT_RG" \
-n "ha-cloud-aks_node_memory_working_set-dev620" \
--scopes "$AKS_ID" \
--severity 3 \
--evaluation-frequency 1m \
--window-size 5m \
--condition "avg Microsoft.ContainerService/managedClusters.node_memory_working_set_percentage greater than 100 with skipmetricvalidation" \
--action "$ACTION_GROUP_ID"
Actual CLI condition used greater-than operator. The portal form rejects angle brackets, so I spell it out here.
Error:
BadRequest: The metric with the name node_memory_working_set_percentage does not exist.
Activity IDs:
4ec17149-f5c3-4132-84fb-40fb8c9889bd
1473fee4-e26f-44c6-a54b-ed50f3ba49da
The same failure occurs from Terraform azurerm_monitor_metric_alert, even with skip_metric_validation set to true.
Things tried:
- Average aggregation
- Maximum aggregation
- with skipmetricvalidation
- no dimension filter
- where node includes wildcard
- where nodepool includes wildcard
- explicit target resource type and region
- same AKS resource with node_memory_rss_percentage
Control test:
A similar alert using node_memory_rss_percentage can be created successfully on the same AKS resource.
Expected behavior:
Because the metric definition exists and the metric is documented for AKS, metric alert creation should accept node_memory_working_set_percentage, or the metric should not appear in list-definitions.
This looks like an Azure Monitor metric alert backend regression for the AKS preview node metric node_memory_working_set_percentage.