An Azure service that provides serverless Kubernetes, an integrated continuous integration and continuous delivery experience, and enterprise-grade security and governance.
Kubernetes Service (AKS): Connectivity issues with the API Server — Investigation and escalation needed
Problem description
I am experiencing connectivity issues with my AKS cluster, specifically with the API Server. Calls such as kubectl logs, exec, port-forward, and admission webhooks are hanging or failing with 'context deadline exceeded'. Restarting konnectivity-agent and running az aks update did not improve the situation. The problem appears to be related to node-local network connection tracking exhaustion on a specific node, aks-largepool-38709133-vmss000002, which causes API-server-to-node operations to break. Despite attempts to cordon and drain the node, the issues persist, and the failure rate has worsened over time.
Environment
Azure Kubernetes Service; Cannot connect to AKS API Server; Microsoft.ContainerService/managedClusters
What I've already tried
Restarted konnectivity-agent; Ran az aks update (reconcile); Cordoned and drained the affected node; Monitored connection tracking on nodes; Confirmed nodes are not an outlier in utilization; Verified no evidence of blocked egress events.
Current status
The issue is ongoing, and the connection exhaustion on the specific node seems to be the root cause. I am seeking escalation to the AKS product group to investigate the health of the konnectivity-server and the private link path for this cluster, as the problem appears to be related to the control plane or private endpoint connectivity.