Hello everyone,
I'm troubleshooting an Azure Local (Azure Stack HCI) environment running on a three-node Dell APEX cluster. I'm hoping to hear from other administrators who have experienced a similar issue and successfully recovered without rebuilding their cluster.
Environment:
Three-node Azure Stack HCI cluster
Dell APEX infrastructure
Azure Resource Bridge deployed
Kubernetes version 1.30.4
kube-vip version 0.8.0
Cluster registration and Azure connectivity are healthy
Existing Hyper-V workloads remain operational
The issue:
Our Azure Resource Bridge Kubernetes API certificate expired on July 11, 2026. At the time of expiration, the kube-vip component began reporting TLS certificate validation failures and subsequently lost its leadership lease.
The Resource Bridge control-plane virtual IP is no longer reachable.
The Kubernetes API responds successfully when accessed directly through the appliance node, including /livez and /readyz.
However, the control-plane virtual IP does not respond to ARP requests, including requests originating from the appliance itself.
Relevant kube-vip logs:
tls: failed to verify certificate:
x509: certificate has expired or is not yet valid
failed to renew lease kube-system/plndr-cp-lock
failed to renew lease kube-system/plndr-svcs-lock
leader lost
lost leadership, restarting kube-vip
The certificate expiration and leadership failure occurred within seconds of each other.
Azure Local update failure:
A subsequent Azure Local update failed during:
UpdateArbAndExtensions
Role: MocArb
Operation: RotateArbCredentials
The failing operation was:
az arcappliance update-infracredentials hci
The operation could not retrieve the Kubernetes configuration because connections to the control-plane virtual IP timed out.
Troubleshooting completed:
Confirmed Azure Stack HCI registration and connectivity are healthy.
Confirmed the Resource Bridge VM is running.
Verified Kubernetes API availability directly through the appliance node.
Confirmed the Kubernetes API server certificate has expired.
Confirmed kube-vip was previously advertising the control-plane virtual IP.
Identified kube-vip leadership failures beginning immediately after certificate expiration.
Collected and reviewed Azure Resource Bridge appliance diagnostic logs.
Confirmed the virtual IP is not currently being advertised.
My questions:
Has anyone experienced an expired Kubernetes API certificate causing kube-vip to lose leadership on Azure Local or Azure Stack HCI?
Were you able to recover the Resource Bridge without rebuilding the entire Azure Stack HCI cluster?
If you successfully recovered, what procedure or tools did you use?
Did recovery require replacing the Resource Bridge appliance, or was certificate renewal sufficient?
Were existing virtual machines and Azure management resources preserved?
I'm particularly interested in hearing about real-world recovery experiences from other Azure Local administrators.
I understand this is a managed appliance and that unsupported manual modifications can introduce additional risks.
Thanks in advance for any experiences or recommendations.
Hello everyone,
I'm troubleshooting an Azure Local (Azure Stack HCI) environment running on a three-node Dell APEX cluster. I'm hoping to hear from other administrators who have experienced a similar issue and successfully recovered without rebuilding their cluster.
Environment:
Three-node Azure Stack HCI cluster
Dell APEX infrastructure
Azure Resource Bridge deployed
Kubernetes version 1.30.4
kube-vip version 0.8.0
Cluster registration and Azure connectivity are healthy
Existing Hyper-V workloads remain operational
The issue:
Our Azure Resource Bridge Kubernetes API certificate expired on July 11, 2026. At the time of expiration, the kube-vip component began reporting TLS certificate validation failures and subsequently lost its leadership lease.
The Resource Bridge control-plane virtual IP is no longer reachable.
The Kubernetes API responds successfully when accessed directly through the appliance node, including /livez and /readyz.
However, the control-plane virtual IP does not respond to ARP requests, including requests originating from the appliance itself.
Relevant kube-vip logs:
tls: failed to verify certificate:
x509: certificate has expired or is not yet valid
failed to renew lease kube-system/plndr-cp-lock
failed to renew lease kube-system/plndr-svcs-lock
leader lost
lost leadership, restarting kube-vip
The certificate expiration and leadership failure occurred within seconds of each other.
Azure Local update failure:
A subsequent Azure Local update failed during:
UpdateArbAndExtensions
Role: MocArb
Operation: RotateArbCredentials
The failing operation was:
az arcappliance update-infracredentials hci
The operation could not retrieve the Kubernetes configuration because connections to the control-plane virtual IP timed out.
Troubleshooting completed:
Confirmed Azure Stack HCI registration and connectivity are healthy.
Confirmed the Resource Bridge VM is running.
Verified Kubernetes API availability directly through the appliance node.
Confirmed the Kubernetes API server certificate has expired.
Confirmed kube-vip was previously advertising the control-plane virtual IP.
Identified kube-vip leadership failures beginning immediately after certificate expiration.
Collected and reviewed Azure Resource Bridge appliance diagnostic logs.
Confirmed the virtual IP is not currently being advertised.
My questions:
Has anyone experienced an expired Kubernetes API certificate causing kube-vip to lose leadership on Azure Local or Azure Stack HCI?
Were you able to recover the Resource Bridge without rebuilding the entire Azure Stack HCI cluster?
If you successfully recovered, what procedure or tools did you use?
Did recovery require replacing the Resource Bridge appliance, or was certificate renewal sufficient?
Were existing virtual machines and Azure management resources preserved?
I'm particularly interested in hearing about real-world recovery experiences from other Azure Local administrators.
I understand this is a managed appliance and that unsupported manual modifications can introduce additional risks.
Thanks in advance for any experiences or recommendations.