Azure Local / Azure Stack HCI – Expired Kubernetes API Certificate Causes kube-vip Failure and Resource Bridge VIP Unreachable

David Gould 0 Reputation points
2026-10-02T17:18:52.17+00:00

Hello everyone,

I'm troubleshooting an Azure Local (Azure Stack HCI) environment running on a three-node Dell APEX cluster. I'm hoping to hear from other administrators who have experienced a similar issue and successfully recovered without rebuilding their cluster.

Environment:

Three-node Azure Stack HCI cluster

Dell APEX infrastructure

Azure Resource Bridge deployed

Kubernetes version 1.30.4

kube-vip version 0.8.0

Cluster registration and Azure connectivity are healthy

Existing Hyper-V workloads remain operational

The issue:

Our Azure Resource Bridge Kubernetes API certificate expired on July 11, 2026. At the time of expiration, the kube-vip component began reporting TLS certificate validation failures and subsequently lost its leadership lease.

The Resource Bridge control-plane virtual IP is no longer reachable.

The Kubernetes API responds successfully when accessed directly through the appliance node, including /livez and /readyz.

However, the control-plane virtual IP does not respond to ARP requests, including requests originating from the appliance itself.

Relevant kube-vip logs:

tls: failed to verify certificate:
x509: certificate has expired or is not yet valid

failed to renew lease kube-system/plndr-cp-lock

failed to renew lease kube-system/plndr-svcs-lock

leader lost
lost leadership, restarting kube-vip

The certificate expiration and leadership failure occurred within seconds of each other.

Azure Local update failure:

A subsequent Azure Local update failed during:

UpdateArbAndExtensions
Role: MocArb
Operation: RotateArbCredentials

The failing operation was:

az arcappliance update-infracredentials hci

The operation could not retrieve the Kubernetes configuration because connections to the control-plane virtual IP timed out.

Troubleshooting completed:

Confirmed Azure Stack HCI registration and connectivity are healthy.

Confirmed the Resource Bridge VM is running.

Verified Kubernetes API availability directly through the appliance node.

Confirmed the Kubernetes API server certificate has expired.

Confirmed kube-vip was previously advertising the control-plane virtual IP.

Identified kube-vip leadership failures beginning immediately after certificate expiration.

Collected and reviewed Azure Resource Bridge appliance diagnostic logs.

Confirmed the virtual IP is not currently being advertised.

My questions:

Has anyone experienced an expired Kubernetes API certificate causing kube-vip to lose leadership on Azure Local or Azure Stack HCI?

Were you able to recover the Resource Bridge without rebuilding the entire Azure Stack HCI cluster?

If you successfully recovered, what procedure or tools did you use?

Did recovery require replacing the Resource Bridge appliance, or was certificate renewal sufficient?

Were existing virtual machines and Azure management resources preserved?

I'm particularly interested in hearing about real-world recovery experiences from other Azure Local administrators.

I understand this is a managed appliance and that unsupported manual modifications can introduce additional risks.

Thanks in advance for any experiences or recommendations. Hello everyone,

I'm troubleshooting an Azure Local (Azure Stack HCI) environment running on a three-node Dell APEX cluster. I'm hoping to hear from other administrators who have experienced a similar issue and successfully recovered without rebuilding their cluster.

Environment:

Three-node Azure Stack HCI cluster

Dell APEX infrastructure

Azure Resource Bridge deployed

Kubernetes version 1.30.4

kube-vip version 0.8.0

Cluster registration and Azure connectivity are healthy

Existing Hyper-V workloads remain operational

The issue:

Our Azure Resource Bridge Kubernetes API certificate expired on July 11, 2026. At the time of expiration, the kube-vip component began reporting TLS certificate validation failures and subsequently lost its leadership lease.

The Resource Bridge control-plane virtual IP is no longer reachable.

The Kubernetes API responds successfully when accessed directly through the appliance node, including /livez and /readyz.

However, the control-plane virtual IP does not respond to ARP requests, including requests originating from the appliance itself.

Relevant kube-vip logs:

tls: failed to verify certificate:
x509: certificate has expired or is not yet valid

failed to renew lease kube-system/plndr-cp-lock

failed to renew lease kube-system/plndr-svcs-lock

leader lost
lost leadership, restarting kube-vip

The certificate expiration and leadership failure occurred within seconds of each other.

Azure Local update failure:

A subsequent Azure Local update failed during:

UpdateArbAndExtensions
Role: MocArb
Operation: RotateArbCredentials

The failing operation was:

az arcappliance update-infracredentials hci

The operation could not retrieve the Kubernetes configuration because connections to the control-plane virtual IP timed out.

Troubleshooting completed:

Confirmed Azure Stack HCI registration and connectivity are healthy.

Confirmed the Resource Bridge VM is running.

Verified Kubernetes API availability directly through the appliance node.

Confirmed the Kubernetes API server certificate has expired.

Confirmed kube-vip was previously advertising the control-plane virtual IP.

Identified kube-vip leadership failures beginning immediately after certificate expiration.

Collected and reviewed Azure Resource Bridge appliance diagnostic logs.

Confirmed the virtual IP is not currently being advertised.

My questions:

Has anyone experienced an expired Kubernetes API certificate causing kube-vip to lose leadership on Azure Local or Azure Stack HCI?

Were you able to recover the Resource Bridge without rebuilding the entire Azure Stack HCI cluster?

If you successfully recovered, what procedure or tools did you use?

Did recovery require replacing the Resource Bridge appliance, or was certificate renewal sufficient?

Were existing virtual machines and Azure management resources preserved?

I'm particularly interested in hearing about real-world recovery experiences from other Azure Local administrators.

I understand this is a managed appliance and that unsupported manual modifications can introduce additional risks.

Thanks in advance for any experiences or recommendations.

Azure Local
0 comments No comments

2 answers

Sort by: Most helpful
  1. David Gould 0 Reputation points
    2026-10-02T18:19:43.3266667+00:00

    Thanks for the response. We've already confirmed the Kubernetes API certificate expired on July 11, and the kube-vip logs show it immediately began failing TLS validation and subsequently lost leadership.

    The Kubernetes API itself is still responding directly through the appliance node, but the control-plane VIP is no longer being advertised.

    We've also confirmed this is affecting the Azure Resource Bridge update process, specifically UpdateArbAndExtensions during RotateArbCredentials.

    What I'm really trying to determine is whether anyone has successfully recovered from this situation without rebuilding the entire Azure Stack HCI cluster.

    Have you personally encountered this issue, or are you aware of a supported procedure to renew the Kubernetes API certificate on an existing Azure Local Resource Bridge appliance?

    Was this answer helpful?

    0 comments No comments

  2. Ankita 20 Reputation points
    2026-10-02T17:48:12.41+00:00

    This sounds like the expired API certificate is affecting the kube-vip control-plane VIP rather than the Kubernetes API itself, especially since the API is still reachable directly from the appliance node. I’d first check the certificate status and kube-vip logs across all three nodes, then verify whether renewing/rotating the Resource Bridge Kubernetes API certificate restores the kube-vip leadership and VIP connectivity. I’d also avoid rebuilding the cluster until the certificate and kube-vip state have been confirmed. If anyone has gone through the same Azure Stack HCI scenario, it would be useful to know the supported certificate renewal/recovery steps.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.