Transient DNS resolution failure to platform endpoint 168.xx.xx.xx— VM

Archana Chandran 0 Reputation points
2026-10-09T06:35:58.65+00:00

On 2026-10-09 between 04:35:05 and 04:49:53 UTC, VM experienced continuous DNS resolution timeouts. The VM's only configured nameserver is 168.XX.XXX.XX. We have ruled out disk exhaustion, memory/OOM, connection tracking table exhaustion, and ARP/neighbor table overflow on the guest OS. Resource Health for this VM shows no platform fault during this window — only our own subsequent manual reboot at 04:49–04:51 UTC, which resolved the issue. We would like to understand whether there was a host-side or SDN-level event affecting this VM's connectivity to the platform DNS relay during this window.

Secondary/lower-priority note: the VM has also logged recurring page_pool_release_retry() stalled pool shutdown kernel messages on the hv_netvsc interface since approximately October 2, continuously and unchanged — unrelated to the acute incident but possibly worth a look.

We were able to fix the issue by restarting the server but would like to understand why a transient loss of connectivity to Azure's platform DNS relay happened.

Azure Virtual Machines
Azure Virtual Machines

An Azure service that is used to provision Windows and Linux virtual machines.

0 comments No comments

Answer accepted by question author
Alex Burlachenko 25,370 Reputation points MVP Volunteer Moderator
2026-10-09T11:26:56.4766667+00:00

Hi Archana Chandran & thx for join me at Q&A platform,

if the address is 168.63.129.16, that's Azure's platform virtual IP used for Azure-provided DNS and other infrastructure services. A timeout reaching it can indicate a problem in the VM's networking path, but it doesn't automatically mean there was an Azure host or SDN outage. One important Resource Health showing no platform fault doesn't rule out a transient networking issue. It also doesn't confirm one. And the fact that a reboot fixed it could point to either a temporary platform-side condition or something stuck in the guest's networking stack.

I'd pay some attention to those hv_netvsc messages, even tho they started earlier. They're related to the Hyper-V network driver, and recurring page_pool_release_retry() warnings are worth investigating alongside the DNS timeouts. I'd check the kernel logs around 04:35 04:50 UTC for interface resets, driver warnings, link changes, or packet drops. Also worth checking whether other outbound connections failed during that window or only DNS queries.

If this happens again, a packet capture showing DNS requests to 168.63.129.16 leaving the VM without responses would be especially useful. Comparing UDP and TCP port 53 could help narrow it down too.

For the October 9 incident, I'd open an Azure VM networking support case and ask Microsoft to correlate 04:35:05–04:49:53 UTC with host networking, SDN, and platform DNS telemetry. Include the VM resource ID, region, guest OS/kernel version, and relevant hv_netvsc logs.Unfortunately, there's no customer accessible log that can definitively confirm a host-side DNS relay failure after the fact. That part needs Azure engineering to investigate. I wouldn't treat the reboot as proof that the root cause was fixed, either.

rgds,

Alex

Was this answer helpful?

1 person found this answer helpful.

0 additional answers

Sort by: Most helpful

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.