Critical Intermittent Azure VM Connectivity Issue – Unable to Connect via Remote Desktop Until VM Restart

Sanjay 20 Reputation points
2026-04-23T11:55:15.31+00:00

Hello Azure Support Team,

We are facing a critical intermittent connectivity issue with one of our Azure Virtual Machines.

VM Details:

<REDACTED VM info>

To protect your confidentiality, sensitive information has been removed from this public thread and securely shared with you via Private message option in this thread, ensuring visibility only between you and the Microsoft engineer assisting you.

Issue Description:

At random times, we are unable to connect to the VM using Remote Desktop (RDP). The server becomes inaccessible, and the RDP connection fails or times out.

Temporary Resolution:

Whenever we restart the VM from Azure Portal, connectivity is restored and we are able to connect again.

Business Impact:

This is a critical production issue because the VM becomes unreachable unexpectedly, causing operational disruption and requiring manual restart intervention.

Request:

Please help us investigate the root cause and suggest a permanent solution. Kindly check whether this could be related to:

  • Azure host / underlying node issue
  • NIC / networking issue
  • Guest VM agent problem
  • Windows RDP service freezing
  • Resource exhaustion or memory leak
  • Storage / disk performance issue
  • Recommended diagnostics and monitoring steps

Additional Note:

Please let us know if you need exact timestamps, logs, or permission to run diagnostics.

Kindly treat this as high priority.

Regards, Sanjay Yadav

Azure Virtual Machines
Azure Virtual Machines

An Azure service that is used to provision Windows and Linux virtual machines.


4 answers

Sort by: Most helpful
  1. Sanjay 20 Reputation points
    2026-06-08T05:53:57.4266667+00:00

    I created E series VM with 128 GB Ram and still getting same issue can you please schedule , i have to restart VM again

    Was this answer helpful?


  2. Sanjay 20 Reputation points
    2026-06-02T05:46:56.78+00:00
    1. Hi Ankit / Nikhil, Thank you for the information regarding the B-series VM limitations. We are evaluating a move to an E-series VM. We currently have the following options available: E8as_v4 (128 GB RAM) E16-8as_v4 (128 GB RAM) E16as_v4 (128 GB RAM) Since all three options appear to have similar costs in our subscription, could you please review our current workload and recommend which VM size would be most appropriate? Additionally, can you confirm whether you see any evidence of CPU credit exhaustion, resource throttling, memory pressure, or other indicators in Azure diagnostics that would support the conclusion that the B-series VM is the root cause of the RDP connectivity issue? We would like to select the most suitable VM size based on Microsoft's recommendation rather than simply increasing resources without understanding the underlying cause. Regards, Sanjay Yadav

    Was this answer helpful?


  3. Nikhil Duserla 9,950 Reputation points Microsoft External Staff Moderator
    2026-04-27T16:19:49.05+00:00

    Hello @Sanjay ,

    In addition to Ankit’s response, he has shared a link via private message if you click on it and create a ticket from your end. That enables backend log collection to help identify the root cause of the issue. Reviewing these logs will provide better visibility into the underlying behavior.

    Furthermore, I would like to highlight an important consideration regarding the VM configuration. The current workload is running on a B-series VM, which is designed for burstable performance and is generally not recommended for sustained production workloads. B-series VMs operate on a CPU credit model, where performance is throttled once credits are exhausted.

    This behavior does not necessarily require a consistently high workload intermittent spikes, background processes, or short bursts of activity can consume CPU credits. As a result, the issue may appear to occur “suddenly,” even without any deliberate change in workload patterns.

    Microsoft does not recommend using B-series VMs for production workloads. I recommend upgrading the VM to a higher tier to achieve more consistent and reliable performance.

    Azure B‑series VMs are explicitly designed for workloads that don’t need constant full CPU and benefit from burst capacity. For reference- https://docs.azure.cn/en-us/virtual-machines/sizes/general-purpose/b-family

    Unlike standard Azure VM types that provide fixed CPU performance, B-series VMs operate on a CPU credit model. These VMs accumulate CPU credits when operating below a defined baseline and consume credits when bursting above that baseline. Once the credits are exhausted, the VM is throttled back to its base performance level until sufficient credits are regained. This behavior can lead to performance variability under sustained workloads.

    For the better performance upgrade to another virtual machine size- [Virtual machine sizes overview - Azure Virtual Machines https://docs.azure.cn/en-us/virtual-machines/sizes/overview?tabs=breakdownseries%2Cgeneralsizelist%2Ccomputesizelist%2Cmemorysizelist%2Cstoragesizelist%2Cgpusizelist%2Chpcsizelist

    Was this answer helpful?

    0 comments No comments

  4. AI answer

    2026-04-23T11:55:43.2266667+00:00

    Use a structured approach to isolate whether the issue is platform, network, or guest-OS related, and add monitoring so the next occurrence can be correlated with logs and metrics.

    1. Check Azure-side health and connectivity
    2. In the Azure portal, open the VM and check Resource health under the Help section to confirm whether Azure reports any underlying host or platform issues at the times when RDP fails.
    3. Verify that the VM is in Running state (not in Failed/Deallocating) when RDP is unavailable.
    4. Use Boot diagnostics to review console logs and screenshots around the time of the issue for OS boot, crash, or hang indications.

    References:

    1. Validate network path and NIC state when the issue occurs
    2. When RDP is failing, use Network Watcher IP flow verify against the VM’s NIC to confirm that an inbound Allow rule exists and no NSG rule is blocking TCP 3389.
    3. Review effective NSG rules on the NIC and subnet to ensure there is a valid inbound rule for RDP and no higher-priority deny.
    4. If all NSG rules are correct but the VM is unreachable on all ports, check whether the NIC inside the guest is disabled:
      • Use Serial Console (CMD or PowerShell) to run:
             netsh interface show interface
        
      • If the primary interface is disabled, re-enable it:
             netsh interface set interface name="<interface name>" admin=enabled
        
      • Re-check with netsh interface show interface and then test RDP again.
    5. If Serial Console is unavailable, use the documented Reset network interface operation from the portal to move the NIC to a new IP in the subnet.

    This addresses intermittent loss caused by NIC being disabled or misconfigured inside the guest.

    References:

    1. Confirm RDP service and listener health

    When the VM is in the failed state (using Serial Console or remote tools):

    1. Verify that the RDP listener is enabled:
         reg query "HKLM\SYSTEM\CurrentControlSet\Control\Terminal Server\Winstations\RDP-Tcp" /v fLogonDisabled
      
      • If it returns 1, enable it:
         reg add "HKLM\SYSTEM\CurrentControlSet\Control\Terminal Server\Winstations\RDP-Tcp" /v fLogonDisabled /t REG_DWORD /d 0 /f
      
    2. Ensure Remote Desktop is enabled via the fDenyTSConnections registry value as described in the RDP reset article referenced from Source 4.
    3. If RDP still fails, follow the guidance to:
      • Ensure Remote Desktop Services is running.
      • Reset the RDP listener if it is missing or misconfigured.

    References:

    1. Investigate guest OS events and resource exhaustion

    Because a restart temporarily fixes the problem, intermittent issues can be caused by:

    • RDP service or LSASS hangs
    • Resource exhaustion (CPU, memory, handles, or disk space)
    • Profile or logon issues

    Actions:

    1. Use remote tools (Serial Console, or attach OS disk to a recovery VM) to access Event Viewer on the affected VM and correlate the time of RDP failures with:
      • System and Application logs for service crashes, TermService issues, or disk/IO errors.
      • Security logs for logon failures or lockouts. Guidance for using event IDs to troubleshoot RDP failures is provided in the event-based article.
    2. Check for low disk space on OS and data disks. If disks are nearly full, expand them and configure alerts; low disk can cause severe slowdowns and service instability.
    3. If logon is extremely slow or stuck at profile-related messages when RDP does work, consider user profile corruption or insufficient resources and test with a new local admin account.

    References:

    1. Hardening and permanent mitigation steps
    2. Implement the basic RDP troubleshooting steps as a baseline:
      • Use the Reset password blade in the portal with Mode = Reset configuration only to reapply RDP configuration and firewall rules if corruption is suspected.
      • Verify NSG rules and public IP configuration.
    3. Enable and retain Boot diagnostics and Serial Console for faster recovery and deeper diagnostics when the issue recurs.
    4. Configure Azure Monitor alerts on:
      • VM CPU, memory (via Guest metrics), and disk space.
      • Platform health events and VM reboot/crash events.
    5. If investigation shows no guest or network cause and Resource health reports intermittent platform issues, open a support case with timestamps so Azure Support can check host/node-level logs.

    References:


    References:

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.