An Azure networking service that is used to provision private networks and optionally to connect to on-premises datacenters.
The deterministic source-port behavior is useful evidence, but I don't think the available information is sufficient to attribute this specifically to a stale Azure SDN/host/ARM frontend flow.
Because the VM has an instance-level public IP, the outbound traffic uses that public IP, and this configuration is implemented as stateless 1:1 NAT. The VM therefore isn't subject to the Azure Load Balancer SNAT-port allocation behavior that normally causes SNAT exhaustion.
Run Network Watcher Connection Troubleshoot against the affected ARM destination on TCP/443, specifying both:
- a source port that consistently fails
- a source port that consistently succeeds
Microsoft's Connection Troubleshoot supports specifying a source port and can identify conditions including NSG/route problems and SourcePortInUse.
Since you've already captured SYNs leaving the VM, preserve simultaneous packet captures for a known-good and known-bad source port, along with:
- source private/public IP
- resolved ARM destination IP
- source and destination ports
- UTC timestamps
- Connection Troubleshoot results
- effective NSG and route information
If the same destination consistently succeeds or fails solely according to the source port, while Network Watcher shows no customer-controlled NSG/route issue, that is strong evidence for escalation.
At that point, open an Azure support case and provide the successful and failed flow tuples plus their UTC timestamps so Microsoft can correlate the traffic with platform-side telemetry.
References:
Connection Troubleshoot overview
Troubleshoot outbound connections with Connection Troubleshoot
Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.