Azure Firewall - High SNAT Port Utilization

$@chin 380 Reputation points
2026-10-05T18:43:37.17+00:00

Observing high SNAT port utilization on Azure Firewall and would like to understand whether Azure Firewall’s built-in auto-scaling can help resolve the issue.

As I understand it, Azure Firewall has a minimum of two instances, and a single instance can provide up to 2,496 SNAT ports per public IP. Azure Firewall can also automatically scale out additional instances based on traffic load.

I have the following questions:

1. Azure Firewall Auto-Scaling and SNAT Capacity

  • If the firewall is experiencing high SNAT port utilization, will the built-in auto-scaling mechanism automatically add firewall instances based on the increased traffic?

Does scaling out additional firewall instances increase the overall available SNAT port capacity proportionally?

Is there any delay between detecting high SNAT utilization and scaling out that could potentially result in SNAT port exhaustion before the additional instances become available?

If the firewall scales from two instances to four instances, can we expect the available SNAT capacity to increase accordingly, assuming sufficient public IPs and traffic distribution?

2. Identifying the Traffic Causing High SNAT Utilization

More importantly, I would like to understand how we can identify which traffic or workloads are consuming the highest number of SNAT ports.

Is there a recommended way to perform a detailed analysis of SNAT port consumption on Azure Firewall?

For example:

How can we identify the source private IPs/VMs generating the highest number of SNAT connections?

Can we identify the destination IPs/FQDNs and ports associated with the high SNAT consumption?

  • Is there a way to determine which application/workload is responsible for creating a large number of outbound connections?

Is there a recommended KQL query that can show the top source IPs by number of outbound connections/SNAT usage?

Can we determine whether the issue is caused by a high number of concurrent connections, short-lived connections, or a particular destination/service?

3. What Should Be Checked During a SNAT Spike?

What would be the recommended troubleshooting methodology and key parameters that should be checked when high SNAT utilization is observed?

Should the analysis primarily focus on the exact time period when SNAT utilization peaked, rather than looking only at the total/aggregate values for the entire day?

For example, if SNAT utilization reached 90–100% at a particular time, should we analyze the traffic and firewall metrics specifically around that period?

Also, when reviewing the SNAT port utilization metric in Azure Monitor, which aggregation should be used for troubleshooting - Maximum (Max) or Average (Avg) ?

Should Max be used to identify short-duration SNAT exhaustion/spikes?

Should Average be used to understand sustained SNAT utilization and overall capacity trends?

Is there a recommended time granularity, such as 1 minute, 5 minutes, or 15 minutes, for this analysis?

Which metrics should be correlated during the spike, such as:

  • SNAT port utilization

Also, should we look at the peak/max SNAT utilization at a specific point in time, or is there value in analyzing the overall utilization trend across the day/week to determine whether this is a recurring capacity issue?

I would also like to understand how to differentiate between:

A temporary traffic spike

A specific VM/workload generating excessive outbound connections

A specific destination/service consuming a large number of SNAT ports

A large number of short-lived connections

Long-lived connections consuming SNAT ports

  • Insufficient SNAT/public IP capacity Recommended Remediation

Once the source of the SNAT exhaustion is identified, what would be the recommended remediation?

For example, should we consider:

Allowing Azure Firewall to scale out automatically

Adding additional Azure Firewall public IP addresses

Reviewing the workload/application connection behavior

Using a different outbound connectivity architecture for specific workloads

Any other Azure Firewall SNAT best practices

I would particularly appreciate guidance on the best method to troubleshoot and identify the top SNAT consumers before making any configuration changes.

Azure Firewall
Azure Firewall

An Azure network security service that is used to protect Azure Virtual Network resources.

0 comments No comments

1 answer

Sort by: Most helpful
  1. SHOUMIK CHAKRAVARTY 1,150 Reputation points
    2026-10-07T08:01:04.3233333+00:00

    Autoscale won't help here, and that's the main thing worth knowing. Scale-out doesn't key on SNAT at all. From Azure Firewall performance: "Azure Firewall gradually scales out when the average throughput and CPU consumption reach 60% or if the number of connections usage reaches 80%. Scale out takes five to seven minutes."

    SNAT utilization isn't in that list. The monitoring reference says: "when the firewall scales out for different reasons (for example, CPU or throughput) more SNAT ports also become available." Extra ports are a side effect of scaling, not a response to running out of them.

    So to your four questions: no, high SNAT utilization alone won't trigger a scale-out. Yes, instances multiply capacity, since the allocation is per instance: "Azure Firewall provides 2,496 SNAT ports per public IP address configured per backend virtual machine scale set instance (Minimum of two instances), and you can associate up to 250 public IP addresses." And yes there's a lag of five to seven minutes, but only once something other than SNAT crosses its threshold.

    Worth knowing what exhaustion actually does, because it isn't a hard cliff: "If SNAT ports are used more than 95%, they're considered exhausted and the health is 50% with status=Degraded and reason=SNAT port. The firewall keeps processing traffic and existing connections aren't affected. However, new connections might not be established intermittently."

    Both documented fixes are deliberate rather than automatic. Add public IP addresses, each worth 2,496 ports per instance, or put a NAT gateway in front: "Use a NAT gateway when you need dynamic SNAT port allocation across the subnet. It provides up to 64,512 SNAT ports per public IP address." The NAT gateway also allocates dynamically across the subnet rather than fixing ports per instance, which suits spiky traffic better.

    For finding what's consuming the ports, enable the Top Flows log. It "shows the top connections that are contributing to the highest throughput through the firewall", it's off by default and turned on through PowerShell, and it lands in the AZFWFatFlow table with source IP, destination and port. That gets you the "which VM, which destination" answer directly.

    Help make this community better for everyone: if this answer resolved your issue, please accept it or leave an upvote. If not, share more details in a comment so we can continue the discussion and find the right solution.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.