Hello Rodrigo R,
Welcome to the Microsoft Q&A and thank you for posting your questions here.
I understand that you are having random loss of connectivity after NVA failover.
For more explanation and comprehensive diagnosis:
The DNS and Azure Firewall TCP explanations do not apply to YOUR scenario. The behavior strongly indicates that ICMP session state is not being preserved correctly during the FortiGate HA transition. FortiGate maintains state for ICMP traffic, but ICMP and UDP sessions require connectionless session synchronization for continuity across HA members. - https://docs.fortinet.com/document/%20fortigate/7.6.5/administration-guide/955521/session-pickup, and https://community.fortinet.com/fortigate-3/technical-tip-ha-session-failover-session-pickup-93706
What you can do is to:
- Verify that both FortiGate nodes run the same supported FortiOS build.
- Enable HA session pickup and connectionless session pickup.
- Confirm that the affected ICMP sessions are synchronized to the secondary node.
- Clear only the affected stale ICMP sessions after applying the correction.
- Perform one controlled failover while capturing traffic on both FortiGate interfaces and the Azure VM NIC.
- Escalate to Fortinet if the ICMP reply reaches the external FortiGate interface but is not forwarded internally.
- Escalate to Microsoft only if Azure packet evidence demonstrates loss outside the FortiGate guest operating system.
Configure HA session synchronization in the global FortiGate context:
config system ha
set session-pickup enable
set session-pickup-connectionless enable
end
Fortinet documents that session-pickup-connectionless synchronizes ICMP and UDP sessions so they can be maintained during failover. After enabling it, verify the cluster and session synchronization:
get system status
get system ha status
diagnose sys ha checksum cluster
diagnose system session list | grep synced
Access the secondary node and confirm the corresponding sessions have the synchronized-session flag:
execute ha manage 1
diagnose system session list | grep syn_ses
Fortinet documents the synced flag on the primary and syn_ses on the secondary as evidence that sessions were replicated between cluster members. If connectivity remains broken after the correction, clear only the affected ICMP session rather than disabling the entire firewall policy for 15 to 20 minutes:
diagnose sys session filter clear
diagnose sys session filter proto 1
diagnose sys session filter src <MONITORING-SERVER-IP>
diagnose sys session filter dst <AFFECTED-DESTINATION-IP>
diagnose sys session list
diagnose sys session clear
diagnose sys session filter clear
Review the output from diagnose sys session list before running the clear command. This prevents unrelated production sessions from being removed. During the next controlled failover, capture the affected flow on the FortiGate:
diagnose sniffer packet any \
"host <MONITORING-SERVER-IP> and host <AFFECTED-DESTINATION-IP> and icmp" \
6 0 l
Run Azure Network Watcher Packet Capture on both FortiGate VMs at the same time. Whatever evidence you found, interpret the evidence as follows:
- If the ICMP reply reaches the FortiGate external interface but does not leave the internal interface, the fault is in FortiGate session, NAT, offload, or HA synchronization. Open a Fortinet case.
- If the destination receives the request and returns a reply, but Azure does not deliver it to the FortiGate NIC, open a Microsoft Azure support case.
- If the destination never receives the request, verify the Public IP association and Azure Activity Log immediately after failover.
- If clearing only the filtered ICMP session restores connectivity, stale or unsynchronized FortiGate session state is confirmed.
- If enabling connectionless pickup makes repeated failovers successful, the configuration deficiency is resolved.
Do not upgrade solely on assumption. First identify the installed FortiOS build. If it is affected by a documented HA session synchronization defect, upgrade both nodes through Fortinet’s supported upgrade path to a corrected, vendor-recommended mature release. Fortinet documents HA synchronization issue 1064728 as corrected in 7.6.1 and scheduled for correction in 7.4.7. After enabling connectionless session pickup, removing the affected stale sessions, and confirming synchronization on both HA members, ICMP monitoring should continue through failover without waiting 15 to 20 minutes or disabling the production policy.
Use the following resources for the configuration and steps:
I hope this is helpful. Please! Do not hesitate to let me know if you have any other questions, steps or clarifications.
Please do not close the thread by upvoting and accepting the answer if any part of it is helpful.