Persistent nonpaged pool leak (tag RTLF, mpsdrv.sys) growing to several GB — root allocator not found after Driver Verifier tracing on mpsdrv.sys, netio.sys, and tcpip.sys

Deepak B 10 Reputation points
2026-07-30T19:09:44.6533333+00:00

System: HP Victus laptop, Windows 11 (build 26200.8875)

Since a recent period, I've observed nonpaged pool usage growing from 0 (fresh boot) to 3-4GB over several days of normal use, eventually impacting system performance. Using PoolmonX, I identified the "RTLF" pool tag (documented as "MPSDRV filter", associated with mpsdrv.sys, the Windows Defender Firewall driver) as responsible for the vast majority of this growth — consistently 80-90%+ of total nonpaged pool consumed, with a poor free/alloc ratio (e.g., 47,551 allocs vs only 5,905 frees in one sample).

What I've already investigated and ruled out:

1. Third-party AV/firewall software: fully uninstalled (MCPR used to remove McAfee remnants), driverquery confirms no leftover drivers.

2. VPN clients: no active VPN in use. However, I did find and remove a leftover WireGuard kernel driver (wireguard.sys, service + file), likely a remnant from a NordVPN installation years ago. Removing it did not eliminate the leak.

3. HP Omen Gaming Hub's "Network Booster" feature was found to be pinging 8.8.8.8 repeatedly, which correlated with a large spike in aleConnectedEndpoints entries (via netsh wfp show state) attributed to the "System" process. Disabling this feature reduced the growth rate significantly (roughly 4-5x slower), but did not stop the leak entirely — growth continues at a slower, usage-correlated rate (heavier browsing/network use = faster growth).

4. Comparing netsh wfp show filters dumps over time: total filter count remains stable (~786 filters), ruling out accumulating filter/policy objects as the cause.

5. Comparing netsh wfp show state dumps: aleConnectedEndpoints count does fluctuate with usage but does not show a 1:1 correlation with the scale of pool growth (e.g., pool grew ~1GB in one heavy-use day while the endpoint snapshot count was actually lower than a prior, lower-pool-usage snapshot) — suggesting the leak relates to allocate/free churn (connection setup/teardown) rather than currently-open connections.

Driver Verifier findings (Pool Tracking only, via LiveKd to avoid full boot debug mode where possible):

  • mpsdrv.sys: Verifier-tracked allocations show only 1-2 small allocations (tag FLTR, ~32 bytes, via mpsdrv!CreateFilterList), nowhere near the RTLF volume.
  • netio.sys: tracked allocations use other tags entirely (WfpA, WfpH, WfpL), no RTLF entries found.
  • tcpip.sys: tracked allocations (1,792 entries sampled) show tags like TcTW, TcBW, Fwpp, AlsT, AleS, etc. (many via tcpip!WfpPoolAllocNonPaged), but zero RTLF entries.

So despite RTLF being labeled "MPSDRV filter, Binary: mpsdrv.sys" in the pooltag reference, Driver Verifier's per-module pool tracking on mpsdrv.sys, netio.sys, and tcpip.sys does not show any allocations under this tag from those modules. I suspect the actual allocation may occur in fwpkclnt.sys or another statically-linked/shared WFP component not directly selectable in Driver Verifier, or that the tag is being passed through in a way that isn't attributed to the calling module's own image range.

Questions:

  1. Is this a known issue with mpsdrv.sys / the WFP connection-tracking layer in recent Windows 11 builds?
  2. Is there a way to get Driver Verifier (or another tool) to properly attribute RTLF-tagged allocations to their true caller, given it's not showing up under the three modules I've tried?
  3. Is there a supported way to reduce mpsdrv's per-connection state tracking overhead (e.g., disabling specific protocol analyzers like FTP/PPTP inspection) that might reduce this leak, assuming it's tied to general ALE connection lifecycle handling rather than a specific protocol?

Happy to provide additional !poolused/!verifier output, netsh wfp show state/filters XML dumps, or a kernel dump if useful.

Windows for business | Windows Client for IT Pros | Performance | System performance
0 comments No comments

2 answers

Sort by: Most helpful
  1. Akane 0 Reputation points
    2026-10-11T02:25:59.3333333+00:00

    Hi Ivy, and thank you Deepak for the original report.

    I can add a second, independent machine reproducing the same RTLF leak, and — more importantly — I have isolated the exact trigger condition, which I believe is the missing piece here.

    System: HP OMEN 17-ck1xxx, Intel i7-12700H, 32 GB RAM, Windows 11 26H2 Build 26300.9457 (newer than the 26200 build in the original post). Secure Boot enabled, VBS running.

    1. The leak is triggered by Modern Standby (S0), not by network usage

    I sampled RTLF every 30 seconds for 90 minutes and cross-referenced every value against Kernel-Power events 506/507 (Modern Standby entry/exit). The correlation is exact:

    | Interval | Duration | State | Δ RTLF | Rate |

    |---|---|---|---|---|

    | 08:52:59 → 09:13:05 | 20.1 min | Awake | 0.00 MB | 0.00 MB/min |

    | 23:52:49 → 00:43:56 | 51 min | Modern Standby | +474.93 MB | 9.31 MB/min |

    | 00:45:48 → 03:24:01 | 158 min | Modern Standby | +1,108.15 MB | 6.92 MB/min |

    | 03:25:13 → 08:12:05 | 287 min | Modern Standby | +1,678.06 MB | 5.82 MB/min |

    The machine was not rebooted during the overnight window (uptime 11h14m), so these are three separate standby windows on one boot. Every RTLF increment maps 1:1 onto a Modern Standby window; every awake window shows exactly zero growth, including a 20-minute awake window in which Chrome was actively transferring at an average 539 KB/s with peaks over 5,000 KB/s.

    This means the trigger is not general network traffic or connection churn. It is specifically the standby transition. I can reproduce on demand with psshutdown -x -t 0, and a single transition reliably adds tens to hundreds of MB.

    2. Peak pool tag data (equivalent to the PoolMon output you requested)

    NtQuerySystemInformation(SystemPoolTagInformation), class 22:

    
    Tag     NonPaged(MB)   Allocs     Frees      Outstanding
    
    RTLF       10,789.47   189,678     1,417        188,261    <- 94% of nonpaged pool
    
    HalD          202.23     1,827     1,583            244
    
    EtwB          198.29     6,193     3,092          3,101
    
    ConT          127.88     4,443     4,239            204
    
    NVRM           55.01 11,431,707 11,407,795        23,912
    
    Total nonpaged pool: 11.50 GB
    
    

    Supporting counters at peak: Pool Nonpaged Bytes 12,527 MB, Committed Bytes 25,007 MB, Available MBytes 0 MB, Page Faults/sec 121,092, Pages Input/sec 0.00 and Pages Output/sec 0.00 — the memory is unpagedable and pinned, not paging activity.

    Critically, the sum of all user-mode process private commits was 8.34 GB, largest single process 0.86 GB. So the leaked memory is not held by any user-mode process.

    3. On the tag attribution problem

    Two findings that may help narrow where RTLF is actually issued from:

    • I exhaustively scanned every .sys and .dll under System32, DriverStore and Program Files — the literal string RTLF does not exist in any binary on disk. This is consistent with Deepak's finding that Verifier cannot attribute it to mpsdrv.sys / netio.sys / tcpip.sys.
    • RTLF is documented as MPSDRV filter, but the absence of the literal in the binary suggests either runtime construction of the tag, or that the allocating component is not a separately-loadable image.

    4. Why customers cannot self-diagnose this (a product gap worth addressing)

    On a locked-down but otherwise fully supported configuration, every path to attributing the allocation is blocked:

    • bcdedit /debug on → "This value is protected by Secure Boot policy and cannot be modified or deleted" (as Administrator)
    • kd.exe -kl → "Local kernel debugging is disabled by default"
    • LiveKD → blocked by VBS

    The only workaround is entering firmware setup and disabling Secure Boot, which weakens the boot chain and may require BitLocker recovery. That is not a reasonable ask for a customer reporting a bug. Question: what is the supported method to attribute a pool tag to a call stack on a Secure Boot + VBS system?

    5. Data I can provide

    I have already captured a ~821 MB WPR Pool-profile ETW trace (wpr -start Pool -filemode) on the affected machine. Note it was captured while awake (so it will not contain the leaking allocations). I am happy to capture a fresh trace spanning a Modern Standby transition, since the trigger is now known and reproducible on demand — please confirm and I will upload it.

    I can also provide the full 30-second-resolution time series (RTLF MB, outstanding allocations, nonpaged pool MB, available MB), and the complete Kernel-Power event log cross-referenced with pool measurements.

    Adding a fourth question to Deepak's list: given that the allocations cluster at standby entry/exit rather than accruing continuously during standby, is there a specific suspend/resume code path in the WFP connection-tracking layer that this points to?

    Thank you — happy to run any additional instrumentation you specify.

    Was this answer helpful?

    0 comments No comments

  2. Ivy Bui (WICLOUD CORPORATION) 515 Reputation points Microsoft External Staff
    2026-07-31T03:53:55.3466667+00:00

    Hello Deepak B,

    Thank you for the detailed information and troubleshooting results you have shared.

    Based on the findings so far, the RTLF pool tag is possibly responsible for a significant portion of the nonpaged pool growth. However, the current data is not sufficient to conclusively identify the component that is actually performing the allocations, as Driver Verifier did not attribute the RTLF allocations to mpsdrv.sys, netio.sys, or tcpip.sys during the tracing performed.

    To further investigate the allocation path behind these allocations, we would appreciate your help collecting additional diagnostic data while the nonpaged pool usage is elevated:

    • PoolMon output showing the RTLF tag growth
    • Output from: verifier /query verifier /querysettings
    • A WPR pool trace collected during reproduction: wpr -start -pool -filemode After allowing the issue to reproduce: wpr -stop C:\Temp\pool.etl
    • If feasible, a kernel memory dump captured when nonpaged pool usage is high

    At this time, we are not aware of a documented Windows 11 issue specifically associated with RTLF or mpsdrv.sys on the build you are running. We also do not have a supported method to selectively disable WFP connection tracking components for troubleshooting purposes.

    Once the requested data is available, we will review it further to determine whether we can identify the allocation source more precisely.

    Thank you for your cooperation, and please let us know if you have any questions.

    Ivy Bui


    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.