Hi Ivy, and thank you Deepak for the original report.
I can add a second, independent machine reproducing the same RTLF leak, and — more importantly — I have isolated the exact trigger condition, which I believe is the missing piece here.
System: HP OMEN 17-ck1xxx, Intel i7-12700H, 32 GB RAM, Windows 11 26H2 Build 26300.9457 (newer than the 26200 build in the original post). Secure Boot enabled, VBS running.
1. The leak is triggered by Modern Standby (S0), not by network usage
I sampled RTLF every 30 seconds for 90 minutes and cross-referenced every value against Kernel-Power events 506/507 (Modern Standby entry/exit). The correlation is exact:
| Interval | Duration | State | Δ RTLF | Rate |
|---|---|---|---|---|
| 08:52:59 → 09:13:05 | 20.1 min | Awake | 0.00 MB | 0.00 MB/min |
| 23:52:49 → 00:43:56 | 51 min | Modern Standby | +474.93 MB | 9.31 MB/min |
| 00:45:48 → 03:24:01 | 158 min | Modern Standby | +1,108.15 MB | 6.92 MB/min |
| 03:25:13 → 08:12:05 | 287 min | Modern Standby | +1,678.06 MB | 5.82 MB/min |
The machine was not rebooted during the overnight window (uptime 11h14m), so these are three separate standby windows on one boot. Every RTLF increment maps 1:1 onto a Modern Standby window; every awake window shows exactly zero growth, including a 20-minute awake window in which Chrome was actively transferring at an average 539 KB/s with peaks over 5,000 KB/s.
This means the trigger is not general network traffic or connection churn. It is specifically the standby transition. I can reproduce on demand with psshutdown -x -t 0, and a single transition reliably adds tens to hundreds of MB.
2. Peak pool tag data (equivalent to the PoolMon output you requested)
NtQuerySystemInformation(SystemPoolTagInformation), class 22:
Tag NonPaged(MB) Allocs Frees Outstanding
RTLF 10,789.47 189,678 1,417 188,261 <- 94% of nonpaged pool
HalD 202.23 1,827 1,583 244
EtwB 198.29 6,193 3,092 3,101
ConT 127.88 4,443 4,239 204
NVRM 55.01 11,431,707 11,407,795 23,912
Total nonpaged pool: 11.50 GB
Supporting counters at peak: Pool Nonpaged Bytes 12,527 MB, Committed Bytes 25,007 MB, Available MBytes 0 MB, Page Faults/sec 121,092, Pages Input/sec 0.00 and Pages Output/sec 0.00 — the memory is unpagedable and pinned, not paging activity.
Critically, the sum of all user-mode process private commits was 8.34 GB, largest single process 0.86 GB. So the leaked memory is not held by any user-mode process.
3. On the tag attribution problem
Two findings that may help narrow where RTLF is actually issued from:
- I exhaustively scanned every
.sysand.dllunderSystem32,DriverStoreandProgram Files— the literal stringRTLFdoes not exist in any binary on disk. This is consistent with Deepak's finding that Verifier cannot attribute it tompsdrv.sys/netio.sys/tcpip.sys. -
RTLFis documented asMPSDRV filter, but the absence of the literal in the binary suggests either runtime construction of the tag, or that the allocating component is not a separately-loadable image.
4. Why customers cannot self-diagnose this (a product gap worth addressing)
On a locked-down but otherwise fully supported configuration, every path to attributing the allocation is blocked:
-
bcdedit /debug on→ "This value is protected by Secure Boot policy and cannot be modified or deleted" (as Administrator) -
kd.exe -kl→ "Local kernel debugging is disabled by default" - LiveKD → blocked by VBS
The only workaround is entering firmware setup and disabling Secure Boot, which weakens the boot chain and may require BitLocker recovery. That is not a reasonable ask for a customer reporting a bug. Question: what is the supported method to attribute a pool tag to a call stack on a Secure Boot + VBS system?
5. Data I can provide
I have already captured a ~821 MB WPR Pool-profile ETW trace (wpr -start Pool -filemode) on the affected machine. Note it was captured while awake (so it will not contain the leaking allocations). I am happy to capture a fresh trace spanning a Modern Standby transition, since the trigger is now known and reproducible on demand — please confirm and I will upload it.
I can also provide the full 30-second-resolution time series (RTLF MB, outstanding allocations, nonpaged pool MB, available MB), and the complete Kernel-Power event log cross-referenced with pool measurements.
Adding a fourth question to Deepak's list: given that the allocations cluster at standby entry/exit rather than accruing continuously during standby, is there a specific suspend/resume code path in the WFP connection-tracking layer that this points to?
Thank you — happy to run any additional instrumentation you specify.