Hello @Melonia-H ,
Thank you for the detailed write-up. The kernel-side checks made it much easier to narrow down. I tried to reproduce the issue and went through the related documentation. Below is what I found.
1. What I tested
I ran your minimal repro, plus two variants closer to your setup, on:
- Windows Server 2022 Datacenter (Azure Edition), build 20348.5622, on an Azure VM with 4 vCPU / 16 GB, memory compression off (the Server default)
- Windows 11 Enterprise 26200, for comparison
The scenarios:
- Your minimal repro: commit and touch 500 MB, call
SetProcessWorkingSetSizeEx(GetCurrentProcess(), 1 MB, 50 MB, QUOTA_LIMITS_HARDWS_MAX_ENABLE), then keep touching every page in a loop. - Your step-down pattern: hard max 380 MB > 89 MB > 61 MB while the process keeps touching 500 MB.
- A cross-process variant of your MPI layout: a controller starts 10 workers, each touching 400 MB continuously. It applies
QUOTA_LIMITS_HARDWS_MAX_ENABLEto each worker's handle and steps the max down 380 > 89 > 61 MB, sampling every 500 ms.
I read the limits and flags back with GetProcessWorkingSetSizeEx and measured the working set with GetProcessMemoryInfo.
2. Results (identical on Server 2022 and Windows 11)
- The trim happened synchronously inside the call. The working set went from 502 MB to 49 MB before
SetProcessWorkingSetSizeExreturned. - It stayed at or below the maximum while the process kept touching all 500 MB.
- On each sharp step-down it went straight to the new limit (379 > 88 > 60 MB). I did not see the lag you describe.
- In the 10-worker cross-process test, 0 of 10 workers were over the limit at any sample point.
So on a clean 20348.5622 system, the hard maximum behaved as a real cap for ordinary private, pageable memory. It did not behave like a trim target evaluated lazily on the fault path.
3. What the documentation says
- SetProcessWorkingSetSizeEx says that with
QUOTA_LIMITS_HARDWS_MAX_ENABLE, "The working set will not exceed the maximum working set limit." The Remarks say the HARDWS flags "enable you to ensure that limits are enforced." Your reading of the API is correct. The documentation does not describe when a lowered maximum is applied, does not mention any registry setting or policy that affects it, and does not describe any difference between Windows versions. - Working Set states that the working set "contains only pageable memory allocations". AWE and large-page allocations are not included, so they cannot explain a working set above the maximum.
- VirtualLock limits the pages a process can lock to roughly its minimum working set. With a 1 MB minimum, user-mode
VirtualLockon its own is unlikely to cause an overrun of tens of MB. - JOBOBJECT_BASIC_LIMIT_INFORMATION notes that when
JOB_OBJECT_LIMIT_WORKINGSETis in effect, "you cannot use SetProcessWorkingSetSize to change the minimum or maximum working set size of a process in a job object". For nested jobs, the smallest limit in the chain applies. - The thread you linked (1003229) reports the same symptom on Server 2022 only, on different hardware from the 2019 machine. It has no root cause or official answer.
4. Answers to your questions
- Is it still a hard limit on Server 2022? According to the documentation, yes. In my test on a clean Server 2022 (20348.5622), it was enforced as a hard cap and applied immediately.
- Has it changed compared with Server 2019? I couldn't compare against 2019 in this test. What I can say is that the lazy, fault-path-only behaviour you see is not the default behaviour of the 20348 build I tested. Because this symptom has been reported on Server 2022 before (1003229), I'm not ruling out something specific to a build, hardware or workload.
- Is there a way to force a synchronous trim? In my tests, calling
SetProcessWorkingSetSizeExwith the lower maximum already trimmed synchronously.(SIZE_T)-1, (SIZE_T)-1or EmptyWorkingSet were not needed. - Is there a registry setting, policy or update that affects this? I found no documented one.
5. Why it may behave differently on your system
Your minimal repro does not reproduce on a clean box, so the overrun most likely depends on what your worker processes' working set contains or how they are hosted. The candidates I would check first:
- Pages the memory manager cannot trim: The most likely source in an MPI workload is memory registered for RDMA/NetworkDirect, which drivers lock in physical memory and often keep cached for reuse (see Caching Registered Memory). Unlike
VirtualLock, driver locks are not limited by the minimum working set. If such pages are counted in the working set, that would also fit the trim counters staying at 0, since there would be nothing trimmable to remove. I haven't confirmed this; it is a hypothesis. - Shared pages, such as the shared-memory sections used by the MPI transport within a node.
- Job objects: whether the MPI launcher (mpiexec/smpd, HPC Pack or a scheduler) places the workers in a job, possibly nested, with its own working-set limit.
- Another component (the solver, scheduler or an agent) changing the working-set limits again after you set them.
6. What would help next
- The exact build including UBR (
winver, orUBRunderHKLM\SOFTWARE\Microsoft\Windows NT\CurrentVersion). - Whether your minimal repro alone (no MPI job running) still exceeds the limit on that server.
- Which MPI implementation and transport you use (MS-MPI or Intel MPI; shared memory, TCP or NetworkDirect/RDMA).
- Whether the workers run inside a job object, and its limits if so.
- A breakdown of a worker's working set while it is over the limit:
- Sysinternals VMMap, looking at the Locked WS, Shareable WS and Private WS columns, or
-
QueryWorkingSetEx, counting theLocked,SharedandShareCountbits inPSAPI_WORKING_SET_EX_BLOCK.
If the standalone minimal repro does exceed the limit on your build, that points to a build- or platform-specific issue. In that case I'd recommend opening a support case with Microsoft, since a definitive "bug vs. by design" answer needs the Windows team. Please include the repro, the exact build/UBR, the hardware/hypervisor details, and a kernel dump or trace captured during the overrun.
I hope this helps narrow it down. I'm happy to look at the VMMap or QueryWorkingSetEx output once you have it. If you found my response helpful or informative, I would greatly appreciate it if you could follow the instruction here so others experiencing similar behavior can benefit from it as well.
Thank you.