Self-Hosted IR Copy activities hitting OutOfMemoryException and indefinite queuing since Synapse-to-Fabric migration doubled concurrent load

Mohd Anas 20 Reputation points
2026-09-15T15:41:25.6566667+00:00

Here's the paragraph draft for the Q&A portal:


We're mid-migration from Synapse to Microsoft Fabric and are running the same SFTP-based Copy Data pipelines on both platforms concurrently, pulling from a local E:\ drive through a single Self-Hosted Integration Runtime (HA disabled, 16 concurrent job limit). Previously, on Synapse alone, our major run executed roughly 350 activities against E:\ without issue. Since bringing Fabric online with the same pipeline set now also pointed at doubled source directories, we've started seeing two problems: first, child activities failing on the source side with ErrorCode=SystemErrorOutOfMemory, Type=Microsoft.DataTransfer.Common.Shared.HybridDeliveryException, Message=A task failed with out of memory., Source=Microsoft.DataTransfer.TransferTask, ultimately a System.OutOfMemoryException from Microsoft.DataTransfer.ClientLibrary; and second, some pipeline runs on Synapse entering a queued state and never progressing toward completion, even when the SHIR node's concurrent job count is below its configured limit. Since the workload and IR configuration didn't change — only the addition of Fabric running the same pipelines in parallel with roughly double the directory/file volume — we suspect the two platforms are now contending for the same SHIR node's memory and connection capacity.

We'd like guidance on: (1) whether Synapse and Fabric pipelines sharing one Self-Hosted IR node is a known contention point that requires separate IR instances per platform during migration, (2) what's actually driving the OutOfMemoryException on the source side — node memory, SFTP client buffering, or something scaling with concurrent activity count — and how to size or tune around it, and (3) what could cause activities to sit queued indefinitely despite available concurrent-job headroom on the node.

Azure Synapse Analytics
Azure Synapse Analytics

An Azure analytics service that brings together data integration, enterprise data warehousing, and big data analytics. Previously known as Azure SQL Data Warehouse.

0 comments No comments

1 answer

Sort by: Most helpful
  1. Aditya Singh Rathore 195 Reputation points
    2026-09-21T11:21:24.54+00:00

    Hi @Mohd Anas ,

    Thanks for the detail, it makes this much easier to reason about. One thing I'd check first: as far as I can tell, Fabric pipelines can't use a self-hosted IR at all. SHIR is for ADF and Synapse, and Fabric pipelines use the on-premises data gateway instead (comparison here). So the two platforms probably aren't sharing one SHIR node. What they may be sharing is the machine, for example a gateway installed on the same server as the SHIR, or both reading from the box that hosts E:\ and the SFTP service. How is Fabric reaching E:\ in your setup? That would change my answer a bit.

    On your three questions, assuming a shared host:

    1. Separate instances: I'd separate them. The SHIR docs recommend a dedicated machine, say not to put it on the same machine as a Power BI gateway, and suggest keeping it off the machine that hosts the data source so they don't fight over resources.
    2. The OOM: That error comes from the IR node running low on memory while it's executing your activities. The usual advice is to reduce how many run on that node at once, or scale the node up or out. I couldn't find a published formula for per-activity memory, so I can't tell you whether SFTP buffering or activity count matters more. Doubling your file volume could well make each activity's file listing heavier, though. The copy monitoring view shows a "Listing source" stage you can compare between failing and passing runs. Since you're memory-bound rather than slot-bound, I'd try lowering the 16-job limit before raising it, and add RAM or a second node.
    3. Queued with headroom: The node only starts a job when it polls the queue, so a memory-starved node or a worker that recycled after an OOM might leave jobs waiting while the counter looks fine. That's a guess on my part. I'd also check the pipeline-level concurrency setting, the IR metrics, and the logs under C:\ProgramData\Microsoft\Integration Runtime\Logs. If it keeps happening after you split the hosts, a support ticket with the run IDs would be worth it.

    Hope that helps. Let us know what you find on the Fabric side.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.