Cognitive Services-Form Recognizer: Slow Processing of Invoice Batches — Prolonged End-to-End Times

Matias Haller 0 Reputation points
2026-09-05T14:08:11.2733333+00:00

Problem description

I am experiencing significant delays in processing invoice batches using Azure Cognitive Services-Form Recognizer. Since September 1, 2026, most batches, each containing 15 invoice documents, are taking over an hour to complete, with some exceeding 45 hours. The expected processing time was within minutes, but now the process is extremely slow.

Environment

Azure Cognitive Services-Form Recognizer resource in the US East, using the asynchronous "invoice" prebuilt model via batch submission, with default service tier S0

Current status

I am seeking assistance to identify the cause of the slowdown and to find possible solutions to improve processing times.

Azure Document Intelligence in Foundry Tools

1 answer

Sort by: Most helpful
  1. Manish Deshpande 8,215 Reputation points Microsoft External Staff Moderator
    2026-09-05T18:08:56.08+00:00

    Hello @Matias Haller ,

    Thanks for laying out the details so clearly the onset date, the batch size, the tier and the region gave me enough to go on without a lot of back and forth, and I appreciate that.

    Short version:
    I don't think this is something you've done wrong. There are no open Incidents to track this latency on the read and layout paths, and the prebuilt-invoice model runs on top of read and layout, so a workload like yours would inherit the slowdown.

    Here's what I'd like you to do, in priority order.

    1. Deal with the 24-hour retention risk first. This one is time-critical and completely independent of root cause. Your 45-hour batches are crossing a hard service boundary: batch operation results are retained for 24 hours after completion, and the batch operation status is no longer available 24 hours after batch processing completes. The input documents and result files do remain in your storage containers. Separately, the API allows you to access results up to 24 hours after sending the request. So: persist every batch operation ID and request ID to durable storage at submission time, and for any batch approaching the 20-hour mark, pull results straight from the output blob container rather than the status endpoint. Without this, a long batch can quietly become unpollable and a delay turns into rework.
    2. Run a cross-region control test. Stand up a second Document Intelligence resource in another region and submit the same 15-document batch to both East US and the new region simultaneously. Pick the comparison region deliberately we've seen similar latency behaviour in a few other regions recently, so validate your candidate with a small test batch before moving any production load. This isolates everything in one shot: your code, SDK, documents and network all drop out as variables. A large divergence in wall-clock completion time confirms it's region-scoped.
    3. Reproduce in Document Intelligence Studio, in East US, using Microsoft's own sample invoice. This takes your application out of the measurement entirely. In the closest prior occurrence of this pattern, the Studio samples were equally slow and that single data point is what got the case escalated to the product group. Record the end-to-end time; if it's well above baseline, we're conclusively service-side.
    4. Capture latency per page and the shape of the onset. Azure portal → your Document Intelligence resource → Monitoring → Metrics → add the Latency metric, adjust Aggregation, and set the window to cover late August onward. Normalise by total pages, not document count — each page is its own transaction. This matters because the published guidance gives you a threshold you can cite: "If you observe sustained periods (exceeding one hour) where latency per page consistently surpasses 15 seconds, consider addressing the issue." You're far past that. Export the chart for 25 August to present — a step change on 1 September, rather than a gradual ramp, is the signature of a service-side event.
    5. Rule out the storage leg. Batch analysis reads from and writes to blob containers, so storage latency sits inside your end-to-end number. Go to your storage resource → Monitoring → Insights and check both End-to-end (E2E) latency (from when Azure Storage receives the first packet of the request until it receives the client acknowledgment on the last packet of the response) and Server latency (from when Storage receives the last packet of the request until the first packet of the response is returned). This is the one segment fully under your control showing both flat across the onset date makes the escalation much harder to deflect.
    6. Sanity-check your polling. The recommendation is not to call the get analyze response more than once every 2 seconds for a corresponding POST, and the analyze response returns a retry-after header with the wait in seconds. If throttling does show up, the documented backoff is progressive — 2, 5, 13, 34 seconds between retries. Aggressive polling won't cause hour-long batches, but it muddies your telemetry and can hide 429s being swallowed inside an SDK poller. Worth capturing the distribution of status codes your client sees.

    Two things I'd save you the trouble of trying. A TPS quota increase almost certainly won't help — the S0 defaults are 15 analyze transactions/sec, 50 get operations/sec, 5 model-management and 10 list operations/sec, but every one of those surfaces as HTTP 429 (Too many requests), and you're seeing no errors at all. Raising TPS moves the throttling ceiling, not per-page processing speed. Switching to a custom model won't help either — same shared regional capacity, and if Step 3 shows our own Studio samples crawling, the model clearly isn't the variable. Both come up a lot in community threads on this symptom, and neither fits your case.

    One honest expectation to set: the documentation states that Foundry Tools don't provide a Service Level Agreement (SLA) for latency, so there's no latency SLA to claim against. The lever that actually works here is a support escalation joined to the live regional investigation — which is exactly what I want to do with your resource ID and request IDs. For what it's worth, in the earlier occurrences of this pattern I looked at, the condition resolved service-side without any customer action. But mitigations in the current East US series have been followed by recurrence, so I'm not going to promise you that outcome.

    Send me the resource ID and a few request IDs when you get a chance, and start with Step 1 today regardless — that's the one with a clock on it. I'll keep this case attached to the regional investigation and come back to you as soon as there's anything substantive, even if it's just "still open."

    References

    Thanks,
    Manish.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.