Azure Content Understanding intermittently fails with InternalServerError when using prebuilt-documentSearch via REST API

Sofia Isabel Tátá 0 Reputation points
2026-06-30T11:46:24.7233333+00:00

Hello Microsoft Community,

We are experiencing intermittent failures when using Azure Content Understanding to analyze documents through the REST API. The issue is occurring in the Sweden Central Azure region and was first noticed last week on 26/06/2026.

The service works correctly for smaller and simpler documents, for example documents with around 3 pages. However, for larger documents, the analysis often fails with an InternalServerError and sometimes succeeds after multiple retries.

Examples of the behavior observed:

  • A 30-page document only completed successfully on the third attempt, using retry logic after failures.
  • A 131-page document was split into batches of 20 pages, significantly below the documented API limit of 300 pages, but processing of the complete document (all the batches) only succeeded on the seventh attempt.
  • The failures do not appear to be related to authentication or endpoint configuration, since some of the same requests succeed after retries.
  • The issue also does not appear to be related to the API version, as we are using the GA version.

We are using the analyzer prebuilt-documentSearch. One additional point we noticed is that this analyzer does not appear in the Content Understanding playground in the new Foundry experience, although it is mentioned in Microsoft documentation listing the available Content Understanding analyzers.

The error returned by the analysis operation is the following:

{'id': '26ca7acf-3f55-4212-bfdd-0ec01dfe64ce', 'status': 'Failed', 'error': {'code': 'InternalServerError', 'message': 'An unexpected error occurred.'}, 'result': {'analyzerId': 'prebuilt-documentSearch', 'apiVersion': '2025-11-01', 'createdAt': '2026-06-29T17:33:54Z', 'warnings': [], 'contents': []}}

Could you please help us understand the following?

  1. Is there any known issue or instability affecting Azure Content Understanding in the Sweden Central region, particularly since last week or around 26/06/2026?
  2. Are there any known limitations, throttling conditions, or service-side constraints that could cause intermittent InternalServerError responses when processing larger documents, even when the requests are below the documented page limits?
  3. Is the prebuilt-documentSearch analyzer fully supported in the GA API version 2025-11-01?
  4. Is there a reason why prebuilt-documentSearch is documented but not visible in the Content Understanding playground in the new Foundry experience?
  5. Are there any recommended retry, batching, timeout, or request configuration practices for this analyzer when processing larger documents through the REST API?

Any guidance on how to troubleshoot this issue, or confirmation of whether this is a known service-side problem would be greatly appreciated.

Thank you in advance.

Azure Content Understanding in Foundry Tools
0 comments No comments

1 answer

Sort by: Newest
  1. Jerald Felix 18,760 Reputation points Volunteer Moderator
    2026-07-05T16:17:48.42+00:00

    Hello Sofia Isabel Tátá,

    Greetings! Thanks for raising this question in the Q&A forum.

    The root cause of intermittent InternalServerError responses that clear up on retry, especially correlating with larger page counts rather than any fixed limit, is almost always backend capacity or transient processing instability on the Content Understanding service side rather than anything wrong with your request payload or authentication. This pattern, where identical requests succeed after several retries with no change on your end, is a classic signature of a service-side throttling or capacity issue rather than a client-side configuration problem. I checked the public Azure Status page and did not find a currently listed incident specifically tied to Content Understanding in Sweden Central around 26/06/2026, but the public status page frequently does not reflect narrower, service-specific degradations, so this does not rule out a backend issue on Microsoft's side.

    Here is how to proceed on each of your questions:

    Known regional issue: There is no publicly listed incident for Content Understanding in Sweden Central on the Azure Status page for that period. This does not mean nothing happened, since narrower service-level degradations are often only visible through Azure Resource Health on your specific resource, not the public status page. Check Resource Health on your Content Understanding resource in the Azure portal for that date range, and if it shows a platform-level event, that confirms it was not isolated to your account.

    Throttling or service-side constraints on larger documents: Content Understanding in Foundry Tools reached GA relatively recently, at API version 2025-11-01, and larger or more complex documents place more load on the backend OCR, layout, and generative pipeline stages that prebuilt-documentSearch chains together. Even when you are well below the documented 300-page limit, elevated processing time per page increases the odds of hitting a transient capacity ceiling on the service side, which surfaces as InternalServerError rather than a specific throttling error code. This is consistent with what other customers have reported with Content Understanding analyzers on larger or more complex inputs.

    Is prebuilt-documentSearch fully supported in GA 2025-11-01: Yes. prebuilt-documentSearch is documented as a supported RAG-oriented prebuilt analyzer as of the GA release, so this is not an unsupported or preview-only analyzer being called incorrectly.

    Why it does not appear in the Foundry playground: The new Foundry portal introduced a prebuilt analyzers playground alongside the GA release, but playground UI rollout for individual analyzers does not always happen in lockstep with API-level GA support. This looks like a rollout gap in the playground UI rather than an indication that the analyzer itself is unsupported, since your REST API calls to it are otherwise being accepted and processed.

    Recommended retry and batching practices, until this stabilizes:

    Keep your current strategy of splitting large documents into smaller batches, since this has already reduced (though not eliminated) the failure rate for you.

      Implement exponential backoff with jitter on `InternalServerError` specifically, since your evidence shows retries do eventually succeed:
      
      ```python
      import time
    

    import random

    max_retries = 8 for attempt in range(max_retries): result = call_content_understanding_analyze(document_batch) if result.get("status") == "Failed" and result["error"]["code"] == "InternalServerError": wait = min(60, (2 ** attempt) + random.uniform(0, 1)) time.sleep(wait) continue break ```

         Log the `apim-request-id` response header for every failed call. This is required if you escalate to Support, since it lets the Content Understanding backend team trace the exact failed request rather than working from timestamps alone.
         
            Avoid re-submitting the same batch immediately back to back, since bursty retries can compound backend load during a capacity event rather than helping.
            
    

    Given that this reproduces consistently on larger documents, clears on retry, and does not correlate with anything in your request configuration, I recommend opening an Azure Support case and including the apim-request-id values you have already captured, the exact timestamps of the failures, and the region. This gets it in front of the Content Understanding engineering team who can check backend telemetry for Sweden Central around your reported dates, which is not something visible from the client side.

    If this answer helps you kindly accept the answer which will help others who have similar questions.

    Best Regards,

    Jerald Felix.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.