An AI tool in Foundry for analyzing documents and media to classify content, extract entities, and generate structured understanding
Hello Sofia Isabel Tátá,
Greetings! Thanks for raising this question in the Q&A forum.
The root cause of intermittent InternalServerError responses that clear up on retry, especially correlating with larger page counts rather than any fixed limit, is almost always backend capacity or transient processing instability on the Content Understanding service side rather than anything wrong with your request payload or authentication. This pattern, where identical requests succeed after several retries with no change on your end, is a classic signature of a service-side throttling or capacity issue rather than a client-side configuration problem. I checked the public Azure Status page and did not find a currently listed incident specifically tied to Content Understanding in Sweden Central around 26/06/2026, but the public status page frequently does not reflect narrower, service-specific degradations, so this does not rule out a backend issue on Microsoft's side.
Here is how to proceed on each of your questions:
Known regional issue: There is no publicly listed incident for Content Understanding in Sweden Central on the Azure Status page for that period. This does not mean nothing happened, since narrower service-level degradations are often only visible through Azure Resource Health on your specific resource, not the public status page. Check Resource Health on your Content Understanding resource in the Azure portal for that date range, and if it shows a platform-level event, that confirms it was not isolated to your account.
Throttling or service-side constraints on larger documents: Content Understanding in Foundry Tools reached GA relatively recently, at API version 2025-11-01, and larger or more complex documents place more load on the backend OCR, layout, and generative pipeline stages that prebuilt-documentSearch chains together. Even when you are well below the documented 300-page limit, elevated processing time per page increases the odds of hitting a transient capacity ceiling on the service side, which surfaces as InternalServerError rather than a specific throttling error code. This is consistent with what other customers have reported with Content Understanding analyzers on larger or more complex inputs.
Is prebuilt-documentSearch fully supported in GA 2025-11-01: Yes. prebuilt-documentSearch is documented as a supported RAG-oriented prebuilt analyzer as of the GA release, so this is not an unsupported or preview-only analyzer being called incorrectly.
Why it does not appear in the Foundry playground: The new Foundry portal introduced a prebuilt analyzers playground alongside the GA release, but playground UI rollout for individual analyzers does not always happen in lockstep with API-level GA support. This looks like a rollout gap in the playground UI rather than an indication that the analyzer itself is unsupported, since your REST API calls to it are otherwise being accepted and processed.
Recommended retry and batching practices, until this stabilizes:
Keep your current strategy of splitting large documents into smaller batches, since this has already reduced (though not eliminated) the failure rate for you.
Implement exponential backoff with jitter on `InternalServerError` specifically, since your evidence shows retries do eventually succeed:
```python
import time
import random
max_retries = 8 for attempt in range(max_retries): result = call_content_understanding_analyze(document_batch) if result.get("status") == "Failed" and result["error"]["code"] == "InternalServerError": wait = min(60, (2 ** attempt) + random.uniform(0, 1)) time.sleep(wait) continue break ```
Log the `apim-request-id` response header for every failed call. This is required if you escalate to Support, since it lets the Content Understanding backend team trace the exact failed request rather than working from timestamps alone.
Avoid re-submitting the same batch immediately back to back, since bursty retries can compound backend load during a capacity event rather than helping.
Given that this reproduces consistently on larger documents, clears on retry, and does not correlate with anything in your request configuration, I recommend opening an Azure Support case and including the apim-request-id values you have already captured, the exact timestamps of the failures, and the region. This gets it in front of the Content Understanding engineering team who can check backend telemetry for Sweden Central around your reported dates, which is not something visible from the client side.
If this answer helps you kindly accept the answer which will help others who have similar questions.
Best Regards,
Jerald Felix.