Cosmos DB: Container Stall — Recurring Partition Latency Issue

nicucern 10 Reputation points
2026-10-07T11:19:07.4233333+00:00

Problem description

I am experiencing intermittent stalls on my Cosmos DB containers that use the SQL API. The issue started with certain partitions showing significant delays. The stalls affect various operations, including point reads, upserts, and queries, especially on partition key range ID 12. I have noticed that some requests time out or return delays of over 8 seconds, set in my app, where normally same operation would take under 100 ms. This has been as well verified, since the same operation is retried with under 100ms. This is the second container, as there was another container with issues on partition key range id 3 earlier that is still ongoing. For the first container these incidents seem to happen with a repeated frequency of either once in 15 minutes or 7 minutes, with spikes up to 4 seconds.

Environment

Azure Cosmos DB account using the API for NoSQL (SQL API) in the Germany West Central region, with containers within a SQL database. The account is configured in serverless mode with no provisioned RU/s or autoscale settings.

What I've already tried

I reviewed the case details and observed the latency patterns. I checked for errors such as 429 or 408 responses and monitored server-side request latency, client-side request latency, and normalized RU consumption. I also examined the operation types and partition key range IDs involved. Additionally, I verified the consistency setting and noted that the account uses session consistency with no recent changes in throughput, indexing, or partition key design.

Current status

I am seeking guidance on diagnosing the root cause of these recurring partition stalls, and recommendations for mitigating or resolving the latency problems.

Azure Cosmos DB
Azure Cosmos DB

An Azure NoSQL database service for app development.

0 comments No comments

1 answer

Sort by: Most helpful
  1. Rukshan edirisinghe 1,075 Reputation points
    2026-10-07T12:29:18.19+00:00

    Hi @nicucern

    The quickest way to find out if these stalls happen inside Cosmos DB or on the way to it is to log the SDK diagnostics for the slow calls only. With the .NET SDK it looks like this:

    1. After each call, check response.Diagnostics.GetClientElapsedTime(). If it's over your threshold (say 1 second), log response.Diagnostics.ToString(). Do the same with cosmosException.Diagnostics in your catch block.
    2. In a slow entry, compare BELatencyInMs with the "Transit time" event. Small BELatencyInMs with a big transit time means the delay is on the network or client side. A big BELatencyInMs means the time was spent on the service.
    3. Also look at systemHistory in the same entry. High CPU or isThreadStarving: True at the time of the stall points to your app host, not the partition.
    4. If the backend latency is high and it keeps hitting the same partition key range, open a support request and include a few of these diagnostics strings with their activity IDs. That gives the Cosmos DB team what they need to look at partition 12 directly.

    Which SDK and connection mode (Direct or Gateway) is your app using?

    If this resolved your issue, please consider accepting it as the answer. If the stalls continue, let me know and I'll be happy to keep helping.

    Reference:

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.