An Azure NoSQL database service for app development.
Hi @nicucern
The quickest way to find out if these stalls happen inside Cosmos DB or on the way to it is to log the SDK diagnostics for the slow calls only. With the .NET SDK it looks like this:
- After each call, check
response.Diagnostics.GetClientElapsedTime(). If it's over your threshold (say 1 second), logresponse.Diagnostics.ToString(). Do the same withcosmosException.Diagnosticsin your catch block. - In a slow entry, compare
BELatencyInMswith the "Transit time" event. Small BELatencyInMs with a big transit time means the delay is on the network or client side. A big BELatencyInMs means the time was spent on the service. - Also look at
systemHistoryin the same entry. High CPU orisThreadStarving: Trueat the time of the stall points to your app host, not the partition. - If the backend latency is high and it keeps hitting the same partition key range, open a support request and include a few of these diagnostics strings with their activity IDs. That gives the Cosmos DB team what they need to look at partition 12 directly.
Which SDK and connection mode (Direct or Gateway) is your app using?
If this resolved your issue, please consider accepting it as the answer. If the stalls continue, let me know and I'll be happy to keep helping.
Reference: