A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference
Hello @Julius Dietmar ,
Welcome to Microsoft Q&A .Thank you for reaching out to us.
Thank you for sharing the telemetry and details regarding the embed-v-4-0 latency increase.
Based on the information provided, the behavior indicates a significant latency anomaly compared with the reported normal response time.
The approximately 60-second plateau closely matches the configured one-minute application timeout. Therefore, these values may represent requests reaching the client/application timeout rather than exactly 60 seconds of model-processing time.
Regarding the Central US and deployment type - The deployment SKU/type should first be confirmed before associating the issue with Central US.
Foundry deployment types have different processing and performance characteristics. Global deployments can dynamically route inference across available Azure datacenters, while Data Zone and geography-based deployments use different processing boundaries.
If the affected deployment is confirmed as GlobalStandard, the Central US location of the resource does not by itself establish that inference was processed in Central US. Global Standard can also experience greater latency variability at high, consistent workload volumes. Provisioned deployment types are intended for workloads requiring more predictable throughput and lower latency variance.
Regarding Service Health showing no degradation - The absence of a Service Health incident does not rule out a workload-specific performance regression.
Service Health covers broader service issues, planned maintenance, and health advisories, while Azure Monitor can detect workload-specific regressions such as increased errors or latency
The following are the troubleshooting steps
- Please confirm the deployment and incident details
- Resource ID and deployment name
- embed-v-4-0 model/version
- Deployment SKU/type
- Resource region and endpoint/API
- Exact start and end timestamps in UTC
- Configured application timeout
- Exact metric represented in the supplied latency graph
- Whether the latency returned to normal without an application or deployment change
- Reviewing Azure Monitor for the same UTC interval Kindly check for :
- ModelRequests, filtered or split by StatusCode
- ModelAvailabilityRate
- HTTP 429, other 4xx, and 5xx responses
- Latency/performance metrics available for the applicable deployment type
- Region, ModelDeploymentName, ModelName, and ModelVersion dimensions where available
- Request/response and trace logs, if diagnostic logging was already enabled
Please check if the following workarounds help-
- If sustained 429 throttling occurs, use exponential backoff and honor the Retry-After header.
- If request volume or concurrency increased during the affected period, reduce the workload temporarily and compare latency against the normal baseline.
- For consistently high workloads where low latency variance is important, consider a provisioned deployment type where supported. Provisioned deployments provide reserved processing capacity and more predictable throughput
- For resiliency-sensitive production workloads, a multi-region deployment design can provide an alternate path when a deployment or region becomes unhealthy
The following references might be helpful , please check them out
- Understanding deployment types in Microsoft Foundry Models - Microsoft Foundry | Microsoft Learn
- Stay informed about service health regressions - Microsoft Foundry | Microsoft Learn
- Monitor Model Deployments in Microsoft Foundry Models - Microsoft Foundry | Microsoft Learn
Please let us know if the response was helpful
Thank you