An Azure machine learning service for building and deploying models.
For Azure Machine Learning online endpoints, HTTP 424 means the model container returned a non-200 response, and Azure surfaced that as 424 Model error. The documented first checks are:
- Check the endpoint’s Azure Monitor Metric Explorer and inspect the Model Status Code dimension under Requests Per Minute.
- Inspect the response headers
ms-azureml-model-error-statuscodeandms-azureml-model-error-reasonfor the failing request. - If
424occurs together with readiness or liveness probe failures, adjust ProbeSettings to allow more time.
Because the failing request produces no container logs, the same troubleshooting article also states that when get-logs does not produce logs, it usually means the container failed to start. In that case, use local deployment to test and debug the model locally before cloud deployment, and use container logs when available.
For intermittent failures that succeed on retry, the supported guidance in the provided material is to add lightweight retry handling with exponential backoff. That pattern is explicitly recommended in Microsoft guidance for transient failures, and the Azure ML endpoint troubleshooting article also recommends exponential backoff for request overload scenarios.
A practical path for this case is:
- Capture the failing response headers, especially
ms-azureml-model-error-statuscode,ms-azureml-model-error-reason, and the request tracing headers. - Send a client tracing header such as
x-ms-client-request-idon every pytest request so the failing call can be correlated. - Record the server-generated
x-request-idfrom the failed response. For failed requests, that ID should be attached to a support ticket, along with the region and endpoint name. - Check Azure Monitor metrics for the endpoint at the failure time, especially Requests Per Minute with the Model Status Code dimension.
- If there is any evidence of probe instability, increase readiness/liveness probe timing.
- Add retry logic in the test harness with exponential backoff for this intermittent
424, since the failure is rare, non-deterministic, and clears on rerun.
Given the described pattern—one random failure, no scoring-container logs, and success on rerun—the most grounded conclusion is that the next actionable step is correlation and escalation using x-request-id, rather than assuming a deterministic score.py defect.
References: