Intermittent 424 Client Error: Failed Dependency on Azure ML Managed Endpoint during pytest HTTP suite

Jaiswal, Snehil 0 Reputation points
2026-09-29T11:44:21.5233333+00:00

We are observing an intermittent requests.exceptions.HTTPError: 424 Client Error: Failed Dependency when making HTTP requests to our Azure ML Managed Endpoint

The endpoint is invoked as part of an automated testing suite using pytest and the Python requests library.

  • Behaviour: Out of all the HTTP tests executed sequentially, exactly one test fails randomly with a 424 Failed Dependency status code. The remaining test requests succeed without any issue.
  • Frequency: This issue has occurred twice over the past month. It is non-deterministic and fixes itself automatically when the flow is re-run after a short period.
  • Log Findings: Upon investigating the deployment logs in Azure ML Studio, no logs or traces are generated for the specific request that receives the 424 status code. The request appears to be rejected at the Azure ML Frontend Router/Gateway level before reaching the scoring container or executing the score.py entry script.
  • Documentation Check: We reviewed the Azure ML Online Endpoint troubleshooting documentation regarding HTTP 424 status codes, but our scenario does not match any listed cases (e.g., scoring script crash, explicit upstream timeouts, or container startup failures). All external dependencies in our infrastructure remain healthy during these runs.

Request Details / Diagnostics:

Region: North Europe

Client Implementation: Python (pytest test runner using requests.post())

Observed HTTP Status: 424 Client Error: Failed Dependency

Azure Machine Learning
0 comments No comments

1 answer

Sort by: Most helpful
  1. AI answer

    2026-09-29T11:44:42.84+00:00

    For Azure Machine Learning online endpoints, HTTP 424 means the model container returned a non-200 response, and Azure surfaced that as 424 Model error. The documented first checks are:

    1. Check the endpoint’s Azure Monitor Metric Explorer and inspect the Model Status Code dimension under Requests Per Minute.
    2. Inspect the response headers ms-azureml-model-error-statuscode and ms-azureml-model-error-reason for the failing request.
    3. If 424 occurs together with readiness or liveness probe failures, adjust ProbeSettings to allow more time.

    Because the failing request produces no container logs, the same troubleshooting article also states that when get-logs does not produce logs, it usually means the container failed to start. In that case, use local deployment to test and debug the model locally before cloud deployment, and use container logs when available.

    For intermittent failures that succeed on retry, the supported guidance in the provided material is to add lightweight retry handling with exponential backoff. That pattern is explicitly recommended in Microsoft guidance for transient failures, and the Azure ML endpoint troubleshooting article also recommends exponential backoff for request overload scenarios.

    A practical path for this case is:

    1. Capture the failing response headers, especially ms-azureml-model-error-statuscode, ms-azureml-model-error-reason, and the request tracing headers.
    2. Send a client tracing header such as x-ms-client-request-id on every pytest request so the failing call can be correlated.
    3. Record the server-generated x-request-id from the failed response. For failed requests, that ID should be attached to a support ticket, along with the region and endpoint name.
    4. Check Azure Monitor metrics for the endpoint at the failure time, especially Requests Per Minute with the Model Status Code dimension.
    5. If there is any evidence of probe instability, increase readiness/liveness probe timing.
    6. Add retry logic in the test harness with exponential backoff for this intermittent 424, since the failure is rare, non-deterministic, and clears on rerun.

    Given the described pattern—one random failure, no scoring-container logs, and success on rerun—the most grounded conclusion is that the next actionable step is correlation and escalation using x-request-id, rather than assuming a deterministic score.py defect.


    References:

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.