Azure Databricks Model Serving: Intermittent timeouts when calling scaled-to-zero endpoints

Megaritis, Vlasis 0 Reputation points
2026-10-08T10:03:36.3366667+00:00

I have an Azure Databricks workspace with a deployed model that runs daily on a GPU Small (T4) endpoint, configured for Small with 0–4 concurrency. The endpoint normally scales to zero between runs, and I use a script to wake it before sending requests.

Most days, the job completes successfully in about 7–8 minutes. Recently, though, it has intermittently taken over an hour without succeeding, eventually timing out. The wake-up step sometimes appears to complete, but the subsequent job still hangs or fails.

Does anyone have suggestions on what might cause this kind of intermittent behavior, or which endpoint diagnostics would be most useful to check?

Azure Databricks
Azure Databricks

An Apache Spark-based analytics platform optimized for Azure.


1 answer

Sort by: Most helpful
  1. AI answer

    2026-10-08T10:03:57.64+00:00

    The intermittent pattern matches cold-start and capacity behavior on scaled-to-zero GPU serving endpoints.

    Most relevant causes to check:

    1. Scale-to-zero warm-up latency
      • When an endpoint is scaled to 0, the next request can wait while the model is loaded again and replicas start.
      • For custom LLM endpoints, this can take one to several minutes.
      • For custom model serving more broadly, scaling up from zero usually takes 10–20 seconds, but it can sometimes take minutes, and there is no SLA on scale-from-zero latency.
      • If the wake-up request returns before the endpoint is fully ready, the following job can still hit client-side timeouts or appear to hang.
    2. GPU capacity is not guaranteed when waking from zero
      • On GPU endpoints, when Azure Databricks needs to acquire a GPU during endpoint creation, workload increase, or wake-up from zero, the request can stop responding if the cloud provider has no GPU capacity in the region.
      • The documentation explicitly calls out that GPU capacity is not guaranteed when scaling to zero, and first requests after scale-to-zero can have extra high latency.
      • That aligns well with runs that usually finish in minutes but occasionally stretch far longer or never recover.
    3. Server-side request timeout
      • Model serving requests time out at 597 seconds for CPU or GPU serving endpoints, and that limit cannot be increased.
      • If a request consistently fails near that boundary, it points to a server-side timeout.
      • If it fails earlier, check configuration and service logs for other errors.
    4. Client-side timeout during warm-up or downstream connection timeout
      • If the endpoint is waking up slowly, the client can time out before the endpoint is ready.
      • The guidance is to inspect service logs and inference tables for timeout-related messages such as "timed out" or "timeout".
      • If the serving pipeline connects to another service, connection timeouts in that dependency can also surface as intermittent endpoint failures.

    Most useful diagnostics to check:

    1. Serving endpoint → Events tab
      • Search for "timed out".
      • This helps identify deployment or endpoint timeout events.
    2. Serving endpoint → Logs tab
      • If Events shows timeout-related entries, inspect the logs for the underlying cause.
      • The documentation specifically recommends checking logs for dependency issues, resource constraints, and configuration issues.
    3. Inference tables
      • If enabled, use inference tables for root-cause analysis of failed requests.
      • In the workspace, go to Serving → select the endpoint → find the Inference tables section and copy the fully qualified table name.
      • Then query it from a notebook:
             %sql
             SELECT * FROM my-catalog.my-schema.my-table
        
      • This is useful for confirming whether requests are failing during warm-up, timing out near the 597-second limit, or returning other errors.
    4. Timing of failures
      • Compare the elapsed time of failed requests with the 597-second request timeout.
      • Near-597-second failures suggest server-side timeout.
      • Much earlier failures suggest warm-up, configuration, or dependency issues.

    Most likely mitigation based on this scenario:

    • Disable scale to zero for this production-style daily workload if consistent startup time matters. The guidance explicitly warns that scale to zero should not be used for workloads requiring consistent uptime or guaranteed response times, and Azure Databricks recommends turning it off for production traffic on custom LLM endpoints.
    • If scale to zero must remain enabled, expect cold-start delay and add client retry logic around the wake-up and first inference calls.
    • If the endpoint serves a very large GPU model, also review whether deployment/runtime behavior is being affected by GPU image creation time or regional GPU capacity constraints.

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.