Azure ML Online Deployment Stuck in Deleting state

Phillip Stenger 10 Reputation points
2026-08-14T19:25:56.39+00:00

I have an Azure ML Online Endpoint Deployment stuck in "Deleting". It has been stuck for multiple hours and is blocking our deployment pipeline. After a while the delete operation timed out and I tried to delete again and it is now stuck still.

I have tried the following resolutions with no luck:

  • Checked for a resource lock on the deployment or the endpoint.
  • Ensured Traffic is allocated to 0% for the deployment
  • I have tried deleting the endpoint, but endpoint deletion is dependent on deployment deletion so that is stuck too
  • I have tried using the cli az ml online-deployment delete. The command completes however the deployment is not actually deleted.
Azure Machine Learning

2 answers

Sort by: Most helpful
  1. Manish Deshpande 8,215 Reputation points Microsoft External Staff Moderator
    2026-08-20T01:40:26.51+00:00

    Hello @Phillip Stenger

    The deployment deletion was blocked by a management lock on the workspace's backing storage account advartdevcorect9. Our backend telemetry corroborates it precisely, and I want to give you the full picture — including the timeline, why your own lock check legitimately came back clean, and how to stop this recurring.

    Root cause Confirmed

    A CanNotDelete management lock on storage account advartdevcorect9 in resource group rg-advart-dev blocked Azure Resource Manager from deleting two child objects that the DeleteDeployment workflow must remove: the storage privateEndpointConnectionProxy and the role assignments held against that account. The backend recorded the explicit error at 12:25 UTC on 18 August:

    ResourceConflictException: The scope

    '.../storageAccounts/advartdevcorect9/providers/Microsoft.Authorization/roleAssignments/33eb42dc-280f-4c5b-8cf6-33ba164e93af' cannot perform delete operation because following scope(s) are locked: '/subscriptions/0d3330fc-.../resourceGroups/rg-advart-dev/...'
    

    This is documented ARM behaviour, not a defect in the lock system: a cannot-delete lock on a resource or resource group prevents deletion of Azure RBAC assignments, and a Delete lock blocks the entire operation rather than partially completing it.

    Why your lock check came back clean. You checked the deployment and the endpoint, which is exactly where anyone would look. But locks inherit downward: per the documentation, "When you apply a lock at a parent scope, all resources within that scope inherit the same lock," and "Extension resources inherit locks from the resource to which they're applied." So a CanNotDelete lock on the storage account — or on the resource group — is fully effective against the delete while being completely invisible on the endpoint and deployment blades. Deleting an online deployment has to remove objects hanging off that storage account, including its role assignments, and the docs are explicit: "A cannot-delete lock on a resource or resource group prevents the deletion of Azure RBAC assignments." There was nothing wrong with your troubleshooting; the diagnostics didn't point you anywhere near it.

    It also explains the partial-progress behaviour you saw: "a Delete lock… blocks the whole delete operation. Even if the resource group or other resources in the resource group are unlocked, the deletion doesn't happen. A partial deletion isn't possible."

    On the exception messages — your feedback is fair and I've raised it internally. ARM produces a precise, actionable error naming the locked scope. What surfaced to you was a generic internal error pointing at a troubleshooting guide with no section on failed deletions and no mention of locks. Separately, az ml online-deployment delete returns on ARM's acceptance of the request, not on completion of the backend workflow — which is why the command reported success while nothing had been deleted. I've logged both with the Azure Machine Learning engineering team as a request to surface the underlying ARM dependency error to the caller rather than collapsing it into a generic failure. I can't give you a fix date, but the feedback is in with the specific evidence attached.

    To stop this recurring in the pipeline:

    1. Enumerate every lock in the subscription before a teardown, not just on the target resource.

    az lock list -o table
    

    This returns all locks in the subscription in one call, so inherited resource-group and subscription-scope locks appear alongside resource-scope ones. Checking the deployment and endpoint individually will always miss those. Confirm nothing returned covers the workspace, its resource group, or its storage account, Key Vault, container registry, or Application Insights.

    2. When a delete stalls, go to the Activity Log rather than the CLI or portal message. Filter to the workspace's resource group over the window of the failed operation. Until the error propagation improves, that's the surface that actually carries the blocking scope path.

    3. Handle the lock explicitly in the pipeline, so a re-applied lock doesn't hang the next teardown:

    lockid=$(az lock show --name <LockName> --resource-group <rg> \
      --resource-type Microsoft.Storage/storageAccounts \
      --resource-name <storageaccount> --output tsv --query id)
    az lock delete --ids $lockid
    
    az ml online-deployment delete --name <deployment> --endpoint <endpoint> --yes
    
    az lock create --name <LockName> --lock-type CanNotDelete \
      --resource-group <rg> --resource-name <storageaccount> \
      --resource-type Microsoft.Storage/storageAccounts
    

    One prerequisite worth checking now rather than at 2 a.m.: creating or removing locks requires Microsoft.Authorization/* or Microsoft.Authorization/locks/* — held by Owner and User Access Administrator. A pipeline service principal with Contributor cannot remove a lock, so if that's your identity, grant it the lock permission or exclude the AML dependency resources from the governance policy that applies the lock.

    Reference: Lock your Azure resources to protect your infrastructure

    Thanks again for closing the loop with the root cause — and for the feedback on the error text. It's the more valuable half of this thread.

    Thanks,
    Manish.

    Was this answer helpful?

    0 comments No comments

  2. Phillip Stenger 10 Reputation points
    2026-08-19T16:27:59.7266667+00:00

    Issue was related to the backing storage account having resource lock. The exception messages from the ML services need a lot of work.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.