Serverless SQL pool database stuck in RECOVERY_PENDING even after CMK and identity fixes

Matt Bryson 0 Reputation points
2026-09-09T20:10:11.6033333+00:00

I have a Synapse workspace where one of my serverless SQL databases is stuck in RECOVERY_PENDING, and I’m unable to clear the error through any customer‑accessible method. I also do not have the ability to open a support ticket, so I’m posting here hoping a Microsoft engineer can assist.

Environment details

  • Workspace: synw‑data‑ncas
  • Resource group: rg‑data‑ncas‑core
  • Region: (same region as Key Vault — can add if needed)
  • Encryption: Customer‑Managed Key (CMK) enabled at workspace level
  • Key Vault: kv‑lsst
  • Serverless SQL pool: built‑in
  • Identity: tested with both System‑Assigned MI and a User‑Assigned MI (UAMI)

Workspace keys

Output of az synapse workspace key list:

[
  {
    "name": "default",
    "keyVaultUrl": "https://kv-lsst.vault.azure.net/keys/workspaceEncryption",
    "isActiveCmk": false
  },
  {
    "name": "45988138-245c-4d5b-acad-77bb92e9e00f",
    "keyVaultUrl": "https://kv-lsst.vault.azure.net/keys/tempEncryption",
    "isActiveCmk": true
  }
]

Database states

From SELECT name, state_desc, is_encrypted FROM sys.databases;:

master         ONLINE            False
NCAS_Archive   RECOVERY_PENDING  True

So the built-in master database is healthy but NCAS_Archive cannot be brought online.

Actions I have already tried (complete list)

  1. Switched workspace CMK between the original key (workspaceEncryption) and temporary key (tempEncryption).
  • Synapse only allows versionless Key Vault URLs for workspace keys, so I re‑registered with the correct format.
  1. Confirmed Key Vault configuration
  • Key enabled
  • Purge protection enabled
  • Soft delete enabled
  • Verified validity/expiration dates
  1. Validated Key Vault permissions
  • Workspace identity granted get, wrapKey, unwrapKey
  • Permissions tested under both System‑Assigned MI and a UAMI with Key Vault Crypto Officer*
  1. Switched workspace identity
  • Tested using UAMI instead of System‑Assigned MI
  • Re‑granted KV permissions after switching
  • No change in behavior
  1. Checked serverless connectivity and storage access
  • Storage account reachable
  • Other workspace components operational
  • Only NCAS_Archive fails recovery

Current symptoms

  • master database loads correctly and is ONLINE
  • Only NCAS_Archive is stuck in RECOVERY_PENDING
  • No customer operations (CMK, MI, KV, SQL commands) have changed its state
  • The database cannot be dropped, recovered, or repaired by any exposed tool

Why I'm posting here

There is a known scenario where serverless SQL pools get stuck in RECOVERY_PENDING due to backend issues with workspace metadata or the managed identity certificate, which can only be corrected internally by Microsoft. I believe I may be in that situation. Since I cannot open a support ticket, I’m requesting assistance from a Microsoft engineer to:

  • Inspect the backend state of the serverless SQL pool
  • Validate the workspace encryption metadata
  • Reset any stuck identities/certificates if necessary
  • Manually force recovery of the NCAS_Archive database if required

Any help would be greatly appreciated. I’ve exhausted all user‑accessible recovery options. Thank you.

Azure Synapse Analytics
Azure Synapse Analytics

An Azure analytics service that brings together data integration, enterprise data warehousing, and big data analytics. Previously known as Azure SQL Data Warehouse.

0 comments No comments

1 answer

Sort by: Most helpful
  1. Aditya Singh Rathore 195 Reputation points
    2026-09-21T11:28:32.0066667+00:00

    Hi @Matt Bryson ,

    Thanks for listing everything you've already tried. I'm not a Microsoft engineer and can't touch the backend of a serverless pool, so I can't force recovery for you. A few thoughts that might still help.

    This has been fixed from the backend before. A 2024 thread describes the same symptom on a CMK workspace, with serverless databases stuck in Recovery Pending. The cause was a backend service that failed to rotate the system-assigned managed identity certificate, and support fixed it by manually resetting the certificate. I can't say that's your cause, but it fits your suspicion that Microsoft needs to step in. Microsoft Learn

    Documented causes worth ruling out first:

    • Old key versions: The docs say to keep old keys and key versions enabled, with a future expiry date, until every pool is back online and re-encrypted, because a disabled or expired old version blocks decryption. In the same thread a Microsoft moderator listed a disabled key version that was still in use as a cause, fixed by re-enabling it. You've checked the key, but have you checked every version of workspaceEncryption, including soft-deleted ones? Microsoft LearnMicrosoft Learn
    • Key Vault firewall with a UAMI: That combination isn't supported, and the suggested mitigation was to remove the vault firewall, reactivate the pool, then switch to the system-assigned identity. If your vault has a firewall or private endpoint, this could matter. It would also help to know which identity was active when the database went pending. Microsoft Learn
    • Permission model: With access policies, the docs say to choose the "Application-only" option for the workspace identity. Also make sure the vault is on either access policies or RBAC, because roles granted under the other model are ignored. GitHub

    On the ticket. Two different things can block you. Creating a request needs Owner, Contributor or Support Request Contributor at subscription level. If that's the problem, someone with Owner can file it or grant you the role. Billing, subscription and quota requests are free, but technical support needs a paid plan. You could also try @AzureSupport on X, which a moderator has suggested to others who couldn't raise a request. You could ask a Microsoft-affiliated moderator here whether they can route it internally too, though I can't promise they can. Microsoft LearnMicrosoft Learn

    When you do get a ticket, include the workspace resource ID, database name, the time it went pending, and the exact time and order of each key and identity change. Don't post subscription IDs in this public thread.

    I'd hold off on more key or identity switching until you hear back, since each change adds variables. As far as I know, serverless databases hold metadata, not the data itself. It's worth having scripts for the objects in NCAS_Archive in case rebuilding becomes the fallback.

    Hope that gets you unstuck.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.