Eine Apache Spark-basierte Analyseplattform, die für Azure optimiert ist
Hi @rene bauer
In addition to the information above, I appreciate the detailed analysis and clear reproduction steps - they were very helpful.
From the behavior observed and the backend validation, this appears to be related to how Azure Databricks currently handles storage resolution for certain workspace-managed/serverless scenarios rather than a metastore configuration failure itself.
Your metastore root storage configuration is valid and correctly reflected in DESCRIBE METASTORE. However, for:
- auto-created workspace catalogs
- Serverless Lakeflow Pipelines
- Preview channel serverless workloads
managed tables may still be created on Databricks-managed internal storage instead of automatically inheriting the metastore root location.
So although the general Unity Catalog hierarchy is documented as: Schema → Catalog → Metastore
there are currently scenarios where serverless/workspace-managed assets do not fully follow that fallback behavior and instead use Databricks-managed storage.
This also explains the 403 behavior on classic clusters:
- the tables are stored on Databricks-managed internal storage
- RBAC permissions for that storage cannot be managed directly from your Azure subscription
- therefore USER_ISOLATION classic compute cannot access those tables unless the data is migrated
For future workloads, we strongly recommend explicitly configuring a managed location at the catalog level:
ALTER CATALOG <catalog_name>
SET MANAGED LOCATION 'abfss://<container>@<storage>.dfs.core.windows.net/<path>/';
This is the most reliable way to ensure that newly created managed tables - including those created through Serverless Lakeflow Pipelines - use your intended ADLS Gen2 storage location.
For the existing tables already placed on Databricks-managed storage, migration would still be required. Common approaches are:
- DEEP CLONE
- CTAS (CREATE TABLE AS SELECT)
- recreating managed tables into the new catalog location
At present, directly granting RBAC access to the Databricks-managed internal storage account is generally not possible because the storage exists within a Databricks-managed subscription.
I understand the documentation can make the fallback behavior appear broader than what is currently observed with some serverless/preview scenarios, and your feedback here is valuable.
Please let us know if you would like guidance on the safest migration approach for the existing ~744 GB of data.