Hi Jones Oscar •
I think this is most likely a combination of warehouse sizing and data layout rather than a clustering issue alone. When a query starts spilling to local and especially remote storage, it indicates that the join operation requires more memory than the warehouse can provide, which leads to a significant performance penalty.
My first recommendation would be to temporarily scale the warehouse up one or two sizes and rerun the same query. If the spill volume drops substantially or disappears, you've confirmed that memory pressure is a major contributor. At the same time, review the join keys and filtering columns in your query profile. Clustering keys tend to provide the most benefit when they align with frequently used filter predicates and help reduce the amount of data scanned before the join occurs.
It's also worth checking for data skew. If a small number of join key values account for a large percentage of rows, a larger warehouse alone may not completely solve the problem. In those cases, query rewrites, pre-aggregation, or breaking large joins into smaller stages can often produce better results than simply adding compute resources.
As a general rule, addressing the spills first by validating warehouse sizing, then evaluating whether the clustering strategy is actually helping prune data for your most common workload patterns. The query profile should tell you which of these factors is having the biggest impact.
If you find this response helpful, please click "Accept Answer" so it can help other community members facing similar performance issues as well.