I’m following up to check whether the issue has been resolved. Feel free to reply if you need further information. If the information provided was helpful, please click "Accept Answer" to help others in the community. Thank you!
S2D cluster SSD wear and replacement
Hi guys
- Inspect drive wear metrics : Failed
- Attempt live SSD replacement :Failed
- Run command :Failed
Get-PhysicalDisk | Select FriendlyName, MediaType, Wear
Get-StorageReliabilityCounter -PhysicalDisk <DiskName> Check detailed endurance and write‑life statistics.
Update-StorageFirmware -PhysicalDisk <DiskName> Apply firmware updates if supported.
So, any other ways to perform zero‑downtime drive replacements ?
Windows for business | Windows 365 Business
2 answers
Sort by: Most helpful
-
Jason Nguyen Tran 27,050 Reputation points Independent Advisor
2026-09-29T05:43:40.2966667+00:00 -
Jason Nguyen Tran 27,050 Reputation points Independent Advisor
2026-09-11T09:51:18.98+00:00 Hello HAILEY PORTER,
In Storage Spaces Direct (S2D), zero‑downtime drive replacement can be tricky when wear metrics fail or firmware updates don’t apply. Normally, the supported workflow is to retire the disk gracefully and let the cluster rebuild data onto healthy drives before physically swapping it out.
If
Set-PhysicalDisk -Usage Retireddoes not succeed, another option is to useRemove-PhysicalDiskin combination withRepair-ClusterStoragePoolto trigger data redistribution. This ensures the pool maintains redundancy while the faulty drive is being removed. You can also useSuspend-ClusterNodetemporarily to drain workloads before maintenance, though this is more disruptive. For firmware or health checks,Get-StorageReliabilityCounterremains the best way to confirm endurance statistics before deciding on replacement.Best practice is to always validate that sufficient capacity exists in the pool before retiring a disk, otherwise repair jobs may stall. It’s also recommended to stagger replacements and monitor
Get-StorageJobto confirm completion before proceeding with additional changes. In environments where live replacement consistently fails, scheduling maintenance windows and scripting retire/repair operations is often the most reliable path.I hope the response provided some helpful insight. If you find this answer useful, please hit “accept answer” so I know it addressed your concern.
Jason