Files
Artem Torubarov 0c27a95097 osd: cleanup destroyed osd deployment if replacement failed
If osd-replacement fails during provisioning of the new device by the
osd-prepare job (can be the case if the new device is faulty), then
ceph-volume will run its internal rollback logic, which purges the
reserved osd id and its CRUSH position. After the faulty device is
replaced with a working one, rook will reprovision it correctly and
reuse the metadata device slot, but as a new OSD, so data will likely
be rebalanced. Ceph hands out the lowest free id, so the new OSD often
takes the same number back, and then the create path recreates the
deployment by itself. Otherwise, and until a working device arrives,
the downscaled deployment of the destroyed osd is left behind. This
commit handles that case by deleting downscaled ready-for-swap osd
deployments without a corresponding OSD in the osd dump.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-08-04 11:28:21 +02:00
..

Rook Feature Proposal

Please follow the design section guideline.