forked from rook/rook
The prepare job runs "ceph-volume raw prepare" per device and then trusts
a no-argument "ceph-volume raw list" to report the OSDs it just created.
That listing can come back without a freshly prepared device. On Ceph
v20.2.0 and v20.2.1, get_devices() enumerates through udev data under
/run/udev/data and returns nothing when that data is empty or not yet
populated (https://tracker.ceph.com/issues/77968). On every released
v20.2.x, a block device smaller than 4KiB in the scan, such as the 1KiB
extended-partition node of an MBR-partitioned OS disk, crashes the
batched ceph-bluestore-tool call behind the listing, which ceph-volume
reports as an empty result (https://tracker.ceph.com/issues/76354).
Either way the no-argument list comes back empty or incomplete even
though the OSD exists. Rook then silently reported fewer OSDs than it
prepared, the operator never created the missing OSD deployment, and the
cluster was left with OSDs in the osdmap but no daemons running.
Verify that every device prepared in raw mode is present in the listing.
For any that is missing, list it by name ("ceph-volume raw list
<device>"), which is immune to both failure modes and reports the OSD.
If even the per-device listing is empty, return an error so the prepare
job exits non-zero and Kubernetes restarts it with a freshly re-scanned
environment, rather than silently under-reporting the OSDs.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit 09ce9e26cc)