osd-replacement was relying on exising ceph.rook.io/do-not-reconcile
label for fencing osd destroy process owned by osd health goroutine from
controller. However, this label can be used but other components like
rook krew maintenance plugin and cannot be owned by osd-replacement
process. Added a separate osd.rook.io/replace-in-progress annotation for
that purpose.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
if the account cr has a force delete annotation, forcefully
remove the account, even if it contatins data in it, using the
purge data flag
Signed-off-by: parth-gr <partharora1010@gmail.com>
Added failureDomains and osdsPerDomain fields to the
ErasureCodedSpec, which maps to the crush-num-failure-domains and
crush-osds-per-failure-domain Ceph EC profile parameters. This enables
creating EC pools that distribute chunks across fewer, larger hosts
without needing one host per data chunk.
Signed-off-by: Prabhala Tara Aasrita <taraprabhala@Prabhalas-MacBook-Pro.local>
The daemon only flagged a device change from in-use to empty, never the
reverse, so an in-use disk was never recorded as non-empty and a later
disk swap produced no ConfigMap update or reconcile. This is needed for
OSD replacement to work automatically: without it a destroyed OSD slot is
never reprovisioned after the disk is swapped. Compare the empty state in
both directions so the swap becomes a real transition that triggers the
reconcile.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
in case of multiple destroyed OSDs or multiple available
blank devices, Rook will try to match them by the same
device class if possible. If no matching DC, then DC of
destroyed OSD will be used
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
Destroyed OSDs awaiting replacement from osd-replacement flow should
be ignored by managed PDB to avoid the following situation:
replaced OSD is already destroyed, its deployment is downscaled zero
and it awaits the physical swap (can take days). This situation does
not trigger blocking PDB in happy case because OSD does not holds PGs.
However, if PGs go unclean during that window for an unrelated reason,
the destroyed OSD gets counted as a down OSD and can grab the single
"draining failure domain" slot - putting blocking PDBs on every other FD.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
A state machine for osd destroy phase of osd-replacement flow.
The state machine is implemented in OSD health goroutine, where each
step is handled on a separate tick. State machine is responsible
for draining, destroying OSD, and reserving its CRUSH position,
downscalling its deployment, and adding "readyForSwap" annotation.
Destroy phase does not zap the device on purpose. On disk signature
is kept to avoid automatic reprovisioning of destroyed OSD and to
detect physical swap.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
defines a new entry point "close-encrypted-devices" in osd job.
The job is used by osd-replacement flow to close dm-crypt mappings
on host for destroyed encrypted OSDs. The job is owned by OSD
health goroutine implementing destroy phase of osd-replacement.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
contains osd prepare job edits to support osd replacement flow:
- removes destroyed OSD IDs from node status CM to stop cluster
controller from reprovisioning destroyed OSD
- adds provisioning logic for OSD replacement to match available
blank devices to destroyed OSD slots
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
Add coverage for GetPort defaults and custom values, Service and probe
port wiring, Ganesha NFS_Port config, and create-or-update when
spec.server.port changes on an existing Service.
Signed-off-by: raaizik <raaizik@yahoo.com>
When NFS runs with host networking, port 2049 may already be in use.
Add CephNFS spec.server.port (default 2049), wire it through Ganesha
config, the operator Service, and probes, and regenerate CRDs.
Signed-off-by: raaizik <raaizik@yahoo.com>
the .nvmeof pool was incorrectly tagged with the rbd
application and initialized as an rbd pool. add .nvmeof
to the built-in pool switch so it gets the nvmeof-meta
application tag instead, and skip rbd pool init.
Signed-off-by: Oded Viner <oviner@redhat.com>
if the csi secret is updated or a client profile is updated
we need a re-reconile of rns
So added the watcher for it
Signed-off-by: parth-gr <partharora1010@gmail.com>
retrieveMultisiteZone is meant to gate the object-store reconcile on the
backing Ceph zone existing: it runs "radosgw-admin zone get" and, when
that fails, returns a non-nil error so the caller requeues. That gate is
dead code. The ENOENT check declares an inner err from exec.ExtractExitCode
that shadows the outer command error, and ExtractExitCode returns a nil
error for the ordinary exit failures radosgw-admin produces. Both the
ENOENT branch and the else branch then wrap that shadowed nil, and
errors.Wrapf(nil, ...) is nil, so the function returns
(waitForRequeueIfObjectStoreNotReady, nil). The caller only propagates the
requeue when the error is non-nil, so the requeue is dropped and reconcile
runs on.
The result is that a failed "zone get" no longer backs off. Reconcile
proceeds to stand up the object store anyway -- the RGW service, the
admin-ops endpoint, the deployment, and the pools radosgw scaffolds as it
comes up -- for a multisite store whose backing zone does not exist. This
is not the "normal multisite bootstrap" transient the original wording
suggested. getMultisiteResourceNames runs immediately before this and
already requeues until the CephObjectZone CR reports Ready, and the zone
controller marks it Ready only after it has created the Ceph zone, so a
healthy bootstrap never reaches this gate with a missing zone. That
CR-Ready gate, not this one, is what actually blocks bootstrap.
Where the dead gate does bite is the cases the CR status cannot cover:
- the Ceph zone deleted or renamed out of band while the CR still reads Ready
- a zone controller that reports Ready without leaving a usable zone behind
- any non-ENOENT "zone get" failure, e.g. a permission or connectivity
error, which the else branch swallows the same way
The check was correct when fc579f4520 introduced it in 2020 with
exec.ExitStatus, which returns (code, ok) and leaves err unshadowed.
bd58790c31 ("ceph: proxy ceph commands when multus is configured", 2021)
swapped it to exec.ExtractExitCode as an unrelated drive-by, inverting the
contract and killing the gate. The sibling realm, zonegroup, and zone
controllers were left untouched and still use exec.ExitStatus today.
Restore exec.ExitStatus so the outer error is no longer shadowed, both
branches wrap the real command error, and the caller requeues until the
zone exists -- matching the sibling controllers.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Changes:
- Use a dedicated .nvmeof pool (via CephBlockPool CR named
builtin-nvmeof) for the NVMe-oF gateway internal state.
- Use a separate nvmeof pool for the StorageClass data.
- Remove the pool field from the CephNVMeOFGateway CRD.
The gateway now always uses the .nvmeof pool, hardcoded
in the operator.
- Add .nvmeof to the allowed CephBlockPool name overrides.
- Create a production example nvmeof.yaml (instances: 2,
replicas: 3) and a CI-only nvmeof-test.yaml (instances: 1,
replicas: 1).
- Update documentation and CI test script accordingly.
Signed-off-by: Oded Viner <oviner@redhat.com>
executeCommandWithTimeout wires a single bytes.Buffer to both cmd.Stdout
and cmd.Stderr and joins the command in a goroutine via cmd.Wait(). On
the timeout kill path it called cmd.Process.Kill() and then read that
buffer without waiting for cmd.Wait() to return. Kill() only signals the
process; it does not wait for os/exec's output-copier goroutines (joined
only by cmd.Wait) to finish, so reading the buffer there races with those
writers, which the race detector flags.
Only read the buffer once the copier goroutines have drained, signaled by
cmd.Wait() on the done channel. Because a killed process can leave an
orphaned descendant holding the output pipe open -- a D-state
cryptsetup/dmsetup child, or the downstream of a "sh -c '... | head'"
pipeline -- cmd.Wait() may never return, so bound the drain with a short
grace period and give up on the captured output rather than blocking the
caller forever. This keeps the bounded return the timeout path is meant
to guarantee. On a failed Kill() the process may still be writing, so
return without reading the buffer.
Add a regression test that drives the kill path with a child that ignores
SIGINT while writing to stdout; it fails under -race before this change.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
startOSDMigration guards against upgrading OSDs while an OSD migration is
still in progress by returning the set of OSDs pending migration, which the
caller then removes from the update queue.
When the last migrated OSD's deployment has not been recreated yet,
isLastOSDMigrationComplete returns (false, nil). startOSDMigration then ran
return nil, errors.Wrapf(err, ...) with err already nil, and
errors.Wrapf(nil, ...) returns nil, so it returned (nil, nil), the same
value it uses to signal that no migration was requested. The caller
therefore skipped removing the pending OSDs from the update queue and could
upgrade and restart them mid-migration, which the pending-migration guard
was meant to prevent.
Return the freshly computed migration config in this case so the caller
still prunes the pending OSDs, without starting a new migration and without
aborting the reconcile. Aborting here would wedge the cluster because the
code that recreates the interrupted OSD runs later in Start, so an early
error return would leave that OSD deleted and block every later reconcile.
Letting the reconcile continue recreates the interrupted OSD and lets the
migration resume on a subsequent reconcile.
Also construct a non-nil error for the sibling unhealthy-PGs guard, which
wrapped a nil err in the same way.
Add a regression test asserting that startOSDMigration returns the pending
migration set, and does not delete another OSD deployment, while a previous
migration is still in progress.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
archiveCrash is meant to silence the RECENT_CRASH health warning for a
removed OSD by archiving its crash, but the early-return guard was
inverted. GetCrashList unmarshals `ceph crash ls` into a non-nil slice on
every success, so `if crash != nil` was always true. The function always
logged "no ceph crash to silence" and the archive loop was unreachable, so
OSD removal never archived the crash and the warning lingered. Guard on
`len(crash) == 0` so only a genuinely empty list short-circuits.
Reviving the loop surfaced two latent bugs it had always masked. On a
non-empty list with no entry for the removed OSD the crash id stayed empty
and the code ran `ceph crash archive ""`, which the mgr rejects with EINVAL
and logs as a spurious error during routine OSD removal, so skip the
archive when no crash matches. The loop also broke on the first match, but
`ceph crash ls` is oldest-first and includes already-archived entries, so
an OSD with several crashes had only its oldest one archived while
RECENT_CRASH persisted; archive every matching entry instead of just the
first.
Extend TestArchiveCrash to cover all four cases: a single matching crash, a
list with several matching crashes that are all archived, an empty list,
and a non-empty list with no matching entry where nothing is archived.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
In createOrUpdateCephUser, when the desired user carries no explicit
keys the reconciler falls back to the live RGW user's keys. If that live
user also has zero keys the code intends to fail, but it built the error
with errors.Wrapf(err, ...) at a point where err is already nil (the
prior SetUserQuota error was handled and returned just above).
errors.Wrapf(nil, ...) returns nil, so the failure was swallowed and
createOrUpdateCephUser returned success on a user the operator itself
flagged as broken.
The reconcile then continued to generateCephUserSecret, which indexes
userConfig.Keys[0] to populate the Kubernetes secret and panicked on the
empty key slice. The reconciler's deferred RecoverAndLogException caught
and logged that panic, so the reconcile was abandoned before it reached
the Ready status update; and because the recovered Reconcile returns a
zero Result with a nil error, the request was not requeued either. The
user was left neither marked Ready nor retried.
Construct the error with errors.Errorf so the intended failure is
surfaced instead of being swallowed. Add a regression test that returns
a keyless live user and asserts a non-nil "no keys set" error.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>