osd-replacement was relying on exising ceph.rook.io/do-not-reconcile
label for fencing osd destroy process owned by osd health goroutine from
controller. However, this label can be used but other components like
rook krew maintenance plugin and cannot be owned by osd-replacement
process. Added a separate osd.rook.io/replace-in-progress annotation for
that purpose.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
if the account cr has a force delete annotation, forcefully
remove the account, even if it contatins data in it, using the
purge data flag
Signed-off-by: parth-gr <partharora1010@gmail.com>
Destroyed OSDs awaiting replacement from osd-replacement flow should
be ignored by managed PDB to avoid the following situation:
replaced OSD is already destroyed, its deployment is downscaled zero
and it awaits the physical swap (can take days). This situation does
not trigger blocking PDB in happy case because OSD does not holds PGs.
However, if PGs go unclean during that window for an unrelated reason,
the destroyed OSD gets counted as a down OSD and can grab the single
"draining failure domain" slot - putting blocking PDBs on every other FD.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
A state machine for osd destroy phase of osd-replacement flow.
The state machine is implemented in OSD health goroutine, where each
step is handled on a separate tick. State machine is responsible
for draining, destroying OSD, and reserving its CRUSH position,
downscalling its deployment, and adding "readyForSwap" annotation.
Destroy phase does not zap the device on purpose. On disk signature
is kept to avoid automatic reprovisioning of destroyed OSD and to
detect physical swap.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
defines a new entry point "close-encrypted-devices" in osd job.
The job is used by osd-replacement flow to close dm-crypt mappings
on host for destroyed encrypted OSDs. The job is owned by OSD
health goroutine implementing destroy phase of osd-replacement.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
Add coverage for GetPort defaults and custom values, Service and probe
port wiring, Ganesha NFS_Port config, and create-or-update when
spec.server.port changes on an existing Service.
Signed-off-by: raaizik <raaizik@yahoo.com>
When NFS runs with host networking, port 2049 may already be in use.
Add CephNFS spec.server.port (default 2049), wire it through Ganesha
config, the operator Service, and probes, and regenerate CRDs.
Signed-off-by: raaizik <raaizik@yahoo.com>
the .nvmeof pool was incorrectly tagged with the rbd
application and initialized as an rbd pool. add .nvmeof
to the built-in pool switch so it gets the nvmeof-meta
application tag instead, and skip rbd pool init.
Signed-off-by: Oded Viner <oviner@redhat.com>
if the csi secret is updated or a client profile is updated
we need a re-reconile of rns
So added the watcher for it
Signed-off-by: parth-gr <partharora1010@gmail.com>
retrieveMultisiteZone is meant to gate the object-store reconcile on the
backing Ceph zone existing: it runs "radosgw-admin zone get" and, when
that fails, returns a non-nil error so the caller requeues. That gate is
dead code. The ENOENT check declares an inner err from exec.ExtractExitCode
that shadows the outer command error, and ExtractExitCode returns a nil
error for the ordinary exit failures radosgw-admin produces. Both the
ENOENT branch and the else branch then wrap that shadowed nil, and
errors.Wrapf(nil, ...) is nil, so the function returns
(waitForRequeueIfObjectStoreNotReady, nil). The caller only propagates the
requeue when the error is non-nil, so the requeue is dropped and reconcile
runs on.
The result is that a failed "zone get" no longer backs off. Reconcile
proceeds to stand up the object store anyway -- the RGW service, the
admin-ops endpoint, the deployment, and the pools radosgw scaffolds as it
comes up -- for a multisite store whose backing zone does not exist. This
is not the "normal multisite bootstrap" transient the original wording
suggested. getMultisiteResourceNames runs immediately before this and
already requeues until the CephObjectZone CR reports Ready, and the zone
controller marks it Ready only after it has created the Ceph zone, so a
healthy bootstrap never reaches this gate with a missing zone. That
CR-Ready gate, not this one, is what actually blocks bootstrap.
Where the dead gate does bite is the cases the CR status cannot cover:
- the Ceph zone deleted or renamed out of band while the CR still reads Ready
- a zone controller that reports Ready without leaving a usable zone behind
- any non-ENOENT "zone get" failure, e.g. a permission or connectivity
error, which the else branch swallows the same way
The check was correct when fc579f4520 introduced it in 2020 with
exec.ExitStatus, which returns (code, ok) and leaves err unshadowed.
bd58790c31 ("ceph: proxy ceph commands when multus is configured", 2021)
swapped it to exec.ExtractExitCode as an unrelated drive-by, inverting the
contract and killing the gate. The sibling realm, zonegroup, and zone
controllers were left untouched and still use exec.ExitStatus today.
Restore exec.ExitStatus so the outer error is no longer shadowed, both
branches wrap the real command error, and the caller requeues until the
zone exists -- matching the sibling controllers.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Changes:
- Use a dedicated .nvmeof pool (via CephBlockPool CR named
builtin-nvmeof) for the NVMe-oF gateway internal state.
- Use a separate nvmeof pool for the StorageClass data.
- Remove the pool field from the CephNVMeOFGateway CRD.
The gateway now always uses the .nvmeof pool, hardcoded
in the operator.
- Add .nvmeof to the allowed CephBlockPool name overrides.
- Create a production example nvmeof.yaml (instances: 2,
replicas: 3) and a CI-only nvmeof-test.yaml (instances: 1,
replicas: 1).
- Update documentation and CI test script accordingly.
Signed-off-by: Oded Viner <oviner@redhat.com>
startOSDMigration guards against upgrading OSDs while an OSD migration is
still in progress by returning the set of OSDs pending migration, which the
caller then removes from the update queue.
When the last migrated OSD's deployment has not been recreated yet,
isLastOSDMigrationComplete returns (false, nil). startOSDMigration then ran
return nil, errors.Wrapf(err, ...) with err already nil, and
errors.Wrapf(nil, ...) returns nil, so it returned (nil, nil), the same
value it uses to signal that no migration was requested. The caller
therefore skipped removing the pending OSDs from the update queue and could
upgrade and restart them mid-migration, which the pending-migration guard
was meant to prevent.
Return the freshly computed migration config in this case so the caller
still prunes the pending OSDs, without starting a new migration and without
aborting the reconcile. Aborting here would wedge the cluster because the
code that recreates the interrupted OSD runs later in Start, so an early
error return would leave that OSD deleted and block every later reconcile.
Letting the reconcile continue recreates the interrupted OSD and lets the
migration resume on a subsequent reconcile.
Also construct a non-nil error for the sibling unhealthy-PGs guard, which
wrapped a nil err in the same way.
Add a regression test asserting that startOSDMigration returns the pending
migration set, and does not delete another OSD deployment, while a previous
migration is still in progress.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
In createOrUpdateCephUser, when the desired user carries no explicit
keys the reconciler falls back to the live RGW user's keys. If that live
user also has zero keys the code intends to fail, but it built the error
with errors.Wrapf(err, ...) at a point where err is already nil (the
prior SetUserQuota error was handled and returned just above).
errors.Wrapf(nil, ...) returns nil, so the failure was swallowed and
createOrUpdateCephUser returned success on a user the operator itself
flagged as broken.
The reconcile then continued to generateCephUserSecret, which indexes
userConfig.Keys[0] to populate the Kubernetes secret and panicked on the
empty key slice. The reconciler's deferred RecoverAndLogException caught
and logged that panic, so the reconcile was abandoned before it reached
the Ready status update; and because the recovered Reconcile returns a
zero Result with a nil error, the request was not requeued either. The
user was left neither marked Ready nor retried.
Construct the error with errors.Errorf so the intended failure is
surfaced instead of being swallowed. Add a regression test that returns
a keyless live user and asserts a non-nil "no keys set" error.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
this commit add support for updating rgw caps with
`user-info-without-key` and `accounts` to cephobjectstoreuser crd
Adding the check for min ceph version from which these caps
are available.
Co-Authoured-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: subhamkrai <srai@redhat.com>
Add annotations to RBD node secrets in
import-external-cluster.sh so that externally
created secrets with
custom names can be discovered at runtime.
Update CreateUpdateClientProfileRadosNamespace to
look up secrets by annotation
instead of using hardcoded names.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
assignMons schedules each mon in its own goroutine and shares a
failedMonSchedule bool to record whether any of them failed. The
resultLock mutex already guards the shared c.mapping.Schedule map
write, but the three failedMonSchedule = true assignments were done
without holding the lock. When two or more mons fail to schedule at
the same time (for example when waitForMonitorScheduling errors, the
node choice is nil, or getNodeInfoFromNode fails) their goroutines
write the flag concurrently, which is a write-write data race.
Guard the failedMonSchedule writes with the existing resultLock, the
same lock the goroutines already use for the map update. The post-Wait
read is left as is since the WaitGroup orders it after every write.
Add a regression test that fails scheduling for multiple mons at once
and confirms assignMons returns an error; run it with -race to catch
the race.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
decodePeerToken logged the whole decoded PeerToken with "%+v" at debug
level. PeerToken embeds Key, the peer cluster's cephx auth secret, so
running the operator with ROOK_LOG_LEVEL=DEBUG wrote that credential to
the operator logs in plaintext (CWE-532).
Log only the non-sensitive fields (fsid, client id, mon host, namespace),
which keep the message useful for diagnostics, and drop Key. Add a
regression test that captures the debug output and asserts the key is
never present.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Port 1024 is the lowest non-privileged port on Linux (privileged ports
are 0-1023), so it should be honored as-is like any port >= 1024. The
dashboardInternalPort guard used `<=` instead of `<`, which pushed a
requested Port of 1024 into the "privileged" branch and reset the
container/targetPort back to the default (7000/8443). That left the
Service Port at 1024 but its targetPort at the default, a mismatch that
contradicts the constant's comment and the identical Port 1025 case.
Change the comparison to `<` and add a boundary case to
TestStartSecureDashboard asserting both Port and targetPort are 1024.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
currently there was a bug in the code where it didnt removed
the account from the ceph cluster during intial intialize
Signed-off-by: parth-gr <partharora1010@gmail.com>
testPodSpecPlacement accepts a pref parameter but never asserts the
preferred anti-affinity count, so the PreferredDuringScheduling branch
of SetNodeAntiAffinityForPod went unverified. The assertion was dropped
in b91f4211c9 when the separate preferred flag was removed from
SetNodeAntiAffinityForPod, leaving pref unused.
Restore the missing assertion so the helper validates both the required
and preferred term counts. Drop the third empty-placement case, which
became an exact duplicate of the second once the two-boolean API was
collapsed and now adds no coverage.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
The capacity-retention case in TestCephStatus asserted
formatTime(now-1m) against formatTime(now-1m), comparing an
expression to itself. That assertion can never fail and never
looks at the value under test, so the branch in
toCustomResourceStatus that keeps the previous capacity when
newStatus.PgMap.TotalBytes is 0 was left unverified.
Seed a known LastUpdated on the current status alongside the
existing TotalBytes, then assert the aggregated status carries
that timestamp through. This checks that the prior capacity,
including its LastUpdated, is retained when total bytes drop to 0.
The sibling cases already assert Capacity.LastUpdated the same way.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
The compareJSON test helper unmarshaled the expected JSON into both
the expected and actual variables, so assert.Equal always compared the
expected value against itself. The actualJSON argument was never read,
which made the assertion a tautology that could never fail and silently
disabled validation of updateCsiClusterConfig's output.
Parse actualJSON for the actual value. Enabling the check surfaced a
stale expectation in the multus sub-test: updateNetNamespaceFilePath
clears rbd.netNamespaceFilePath (holder pods are no longer used with the
ceph-csi-operator), so the expected literal is updated to drop that
field and match the real output.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Add annotations to CephFS and RBD provisioner secrets in
import-external-cluster.sh so that externally created secrets with
custom names can be discovered at runtime.
Update CreateUpdateClientProfileRadosNamespace and
generateProfileSubVolumeGroupSpec to look up secrets by annotation
instead of using hardcoded names.
closes: https://github.com/rook/rook/issues/16956
Signed-off-by: parth-gr <partharora1010@gmail.com>
CephFS reconciliation loop on Filesystem creation was swallowing errors.
The err variable which contained the error was replaced with the resulted error from the
kubernetes update status call on the CRD.
This commit fixes that and aggregate both errors in case there is both a kubernetes status update failure and a creation error.
Signed-off-by: Maxime Bertin <mbertin@luccasoftware.com>