Added failureDomains and osdsPerDomain fields to the
ErasureCodedSpec, which maps to the crush-num-failure-domains and
crush-osds-per-failure-domain Ceph EC profile parameters. This enables
creating EC pools that distribute chunks across fewer, larger hosts
without needing one host per data chunk.
Signed-off-by: Prabhala Tara Aasrita <taraprabhala@Prabhalas-MacBook-Pro.local>
The daemon only flagged a device change from in-use to empty, never the
reverse, so an in-use disk was never recorded as non-empty and a later
disk swap produced no ConfigMap update or reconcile. This is needed for
OSD replacement to work automatically: without it a destroyed OSD slot is
never reprovisioned after the disk is swapped. Compare the empty state in
both directions so the swap becomes a real transition that triggers the
reconcile.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
in case of multiple destroyed OSDs or multiple available
blank devices, Rook will try to match them by the same
device class if possible. If no matching DC, then DC of
destroyed OSD will be used
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
A state machine for osd destroy phase of osd-replacement flow.
The state machine is implemented in OSD health goroutine, where each
step is handled on a separate tick. State machine is responsible
for draining, destroying OSD, and reserving its CRUSH position,
downscalling its deployment, and adding "readyForSwap" annotation.
Destroy phase does not zap the device on purpose. On disk signature
is kept to avoid automatic reprovisioning of destroyed OSD and to
detect physical swap.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
defines a new entry point "close-encrypted-devices" in osd job.
The job is used by osd-replacement flow to close dm-crypt mappings
on host for destroyed encrypted OSDs. The job is owned by OSD
health goroutine implementing destroy phase of osd-replacement.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
contains osd prepare job edits to support osd replacement flow:
- removes destroyed OSD IDs from node status CM to stop cluster
controller from reprovisioning destroyed OSD
- adds provisioning logic for OSD replacement to match available
blank devices to destroyed OSD slots
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
the .nvmeof pool was incorrectly tagged with the rbd
application and initialized as an rbd pool. add .nvmeof
to the built-in pool switch so it gets the nvmeof-meta
application tag instead, and skip rbd pool init.
Signed-off-by: Oded Viner <oviner@redhat.com>
archiveCrash is meant to silence the RECENT_CRASH health warning for a
removed OSD by archiving its crash, but the early-return guard was
inverted. GetCrashList unmarshals `ceph crash ls` into a non-nil slice on
every success, so `if crash != nil` was always true. The function always
logged "no ceph crash to silence" and the archive loop was unreachable, so
OSD removal never archived the crash and the warning lingered. Guard on
`len(crash) == 0` so only a genuinely empty list short-circuits.
Reviving the loop surfaced two latent bugs it had always masked. On a
non-empty list with no entry for the removed OSD the crash id stayed empty
and the code ran `ceph crash archive ""`, which the mgr rejects with EINVAL
and logs as a spurious error during routine OSD removal, so skip the
archive when no crash matches. The loop also broke on the first match, but
`ceph crash ls` is oldest-first and includes already-archived entries, so
an OSD with several crashes had only its oldest one archived while
RECENT_CRASH persisted; archive every matching entry instead of just the
first.
Extend TestArchiveCrash to cover all four cases: a single matching crash, a
list with several matching crashes that are all archived, an empty list,
and a non-empty list with no matching entry where nothing is archived.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Ceph reports bytes_quota in `ceph fs subvolumegroup info` as the JSON
string "infinite" when no quota is set, and as a number otherwise. The
subvolumeGroupInfo.BytesQuota field is typed int64, so json.Unmarshal
fails with "cannot unmarshal string into Go value of type int64" for any
subvolume group without a quota (the default).
That unmarshal error is returned from getCephFSSubVolumeGroupInfo and
swallowed in CreateCephFSSubVolumeGroup: exec.ExitStatus reports ok=false
for the wrapped error, so the non-ENOENT guard does not fire and the
dedicated resize path (which enforces --no-shrink) is skipped.
Add an UnmarshalJSON on subvolumeGroupInfo that accepts either a number
or the "infinite" sentinel, mapping "infinite" to 0 (unlimited) so the
existing CmpInt64 comparison keeps working. Add a table-driven test
covering the finite and infinite cases.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Test_cleanCSIDirs created both the rbd and cephfs CSI driver
directories but only ever checked that the rbd directory was
removed. The second os.Stat call re-checked rbdDriverDir instead
of cephFSDriverDir, so removal of the cephfs CSI directory by
cleanCSIDirs was never verified even though the test set it up
for exactly that purpose.
Point the second assertion at cephFSDriverDir so both directories
that cleanCSIDirs is expected to remove are actually asserted.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
The success log in removeSnapshotSchedule passed its arguments in the
wrong order for the format string "successfully removed snapshot
schedule %q for pool %q". The arguments were poolName followed by
snapScheduleResponse.Interval, so the schedule slot printed the pool
name and the pool slot printed the interval.
Swap the two arguments so the interval prints in the schedule slot and
the pool name in the pool slot. This matches the sibling log in
enableSnapshotSchedule and the order used to build the remove command.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The //nolint directives at these sites omit the linter name, so each
suppresses every linter on its line rather than the one check it needs.
That hides any unrelated errcheck/gosec/govet finding later introduced
on the same line. Name the specific linter for each:
- staticcheck for the two operator sites: SA4004 (the intentional
single-iteration loop in the OSD PVC host lookup) and SA1019 (the
deliberate read of the deprecated S3.Enabled field in the RGW
API-enable builder).
- errcheck for the rbd-mirror deferred token-file cleanup and the
test-framework logging helpers (WriteString / writeHeader).
No behavior change.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The OSDPerfStats struct had no references anywhere in the repo (verified
with repo-wide grep and an exported-symbol audit); the code that consumed
this unmarshal target was removed long ago.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
NewSharedStorageAndWorkerNodesValidationTestConfig was a wrapper that
simply returned NewDefaultValidationTestConfig() and had no callers
anywhere in the repo (verified with deadcode and repo-wide grep, tests
included). The sibling constructors are all referenced by the multus
validation CLI and remain.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Remove four dead functions from pkg/daemon/ceph/client; none had any
callers anywhere in the repo (verified with deadcode, staticcheck U1000,
and repo-wide grep, tests included):
- auth.go: AuthGetOrCreate
- command.go: ExecuteRBDCommandWithTimeout
- filesystem.go: MarkFilesystemAsDown
- status.go: IsClusterCleanError (IsClusterClean itself remains in use)
No imports become unused by these removals.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
createREplicatedPoolForApp creates a base Crush rule (named after the
pool) on every reconcile, even if the pool already exists and uses a
suffixed rule ( eg pool_host_ssd). This causes orphaned rules to
reappear after every operator restart, immediately after
CleanupUnusedCrushRules deletes them.
This PR ensures that the base rule is created only if the pool does not
exist yet.
Signed-off-by: Santosh <sapillai@redhat.com>
this commit adds flag stripe_unit for the command
ceph osd erasure-code-profile set [flags]. This key
increase perfomance for ec pool. The default value
is 4kib/4096 bytes.
Signed-off-by: subhamkrai <srai@redhat.com>
remov restrictions that prevented LVM logical volume from being
provisioned as OSDs. LV devices with a metadatadevice were rejected at
the filtering stage. LV devices without one were incorrectly being
routed to ceph-volume raw mode which was resuting in error. This PR:
- removes metadatadevice restriction on LV devices.
- Adds LVM time check to isSafeToUseRAWMode to prevent raw mode on lvm
Signed-off-by: Santosh <sapillai@redhat.com>
During the initial cluster setup, CleanUnusedCrushRules() runs after MGR
startup but before any pools are created (if user does not create the
.mgr pool along with the cephCluster manifest). It deletes the default
replicated_pool since no pool references it.
This leaves the cluster in a state if Zero crush rules. When the MGR
tries to start the device_health_metris during this window, thne it
fails to find the crush rule and crashes.
Skip the cleanup when no pools exists yet, since there is nothing
meaninful to cleanup.
Signed-off-by: Santosh <sapillai@redhat.com>
Fixes#16834 and #17384. Per-OSD CR deviceClass was ignored in
osd prepare job when reporting the osd config back to the operator,
and in the operator's reconcile override.
Fixed by resolving the per-device class from the CR devices list
in GetCephVolumeRawOSDs and in makeDeployment.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
Clean up stale CRUSH rules after the Ceph mgr starts so rules left behind by pool failure-domain or device-class changes do not accumulate indefinitely.
The cleanup is guarded by a package-level RWMutex. Pool create and update paths hold the read lock while creating and assigning CRUSH rules, while cluster-wide cleanup holds the write lock before listing pools and deleting unused rules. This keeps pool reconciles parallel with each other while preventing cleanup from deleting a rule that another reconcile has just created but not yet attached to a pool.
Keep direct pool-delete cleanup for the pool's current CRUSH rule, make the cluster-wide cleanup best-effort across all unused rules, and add an operator-level ROOK_DELETE_UNUSED_CRUSH_RULES setting for clusters that need to leave unused custom rules in place.
Document the operator and Helm settings, regenerate the Helm chart docs, and add a pending release note for the default cleanup behavior.
Signed-off-by: Asish Kumar <officialasishkumar@gmail.com>
During OSD provisioning, the KMS error log incorrectly uses logger.Error() instead of logger.Errorf(), causing the format specifier %q to appear in the logs. Changed logger.Error to logger.Errorf to fix the same.
Signed-off-by: Prabhala Tara Aasrita <taraasrita@ibm.com>
Ceph incorrectly raises the MON_NETSPLIT health warning in stretch cluster configurations. Added a call to MuteHealthWarning() in postMonStartupActions() which runs `ceph health mute MON_NETSPLIT --sticky` to suppress the warning during reconcile. All errors are logged as warnings and the reconcile proceeds.
Signed-off-by: Prabhala Tara Aasrita <taraasrita@ibm.com>
Support using SSE-S3 encryption with RGW using vault Agent auth.
RGW sends requests to the agent instead of directly to Vault, and the agent transparently injects the authentication token. This eliminates using and managing a static token
Signed-off-by: Santosh <sapillai@redhat.com>