Commit Graph
959 Commits
Author SHA1 Message Date
Travis Nielsen d1a92df02b Merge pull request #17953 from cobaltcore-dev/osd-replacement-2-6
Osd replacement impl
2026-07-29 12:03:23 -06:00
Travis Nielsen 6ed8f3efcd Merge pull request #18028 from taraasrita10/support_EC_with_MSR
pool: support CRUSH MSR rules for EC pools
2026-07-28 13:45:02 -06:00
Prabhala Tara Aasrita 715f38de91 pool: support CRUSH MSR rules for EC pools
Added failureDomains and osdsPerDomain fields to the
ErasureCodedSpec, which maps to the crush-num-failure-domains and
crush-osds-per-failure-domain Ceph EC profile parameters. This enables
creating EC pools that distribute chunks across fewer, larger hosts
without needing one host per data chunk.

Signed-off-by: Prabhala Tara Aasrita <taraprabhala@Prabhalas-MacBook-Pro.local>
2026-07-28 20:23:51 +05:30
Artem Torubarov d18ff0da06 osd: detect in-use to empty device transition so a disk swap triggers a reconcile
The daemon only flagged a device change from in-use to empty, never the
reverse, so an in-use disk was never recorded as non-empty and a later
disk swap produced no ConfigMap update or reconcile. This is needed for
OSD replacement to work automatically: without it a destroyed OSD slot is
never reprovisioned after the disk is swapped. Compare the empty state in
both directions so the swap becomes a real transition that triggers the
reconcile.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:44:01 +02:00
Artem Torubarov 0eb4735ee3 osd: device class matching for osd replacement
in case of multiple destroyed OSDs or multiple available
blank devices, Rook will try to match them by the same
device class if possible. If no matching DC, then DC of
destroyed OSD will be used

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov 11ffd3b717 osd: use helper to check command status code
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov a0d215b428 osd: implement osd replacement health goroutine
A state machine for osd destroy phase of osd-replacement flow.
The state machine is implemented in OSD health goroutine, where each
step is handled on a separate tick. State machine is responsible
for draining, destroying OSD, and reserving its CRUSH position,
downscalling its deployment, and adding "readyForSwap" annotation.
Destroy phase does not zap the device on purpose. On disk signature
is kept to avoid automatic reprovisioning of destroyed OSD and to
detect physical swap.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov c4727a0b21 osd: implement osd replacement crypto close job
defines a new entry point "close-encrypted-devices" in osd job.
The job is used by osd-replacement flow to close dm-crypt mappings
on host for destroyed encrypted OSDs. The job is owned by OSD
health goroutine implementing destroy phase of osd-replacement.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov 887d533001 osd: implement osd replacement prepare-job
contains osd prepare job edits to support osd replacement flow:
- removes destroyed OSD IDs from node status CM to stop cluster
  controller from reprovisioning destroyed OSD
- adds provisioning logic for OSD replacement to match available
  blank devices to destroyed OSD slots

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Oded Viner f095a530c8 nvmeof: set nvmeof-meta application tag on .nvmeof pool
the .nvmeof pool was incorrectly tagged with the rbd
application and initialized as an rbd pool. add .nvmeof
to the built-in pool switch so it gets the nvmeof-meta
application tag instead, and skip rbd pool init.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-07-24 18:11:37 +03:00
Chiman Jain 2034be1f37 build: migrate gopkg.in/yaml to go.yaml.in/yaml
Signed-off-by: Chiman Jain <chimanjain15@gmail.com>
2026-07-22 20:05:39 +05:30
Joshua Hoblitt 9312bbb350 Merge pull request #17980 from anxkhn/patch-10
core: parse subvolumegroup bytes_quota reported as "infinite"
2026-07-21 09:26:28 -07:00
Anas Khan 1630ad1049 osd: archive crash on OSD removal
archiveCrash is meant to silence the RECENT_CRASH health warning for a
removed OSD by archiving its crash, but the early-return guard was
inverted. GetCrashList unmarshals `ceph crash ls` into a non-nil slice on
every success, so `if crash != nil` was always true. The function always
logged "no ceph crash to silence" and the archive loop was unreachable, so
OSD removal never archived the crash and the warning lingered. Guard on
`len(crash) == 0` so only a genuinely empty list short-circuits.

Reviving the loop surfaced two latent bugs it had always masked. On a
non-empty list with no entry for the removed OSD the crash id stayed empty
and the code ran `ceph crash archive ""`, which the mgr rejects with EINVAL
and logs as a spurious error during routine OSD removal, so skip the
archive when no crash matches. The loop also broke on the first match, but
`ceph crash ls` is oldest-first and includes already-archived entries, so
an OSD with several crashes had only its oldest one archived while
RECENT_CRASH persisted; archive every matching entry instead of just the
first.

Extend TestArchiveCrash to cover all four cases: a single matching crash, a
list with several matching crashes that are all archived, an empty list,
and a non-empty list with no matching entry where nothing is archived.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-20 17:45:02 -07:00
Joshua Hoblitt 7fc7c33a0e Merge pull request #17956 from anxkhn/patch-15
test: assert cephfs CSI directory removal in cleanCSIDirs test
2026-07-20 14:47:36 -07:00
Anas Khan 7ee65e13be core: parse subvolumegroup bytes_quota reported as "infinite"
Ceph reports bytes_quota in `ceph fs subvolumegroup info` as the JSON
string "infinite" when no quota is set, and as a number otherwise. The
subvolumeGroupInfo.BytesQuota field is typed int64, so json.Unmarshal
fails with "cannot unmarshal string into Go value of type int64" for any
subvolume group without a quota (the default).

That unmarshal error is returned from getCephFSSubVolumeGroupInfo and
swallowed in CreateCephFSSubVolumeGroup: exec.ExitStatus reports ok=false
for the wrapped error, so the non-ENOENT guard does not fire and the
dedicated resize path (which enforces --no-shrink) is skipped.

Add an UnmarshalJSON on subvolumeGroupInfo that accepts either a number
or the "infinite" sentinel, mapping "infinite" to 0 (unlimited) so the
existing CmpInt64 comparison keeps working. Add a table-driven test
covering the finite and infinite cases.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
2026-07-15 11:34:57 +05:30
Anas Khan 130b868d56 test: assert cephfs CSI directory removal in cleanCSIDirs test
Test_cleanCSIDirs created both the rbd and cephfs CSI driver
directories but only ever checked that the rbd directory was
removed. The second os.Stat call re-checked rbdDriverDir instead
of cephFSDriverDir, so removal of the cephfs CSI directory by
cleanCSIDirs was never verified even though the test set it up
for exactly that purpose.

Point the second assertion at cephFSDriverDir so both directories
that cleanCSIDirs is expected to remove are actually asserted.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
2026-07-14 14:22:00 +05:30
Anas Khan 9e88101c72 rbdmirror: fix swapped log args on snapshot schedule removal
The success log in removeSnapshotSchedule passed its arguments in the
wrong order for the format string "successfully removed snapshot
schedule %q for pool %q". The arguments were poolName followed by
snapScheduleResponse.Interval, so the schedule slot printed the pool
name and the pool slot printed the interval.

Swap the two arguments so the interval prints in the schedule slot and
the pool name in the pool slot. This matches the sibling log in
enableSnapshotSchedule and the order used to build the remove command.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
2026-07-14 14:18:01 +05:30
Artem Torubarov 75fe8db848 osd: rename internal OSD migration field to migrateOSD
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-10 11:45:47 +02:00
weifanglab 4d2705ce28 core: replace Split in loops with more efficient SplitSeq
Signed-off-by: weifanglab <weifanglab@outlook.com>
2026-07-07 00:41:15 +08:00
Joshua Hoblitt c0eb360041 docs: fix typos and grammar in code comments
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:34:11 -07:00
Joshua Hoblitt e5c6d75a73 core: narrow bare //nolint directives to specific linters
The //nolint directives at these sites omit the linter name, so each
suppresses every linter on its line rather than the one check it needs.
That hides any unrelated errcheck/gosec/govet finding later introduced
on the same line. Name the specific linter for each:

- staticcheck for the two operator sites: SA4004 (the intentional
  single-iteration loop in the OSD PVC host lookup) and SA1019 (the
  deliberate read of the deprecated S3.Enabled field in the RGW
  API-enable builder).
- errcheck for the rbd-mirror deferred token-file cleanup and the
  test-framework logging helpers (WriteString / writeHeader).

No behavior change.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 16:32:59 -07:00
Joshua Hoblitt 993be34a24 Merge pull request #17834 from jhoblitt/docs/fix-stale-doc-comment-names
docs: fix function comments to match their declaration names
2026-06-29 12:04:25 -07:00
Blaine Gardner 485529056d Merge pull request #17764 from subhamkrai/mute-ceph-errors
core: add api changes to mute ceph warnings
2026-06-29 09:50:13 -06:00
Joshua Hoblitt 49612461a4 docs: fix function comments to match their declaration names
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-26 12:52:39 -07:00
pengqima 240c431f42 docs: fix function comment to match actual function name
Signed-off-by: pengqima <pengqima@outlook.com>
2026-06-26 23:43:22 +08:00
daixihegu 063dcad054 core: use slices.Contains to simplify code
Signed-off-by: daixihegu <daixihegu@163.com>
2026-06-25 10:17:53 +08:00
subhamkrai 9a6c5842bc core: implement ceph health warning mute
this commit implement api to configure
ceph health warning mute/unmute.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-06-24 21:10:54 +05:30
Joshua Hoblitt ef9761d4df core: remove unused OSDPerfStats type
The OSDPerfStats struct had no references anywhere in the repo (verified
with repo-wide grep and an exported-symbol audit); the code that consumed
this unmarshal target was removed long ago.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 15:40:49 -07:00
Joshua Hoblitt 2433565238 multus: remove unused validation test config constructor
NewSharedStorageAndWorkerNodesValidationTestConfig was a wrapper that
simply returned NewDefaultValidationTestConfig() and had no callers
anywhere in the repo (verified with deadcode and repo-wide grep, tests
included). The sibling constructors are all referenced by the multus
validation CLI and remain.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 15:40:10 -07:00
Joshua Hoblitt 278453e10c core: remove unused ceph client helper functions
Remove four dead functions from pkg/daemon/ceph/client; none had any
callers anywhere in the repo (verified with deadcode, staticcheck U1000,
and repo-wide grep, tests included):

- auth.go: AuthGetOrCreate
- command.go: ExecuteRBDCommandWithTimeout
- filesystem.go: MarkFilesystemAsDown
- status.go: IsClusterCleanError (IsClusterClean itself remains in use)

No imports become unused by these removals.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-09 15:06:17 -07:00
Satoru Takeuchi 4105874d9a Merge pull request #17525 from sp98/lv-osd-with-metadata
osd: allow LVM logical volumes as OSD data devices
2026-05-22 19:44:00 +09:00
Blaine Gardner 645aa741eb Merge pull request #17530 from rook/maint-rm-unused-consts-opus-4.7
core: rm unused consts
2026-05-20 09:01:05 -06:00
Travis Nielsen d66f1e3465 Merge pull request #17547 from Nordix/Sunnatillo/fix-go1.26-printf-check
build: fix gosec and go vet lint errors for Go 1.26
2026-05-15 13:25:41 -06:00
Sunnatillo 0bd299aa60 build: fix gosec and go vet lint errors for Go 1.26
Signed-off-by: Sunnatillo <sunnat.samadov@est.tech>
2026-05-15 21:45:06 +03:00
Santosh Pillai fda5ab6a63 Merge pull request #17535 from sp98/create-base-rule-only-once
ceph: skip base rule creation if pools exit
2026-05-15 19:44:08 +05:30
Santosh df0f4e6c99 pool: skip base rule creation of pools exit
createREplicatedPoolForApp creates a base Crush rule (named after the
pool) on every reconcile, even if the pool already exists and uses a
suffixed rule ( eg pool_host_ssd). This causes orphaned rules to
reappear after every operator restart, immediately after
CleanupUnusedCrushRules deletes them.

This PR ensures that the base rule is created only if the pool does not
exist yet.

Signed-off-by: Santosh <sapillai@redhat.com>
2026-05-15 12:39:10 +05:30
Travis Nielsen 3775bf206a Merge pull request #17407 from cobaltcore-dev/osd-deviceclass-minimal
osd: honor per-device deviceClass in raw-mode prepare and reconcile
2026-05-14 12:04:16 -06:00
Travis Nielsen 2ef74f3189 Merge pull request #17447 from subhamkrai/add-stipe-unit-support
pool: add stripe_unit with erasure-code-profile set
2026-05-14 10:38:06 -06:00
Joshua Hoblitt 9eec9c730e core: rm unused consts
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-05-14 09:34:52 -07:00
subhamkrai 9fb0582dd6 pool: add stripe_unit with erasure-code-profile set
this commit adds flag stripe_unit for the command
ceph osd erasure-code-profile set [flags]. This key
increase perfomance for ec pool. The default value
is 4kib/4096 bytes.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-05-14 21:39:15 +05:30
Santosh c6d48178dd osd: allow LVM logical volumes as OSD data devices
remov restrictions that prevented LVM logical volume from being
provisioned as OSDs. LV devices with a metadatadevice were rejected at
the filtering stage. LV devices without one were incorrectly being
routed to ceph-volume raw mode which was resuting in error. This PR:

- removes metadatadevice restriction on LV devices.
- Adds LVM time check to isSafeToUseRAWMode to prevent raw mode on lvm

Signed-off-by: Santosh <sapillai@redhat.com>
2026-05-14 11:43:45 +05:30
Travis Nielsen 2898acddfc Merge pull request #17494 from immanuwell/docs-fix-duplicated-words
docs: fix duplicated words in docs and comments
2026-05-12 12:41:18 -10:00
Santosh 1c1d6cf261 core: skip crush rule cleanup if no pools exist yet
During the initial cluster setup, CleanUnusedCrushRules() runs after MGR
startup but before any pools are created (if user does not create the
.mgr pool along with the cephCluster manifest). It deletes the default
replicated_pool since no pool references it.

This leaves the cluster in a state if Zero crush rules. When the MGR
tries to start the device_health_metris during this window, thne it
fails to find the crush rule and crashes.

Skip the cleanup when no pools exists yet, since there is nothing
meaninful to cleanup.

Signed-off-by: Santosh <sapillai@redhat.com>
2026-05-12 11:45:04 +05:30
Artem Torubarov b41d63788d osd: honor per-device deviceClass in raw-mode prepare and reconcile
Fixes #16834 and #17384. Per-OSD CR deviceClass was ignored in
osd prepare job when reporting the osd config back to the operator,
and in the operator's reconcile override.
Fixed by resolving the per-device class from the CR devices list
in GetCephVolumeRawOSDs and in makeDeployment.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-05-11 15:37:16 +02:00
immanuwell c3a9f5fdd9 docs: fix duplicated words in docs and comments
Clean up duplicated words and a typo in object storage docs, comments, and user-facing messages.

Signed-off-by: immanuwell <pchpr.00@list.ru>
2026-05-09 12:50:31 +04:00
Asish Kumar 24a35e4e74 pool: clean up unused crush rules
Clean up stale CRUSH rules after the Ceph mgr starts so rules left behind by pool failure-domain or device-class changes do not accumulate indefinitely.

The cleanup is guarded by a package-level RWMutex. Pool create and update paths hold the read lock while creating and assigning CRUSH rules, while cluster-wide cleanup holds the write lock before listing pools and deleting unused rules. This keeps pool reconciles parallel with each other while preventing cleanup from deleting a rule that another reconcile has just created but not yet attached to a pool.

Keep direct pool-delete cleanup for the pool's current CRUSH rule, make the cluster-wide cleanup best-effort across all unused rules, and add an operator-level ROOK_DELETE_UNUSED_CRUSH_RULES setting for clusters that need to leave unused custom rules in place.

Document the operator and Helm settings, regenerate the Helm chart docs, and add a pending release note for the default cleanup behavior.

Signed-off-by: Asish Kumar <officialasishkumar@gmail.com>
2026-04-24 23:57:40 +05:30
Prabhala Tara Aasrita 85ee8631cf osd: fix KMS log error format specifier
During OSD provisioning, the KMS error log incorrectly uses logger.Error() instead of logger.Errorf(), causing the format specifier %q to appear in the logs. Changed logger.Error to logger.Errorf to fix the same.

Signed-off-by: Prabhala Tara Aasrita <taraasrita@ibm.com>
2026-04-23 16:19:33 +05:30
Denis Egorenko b3ffac9e0d mds: fix incorrect behaviour for CephFS health state when no active standby
Set expected standby mds daemons, when changing activeStandby mds parameter.

Resolves #17372

Signed-Off-By: Denis Egorenko <degorenko@mirantis.com>
2026-04-20 13:09:41 +04:00
Prabhala Tara Aasrita eccc8128af mon: mute MON_NETSPLIT health warning in stretch clusters
Ceph incorrectly raises the MON_NETSPLIT health warning in stretch cluster configurations. Added a call to MuteHealthWarning() in postMonStartupActions() which runs `ceph health mute MON_NETSPLIT --sticky` to suppress the warning during reconcile. All errors are logged as warnings and the reconcile proceeds.

Signed-off-by: Prabhala Tara Aasrita <taraasrita@ibm.com>
2026-04-17 19:29:03 +05:30
Santosh c6836c474a rgw: support SSE-S3 with vault agent
Support using SSE-S3 encryption with RGW using vault Agent auth.
RGW sends requests to the agent instead of directly to Vault, and the agent transparently injects the authentication token. This eliminates using and managing a static token

Signed-off-by: Santosh <sapillai@redhat.com>
2026-04-07 12:18:29 +05:30