Commit Graph
4037 Commits
Author SHA1 Message Date
raaizik 8e52af5cfb nfs: add unit tests for custom NFS port
Add coverage for GetPort defaults and custom values, Service and probe
port wiring, Ganesha NFS_Port config, and create-or-update when
spec.server.port changes on an existing Service.

Signed-off-by: raaizik <raaizik@yahoo.com>
(cherry picked from commit b97296c47c)
2026-07-28 14:50:04 +00:00
raaizik 7c4787691d nfs: support custom NFS server port
When NFS runs with host networking, port 2049 may already be in use.
Add CephNFS spec.server.port (default 2049), wire it through Ganesha
config, the operator Service, and probes, and regenerate CRDs.

Signed-off-by: raaizik <raaizik@yahoo.com>
(cherry picked from commit 5e36fa4bc3)
2026-07-28 14:50:04 +00:00
Travis Nielsen 34c8545a9e Merge pull request #18031 from rook/mergify/bp/release-1.20/pr-17785
fix(ceph): deduplicate external MGR endpoints in EndpointSlice (backport #17785)
2026-07-27 15:09:29 -06:00
Travis Nielsen a96bbe74e0 Merge pull request #18004 from rook/mergify/bp/release-1.20/pr-17962
rgw: rgw account deletion should remove account from ceph (backport #17962)
2026-07-22 09:27:35 -06:00
beyondcloud-co 66c9171785 external: deduplicate external MGR endpoints in EndpointSlice
When the active MGR IP already exists at a non-zero position in the
externalMgrEndpoints list, the rotation logic that replaces endpoints[0]
creates duplicate IPs. Also guards against mgr map failure which could
overwrite endpoints[0] with an empty string.

Fixes: #17761

Signed-off-by: beyondcloud-co <87993136+beyondcloud-co@users.noreply.github.com>
(cherry picked from commit 8134d683e3)
2026-07-22 15:06:14 +00:00
Michael Adam 2d05950963 rgw: fix a comment typo
This fixes a comment typo caught bt the codespell CI check.

Signed-off-by: Michael Adam <obnox@samba.org>
(cherry picked from commit c90913051b)
2026-07-21 15:08:20 -06:00
mergify[bot] 9b826b29a4 Merge pull request #18016 from rook/mergify/bp/release-1.20/pr-17968
mon: fix data race on failedMonSchedule during scheduling (backport #17968)
2026-07-21 16:31:16 +00:00
Anas Khan f79696157d mon: fix data race on failedMonSchedule during scheduling
assignMons schedules each mon in its own goroutine and shares a
failedMonSchedule bool to record whether any of them failed. The
resultLock mutex already guards the shared c.mapping.Schedule map
write, but the three failedMonSchedule = true assignments were done
without holding the lock. When two or more mons fail to schedule at
the same time (for example when waitForMonitorScheduling errors, the
node choice is nil, or getNodeInfoFromNode fails) their goroutines
write the flag concurrently, which is a write-write data race.

Guard the failedMonSchedule writes with the existing resultLock, the
same lock the goroutines already use for the map update. The post-Wait
read is left as is since the WaitGroup orders it after every write.

Add a regression test that fails scheduling for multiple mons at once
and confirms assignMons returns an error; run it with -race to catch
the race.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
(cherry picked from commit 0bbd5de84c)
2026-07-21 16:25:00 +00:00
Anas Khan 245c45d536 csi: do not log peer bootstrap token secret
decodePeerToken logged the whole decoded PeerToken with "%+v" at debug
level. PeerToken embeds Key, the peer cluster's cephx auth secret, so
running the operator with ROOK_LOG_LEVEL=DEBUG wrote that credential to
the operator logs in plaintext (CWE-532).

Log only the non-sensitive fields (fsid, client id, mon host, namespace),
which keep the message useful for diagnostics, and drop Key. Add a
regression test that captures the debug output and asserts the key is
never present.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
(cherry picked from commit 1013a8effc)
2026-07-21 00:24:38 +00:00
parth-gr 3010491f03 rgw: rgw account deletion should remove account from ceph
currently there was a bug in the code where it didnt removed
the account from the ceph cluster during intial intialize

Signed-off-by: parth-gr <partharora1010@gmail.com>
(cherry picked from commit cca7020f0b)
2026-07-20 20:12:38 +00:00
Maxime Bertin 084e02a287 file: reveal swallowed errors on cephfs reconciliation
CephFS reconciliation loop on Filesystem creation was swallowing errors.
The err variable which contained the error was replaced with the resulted error from the
kubernetes update status call on the CRD.
This commit fixes that and aggregate both errors in case there is both a kubernetes status update failure and a creation error.

Signed-off-by: Maxime Bertin <mbertin@luccasoftware.com>
(cherry picked from commit 0696b24d9c)
2026-07-10 17:23:17 +00:00
subhamkrai 55526a4e98 mon: update how hostname was set for drbd
currently, drbd hostname was ready using command
but now the change we use kubernetes built-in api
to read hostname from `spec.nodeName`.`

Signed-off-by: subhamkrai <srai@redhat.com>
(cherry picked from commit c49ef5f079)
2026-07-09 15:38:16 +00:00
Anas Khan f49b96f08e core: reconcile when node OSD topology labels change
shouldReconcileChangedNode only diffed the node spec, so changes to node
labels, including rook's own topology.rook.io/* labels, were never detected.
Relabeling a node therefore did not trigger a CephCluster reconcile until some
other condition eventually did.

Trigger a reconcile from the node update predicate when one of the OSD topology
labels rook cares about changes. The whitelist reuses the canonical label set
from topology.GetDefaultTopologyLabels() (the hostname, the Kubernetes
region/zone labels, and the topology.rook.io/* labels), so a reconcile is
scoped to topology changes instead of firing on any label or annotation change.
The UpdateFunc returns true directly for these changes because onK8sNode()
returns false for a node that is already an OSD host, which would otherwise
swallow the relabel.

Resolves #17652

Signed-off-by: Anas Khan <anas@anaskhan.me>
(cherry picked from commit 47b4167b70)
2026-07-07 14:01:58 +00:00
Joshua Hoblitt ac686a02c8 docs: fix typos and grammar in code comments
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit c0eb360041)
2026-07-01 09:23:27 -07:00
Joshua Hoblitt dcd4dfb2bd docs: fix function comments to match their declaration names
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit 49612461a4)
2026-06-29 19:04:45 +00:00
subhamkrai 87e7082d48 core: implement ceph health warning mute
this commit implement api to configure
ceph health warning mute/unmute.

Signed-off-by: subhamkrai <srai@redhat.com>
(cherry picked from commit 9a6c5842bc)
2026-06-29 15:50:37 +00:00
kavirakesh14 8e0e66d324 operator: fail reconcile if rook-config-override lacks trailing newline
Signed-off-by: kavirakesh14 <kavirakesh007@gmail.com>
(cherry picked from commit 619b800b67)
2026-06-24 16:49:17 +00:00
mergify[bot] c6247f44a5 Merge pull request #17816 from rook/mergify/bp/release-1.20/pr-17812
core: Skip node watcher during initial cache sync (backport #17812)
2026-06-23 23:56:40 +00:00
Travis Nielsen d628fc8229 core: skip node watcher during initial cache sync
When the operator starts, there is no need to handle the
node watcher immediately while the controller runtime
cache is being initialized. The cluster will anyway be
fully reconciled and create new OSDs if needed. The node
watcher is only intended for triggering a reconcile when
a node is added or updated at a later time after a sync
would anyway be triggered.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 7803fac10d)
2026-06-23 23:34:25 +00:00
parth-gr de5ff42de7 core: reduce the floating mon reschdule time
currently k8s add the default eviction time as 300 sec
setting it to 5 sec for the quick re-schedule

Signed-off-by: parth-gr <partharora1010@gmail.com>
(cherry picked from commit 9e0b6e5bbe)
2026-06-23 15:26:08 +00:00
Joshua Hoblitt 025a76fe57 osd: do not scan rbd devices in raw activate fallback
When the OSD activate init container cannot match its configured
device, it falls back to a bare 'ceph-volume raw list', which
enumerates every block device on the node and opens each one to read
bluestore labels. Opening a krbd device whose backing cluster is
unhealthy can block in uninterruptible sleep; when that device's I/O
depends on the OSD being activated, activation deadlocks.

Build the fallback scan list explicitly with lsblk and exclude the
network-backed device types that can hang when their backing storage
is unavailable (rbd, nbd, drbd) along with zram (volatile RAM). Rook
never provisions OSDs on any of these, so the fallback can never
legitimately need to find an OSD on them. loop devices are kept,
since OSDs on loop devices are supported for CI and local testing.
Fail fast if the filtered list is empty rather than degenerating back
to a full scan.

Note this hardens only the rename-fallback path. The deadlock observed
in CI (kernel hung tasks under 'ceph-volume raw activate' opening a
stalled /dev/rbd0; actions runs 27325963851, 27333888237, and
27378983111, which reproduced it on this PR's own head) ultimately
originates inside ceph-volume: 'raw activate' runs an unconditional
LUKS2/TPM2 discovery sweep over all block devices even when given an
explicit --device. That requires a ceph-volume fix; this change makes
the Rook-controlled scan safe regardless of ceph version.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit da07f0c476)
2026-06-22 22:44:18 +00:00
subhamkrai e2600812e5 build: update ceph version to latest v20.2.2
updating the ceph version from v20.2.1 to v20.2.2

Signed-off-by: subhamkrai <srai@redhat.com>
(cherry picked from commit ad26ebdab0)
2026-06-17 17:01:15 +00:00
Joshua Hoblitt e64c06f0b4 object: generate URL-safe realm access keys
The realm system user's access and secret keys are generated by
GeneratePassword() and then wrapped in base64. GeneratePassword()
deliberately excludes '/' from the access key character set, but the
base64.StdEncoding wrap reintroduces it: its alphabet contains '/' and
'+', and the 14-character input always produces trailing '='. The
encoded string is the literal key used in S3 requests.

An access key containing '/' breaks AWS SigV4 credential scope parsing
("<access-key>/<date>/<region>/<service>/aws4_request" is split on
'/'), so "radosgw-admin realm pull" against the realm endpoint fails
permanently with "request failed: (22) Invalid argument" (HTTP 400),
and a CephObjectRealm pulling that realm can never reconcile.

This is the dominant cause of the "deploy second cluster rook"
failures in the rgw-multisite-testing canary job: every sampled failure
had a generated access key containing '/' and looped on EINVAL for the
whole 600s wait window, while runs with slash-free keys pulled the
realm successfully.

Encode both keys with base64.RawURLEncoding instead, whose alphabet
(A-Za-z0-9-_, unpadded) is safe in credential scopes, URLs, and shell
arguments. Only newly created realm secrets are affected; existing
secrets are not modified by the reconciler.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit 573e2f81e7)
2026-06-16 21:16:39 +00:00
Joshua Hoblitt b2d0df03b1 k8sutil: remove unused helper functions
Remove dead code from the k8sutil package; none of these symbols had any
callers anywhere in the repo (verified with deadcode, staticcheck U1000,
and repo-wide grep, tests included):

- patcher.go: delete file (Patcher type, NewPatcher, Patch)
- replicaset.go: delete file (DeleteReplicaSet)
- customresource.go: WatchCR (the CustomResource type remains in use)
- deployment.go: WaitForDeploymentImage
- job.go: WaitForJobCompletion
- k8sutil.go: GetK8SVersion
- pod.go: ConfigOverrideMount, ConfigOverrideVolume, GetPodLog
- prometheus.go: DeleteServiceMonitor
- service.go: DeleteService
- volume.go: NodeConfigURI, BinariesMountInfo, YamlToVolumes,
  YamlToVolumeMounts, and the BinariesMountPath const (only used by the
  removed BinariesMountInfo)

Now-unused imports are dropped from each file. The shared
deleteResourceAndWait helper remains (still used by DeleteDeployment).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit b29e89ef67)
2026-06-12 16:10:49 +00:00
subham rai 3afc8f2adb Merge pull request #17721 from rook/mergify/bp/release-1.20/pr-17701
nvmeof: fetch gateway image from ceph config when not in CR (backport #17701)
2026-06-12 12:29:57 +05:30
Joshua Hoblitt 9f867c9b34 k8sutil: remove unused daemonset helpers
daemonset.go was entirely dead: CreateDaemonSet, DeleteDaemonset,
AddRookVersionLabelToDaemonSet, and GetDaemonsets had no callers anywhere
in the repo (verified with grep, deadcode, and staticcheck). Rook does not
manage any DaemonSets via these helpers. Remove the file.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit c8ad81a3cf)
2026-06-11 21:21:46 +00:00
Oded Viner 2a2dc44cb7 nvmeof: fetch gateway image from ceph config when not in CR
Make the NVMe-oF gateway image optional in the CR spec.
When not specified, the operator fetches the image from
the Ceph mon config store using the key
mgr/cephadm/container_image_nvmeof. Falls back to the
hardcoded default if the config is unavailable.

Signed-off-by: Oded Viner <oviner@redhat.com>
(cherry picked from commit 8b4fd67e17)
2026-06-11 18:03:36 +00:00
Joshua Hoblitt 15dbb2aec7 osd: remove unused updateQueue.Clear method
The Clear method on updateQueue had no callers anywhere in the repo
(verified with grep, deadcode, and staticcheck). The queue is emptied via
Remove/Pop in practice. Remove the dead method.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit 50c95d343e)
2026-06-10 21:20:24 +00:00
Joshua Hoblitt 9cc9f52f56 Merge pull request #17703 from rook/mergify/bp/release-1.20/pr-17700
core: delete stale MDS and RGW pdbs (backport #17700)
2026-06-10 14:03:25 -07:00
Joshua Hoblitt 57d21ff406 Merge pull request #17697 from rook/mergify/bp/release-1.20/pr-17671
test: remove unused container test helper (backport #17671)
2026-06-10 10:33:37 -07:00
Santosh d028a7a836 core: delete stale MDS and RGW pdbs
When RGW and MDS intances are reduce to 1, the older RGW and MDS PDB
instance (with maxUnavailabe=1) are still available. This blocks node
drain when the node, on which the single RGW and MDS instances are
running, is drained.

This PR deletes any stale MDS and RGW PDB when the only 1 MDS and 1 RGW
instances are available.

Signed-off-by: Santosh <sapillai@redhat.com>
(cherry picked from commit 3396198e9c)
2026-06-10 16:55:27 +00:00
Joshua Hoblitt ea244dc3c8 test: remove unused label verification helpers
VerifyPodLabels had no callers anywhere in the repo, and it was the only
caller of VerifyAppLabels, which in turn was the only caller of the
checkLabel and combineErrors helpers. Remove the whole dead cluster (and
the now-unused strings, github.com/pkg/errors, and ceph/controller
imports). AssertLabelsContainCephRequirements in the same file remains and
is unaffected. Verified with grep, deadcode, and staticcheck.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit ac75488621)
2026-06-09 18:28:31 +00:00
Joshua Hoblitt 71be697ace test: remove unused container test helper
container.go in the ceph operator test-helper package was entirely dead:
ContainerTestDefinition, its TestContainer method, and the logCommandWithArgs
helper had no references anywhere in the repo (no test or non-test callers,
verified with grep, deadcode, and staticcheck). The package doc comment is
also present on info.go, so removing this file leaves the package documented.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit 9efb7f351f)
2026-06-09 18:27:28 +00:00
Joshua Hoblitt 2ce517a0f3 Merge pull request #17694 from rook/mergify/bp/release-1.20/pr-17669
mgr: remove unused IsModuleInSpec function (backport #17669)
2026-06-09 10:56:01 -07:00
Joshua Hoblitt 1114be7d33 Merge pull request #17693 from rook/mergify/bp/release-1.20/pr-17648
object: remove dead functions and ObjectBuckets sort type (backport #17648)
2026-06-09 09:47:48 -07:00
Joshua Hoblitt f8a15d7401 mgr: remove unused IsModuleInSpec function
IsModuleInSpec has no callers anywhere in the repository (verified with
deadcode and grep). Remove the dead code.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit f8c0d36a50)
2026-06-09 15:58:23 +00:00
Joshua Hoblitt ad4d9927e8 object: remove dead functions and ObjectBuckets sort type
Remove the following unreferenced symbols:
- (*S3Agent).CreateBucketNoInfoLogging
- (*S3Agent).DeleteBucket (wrapper; callers use the SDK client directly)
- GetBucketsStats
- ObjectBuckets type and its Len/Less/Swap sort.Interface methods

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit ae481f0f43)
2026-06-09 15:58:16 +00:00
Joshua Hoblitt 70614ae5ff object: remove dead EjectPrincipals methods and duplicate AllowedActions entry
Remove (*BucketPolicy).EjectPrincipals and (*PolicyStatement).EjectPrincipals,
which had no callers outside policy.go. Also drop the duplicate PutBucketVersioning
entry in AllowedActions.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit df30bf66e8)
2026-06-09 15:58:03 +00:00
Yuxiang He ea2d2b4c50 osd: support using node labels for OSD device class assignment
Add support for the osd.rook.io/device-class node label to assign
CRUSH device classes to all OSDs on a node.
The operator reads the label during OSD provisioning and
deployment reconciliation. If CR-level deviceClass conflicts
with the node label, the operator will skip reconciling the
node and logs an error.

Resolves: #17131
Signed-off-by: Yuxiang He <yuxiang@dropbox.com>
(cherry picked from commit d9357dbfdc)
2026-06-08 22:07:21 +00:00
Sunnatillo 580663e917 osd: set require-osd-release from the health monitor
applyUpgradeOSDFunctionality() may run before all OSDs have registered
their new version, leaving require-osd-release unset and the cluster in
HEALTH_WARN.

Check for a single converged OSD version on each health-monitor tick and
set require-osd-release when it changes. The release name is cached to
avoid redundant ceph commands

Signed-off-by: Sunnatillo <sunnat.samadov@est.tech>
(cherry picked from commit f6fa25e844)
2026-06-04 15:31:21 +00:00
Artem Muterko 3ce302a218 object: clobber bucket policy on modify instead of merging
ModifyBucketPolicy merged the caller's statement into the policy fetched
from the bucket, matching by SID. Besides a missing match flag that
appended a duplicate when a SID matched, this preserved any pre-existing
statements on the bucket and reapplied them on every reconcile.

Overwrite the policy with the provided statements instead, so a managed
bucket always ends up with exactly the intended policy and cannot retain
unexpected statements.

Signed-off-by: Artem Muterko <artem@sopho.tech>
(cherry picked from commit 321911b96c)
2026-06-03 15:47:55 +00:00
Santosh 726a3837a2 osd: add cryptsetup resize for encrypted host based osds
when using encryptedDevice:true with host based (non pvc) osds, resizing
the underlying disk and restarting the OSD didn't expand the OSD because
the LUKS encryption layer was not resized.

This PR adds resize capabilities inside the activate container. Full
device stack must be resized: PV → LV → LUKS

Steps:
1. Run pvresize and lvextend to grow the LVM layers.
2. Retrieve the LUKS key from the Ceph config-key store using the
lockbox crednetials already available in the activate container
3. Pass the key explicitly to cryptsetup resize via --key-file.

Signed-off-by: Santosh <sapillai@redhat.com>
(cherry picked from commit 3e49c9f4cf)
2026-06-01 00:24:38 +00:00
Oded Viner 9f6809985d object: add logger for AWS SDK v2 debug signing output
Wire a rookLogger adapter implementing smithy's Logger
interface so that aws.LogSigning output is emitted through
rook's capnslog when debug mode is enabled.

Signed-off-by: Oded Viner <oviner@redhat.com>
(cherry picked from commit 4b5fc9d16a)
2026-05-28 15:17:50 +00:00
Oded Viner 37e92fd749 object: remove AWS SDK v1 dependency
Remove the AWS SDK v1 (github.com/aws/aws-sdk-go)
dependency entirely. All S3 operations now use AWS SDK v2
exclusively.

- Remove the v1 Client field from S3Agent struct and
  rename ClientV2 to Client
- Remove v1 session/client initialization from NewS3Agent
- Update all call sites referencing ClientV2
- Convert integration tests to use v2 API calling
  conventions (context parameter) and smithy error handling
- Remove aws-sdk-go v1.55.8 from go.mod

Signed-off-by: Oded Viner <oviner@redhat.com>
(cherry picked from commit 254965fff1)

# Conflicts:
#	go.mod
#	go.sum
2026-05-27 10:38:23 -06:00
subhamkrai e10393f5a2 object: add TLS 1.3 cipher suite support for RGW beast frontend
adding new tls ssl_ciphersuites supporting tls 1.3 and the existing,
ssl_cipher supports tls 1.2 and below. Adding, the docs and unit-test
changs as well.

Signed-off-by: subhamkrai <srai@redhat.com>
(cherry picked from commit b70a0c9257)
2026-05-27 15:52:52 +00:00
Joshua Hoblitt d65da7520a Merge pull request #17597 from rook/mergify/bp/release-1.20/pr-17544
helm: update prometheus rules from upstream ceph/ceph; add `cluster` label to all scraped metrics (backport #17544)
2026-05-26 13:50:50 -07:00
Joshua Hoblitt bba4beeae7 core: reference all containers in a spec by name instead of index
This changeset excludes converting _test.go code as there may be a
legitimate reasons (E.g. convenience) for unit test to be sensitive to
ordering.

Related to:
- https://github.com/rook/rook/issues/15291
- https://github.com/rook/rook/pull/16969

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit d059c351fc)
2026-05-26 19:10:47 +00:00
Joshua Hoblitt 6e96e6b07f core: add cluster label to all servicemonitor scraped metrics
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit cc05619dd8)
2026-05-26 19:09:15 +00:00
Travis Nielsen 8abf77d763 Merge pull request #17566 from rook/mergify/bp/release-1.20/pr-17527
object: migrate sns topic provisioner to aws sdk v2 (backport #17527)
2026-05-20 10:36:35 -06:00
Joshua Hoblitt 1efea0123c core: rm unused consts
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
(cherry picked from commit 9eec9c730e)
2026-05-20 15:03:03 +00:00