Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
When the active MGR IP already exists at a non-zero position in the
externalMgrEndpoints list, the rotation logic that replaces endpoints[0]
creates duplicate IPs. Also guards against mgr map failure which could
overwrite endpoints[0] with an empty string.
Fixes: #17761
Signed-off-by: beyondcloud-co <87993136+beyondcloud-co@users.noreply.github.com>
AddCephVersionLabelToDaemonSet had no callers anywhere in the repo
(verified with deadcode and repo-wide grep); rook applies ceph version
labels to Deployments, Jobs, and object metas, but never to DaemonSets.
The shared addCephVersionLabel helper remains in use by the live
siblings.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Strip unnecessary leading and trailing newlines from osdLivenessProbeScript and livenessProbeScript raw string literals in spec.go
Signed-off-by: Santosh <sapillai@redhat.com>
in endpoint slice for the mgr and rgw service
it was harcoded to use the ipv4,
now dynamically assign the type based on the endpoint
ip adrress type
Signed-off-by: parth-gr <partharora1010@gmail.com>
For all of the controllers besides the cluster controller,
the logging now includes the namespaced name of the resource
that is being reconciled. This will help with log troubleshooting
to help analyze logs consistently for the resource being
reconciled.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For Rook v1.19,removing support for Ceph Reef. It is at end of life.
Users on Reef can continue to use Rook v1.18.x or older.
The latest two Ceph versions Squid and Tentacle are supported.
Signed-off-by: subhamkrai <srai@redhat.com>
k8s requires the endpointSlice must have label
`kubernetes.io/service-name` to be associated with
service. Without this label, the k8s service controller
can't link the EndpointSlice to the service.
Signed-off-by: subhamkrai <srai@redhat.com>
currently the extraction ipv6 was not correct
used the net package to extract the host ip from
the ipadress
Signed-off-by: parth-gr <partharora1010@gmail.com>
If the network is encrypted with msgr2, cephfs requires
the mount options to include ms_mode=secure. Now this
will be set by default if the mount options are not
already set in the cephcluster CR or the operator env
var CSI_CEPHFS_KERNEL_MOUNT_OPTIONS.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If the cephfs kernel mount options are not specified in the
cephcluster.spec.csi CR settings, fall back to the setting
in the operator env var CSI_CEPHFS_KERNEL_MOUNT_OPTIONS.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Implement CephX key rotation for Rook's client.admin user.
Admin user rotation is risky, so this has been tested extensively both
in unit tests as well as by manually injecting failures during runtime.
In testing, all failures were able to be recovered by the recovery
routine.
A mutex is also added to help ensure that two simultaneous admin key
rotation processes cannot be running simultaneously for any given
namespace. The mutex is tested in unit tests, and it was verified during
runtime via manual testing.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Cleanup up overall usage of CephX statuses in CephCluster.
Convert all CephX statuses to non-pointer types. There is no specific
need for pointer types, and keeping these as non-pointers means there is
not a risk of nil pointer exceptions in all the various Rook controllers
for upgraded clusters.
Move initialization of CephCluster CephX status items to
`preMonStartupActions()`. This helps resolve 2 issues:
1. 'Uninitialized' state was being set for external CephClusters where
CephX statuses are not relevant.
2. 'Uninitialized' state was being set for upgraded clusters whose key
status cannot be determined.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
rotate the client.rbd-mirror-peer key based on the cephxconfig.
Also update the cephxStatus in the cephcluster as well as the
mirrored cephblockpool.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
add RecoverAndLogException() helper to log panics with stack trace.
added defer call in all rook controller Reconcile() methods for
better error visibility in operator logs
Signed-off-by: Oded Viner <oviner@redhat.com>
This change updates the codebase to replace the deprecated
v1.Endpoints API with discoveryv1.EndpointSlice.
The v1.Endpoints API has been deprecated in Kubernetes v1.33+
Signed-off-by: subhamkrai <srai@redhat.com>
This change addresses a permission issue where mon pods crashloop
on some Kubernetes setups with SELinux enabled, even when
ROOK_HOSTPATH_REQUIRES_PRIVILEGED is set.
To resolve this, a new function makeMonSecurityContext() was introduced
to explicitly set runAsUser: 0 at the pod level for mon pods
when the environment variable ROOK_CEPH_MON_RUN_AS_ROOT is set to true.
Additionally, the function previously named PodSecurityContext() was renamed
to DefaultContainerSecurityContext() to avoid confusion between
container-level and pod-level security context configuration. All container
SecurityContext usages across Ceph daemons were updated to reflect this change.
This ensures the root user configuration is applied
automatically and consistently in environments where it is required
for mon pod startup, while preserving clear separation of container
and pod-level security logic.
Signed-off-by: Patryk Rostkowski <patrostkowski@gmail.com>
This is the first implementation of CephX key rotation in Rook.
Adds key rotation API from design PR 15915.
Also add helper methods for rotating keys, determining when keys
need to be rotated, and for generating CephX key statuses. These will be
reusable for other Rook reconciles beyond object/RGW.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
currently nfs is used as the unique selector
for daemon id, which returns nfs: ocs-storagecluster-cephnfs-a
Updating the has() to function to check the same label value
(n.Name + "-" + id)
Signed-off-by: parth-gr <partharora1010@gmail.com>
This change cleans up and standardizes import statements across the Ceph operator code. It removes redundant or duplicate imports and reorganizes alias names for improved clarity and consistency. Additionally, the ST1019 exception was removed from .golangci.yaml now that the code complies with the rule.
Signed-off-by: Carlos Barria <cbarria@yahoo.com>
This change adds support for skipping reconciliation of CephNFS daemons
that are labeled with `ceph.rook.io/skip-reconcile=true`.
Similar to MDS, MGR, and RGW components, this allows cluster operators
to prevent Rook from modifying specific NFS daemon deployments.
Signed-off-by: Patryk Rostkowski <patrostkowski@gmail.com>
When finalizers are added to the CRs, a follow-up reconcile
will be triggered due to the increased generation on the CR.
Therefore, abort the initial reconcile when adding the finalizer,
and allow the follow-up reconcile to complete the configuration.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
All existing controller runtime watches are converted to use "typed"
handlers and predicates instead of operating on `client.Object`. The
intent is to be bug for bug equivalent with the existing logic while
replacing run time type assertions and switch statements with compile
time type constraints and type casts. In several cases, functions using
assertions were split up such that each function only handles a single
Kind at a time. It is hoped that this will improve readability and
maintainability while facilitating future refactoring such as migrating
some watches to using IndexFields.
Of particular note is that the massive switch statement in
`WatchControllerPredicate()` from
`pkg/operator/ceph/controller/predicate.go` has been replaced with
generics, reflection, and splitting the obc logic into its own predicate
function. There are still many helper functions operating on
`client.Object`. These were not updated unless required by the compiler
in order to limit the size of this change. The type safety of these
funcs should be improved as followup work.
It is strongly suggested that going forward, handlers and predicates
only handle a single Kind (generic or not) and that switches / type
assertions are heavily discouraged or forbidden. This PR removed all
but a single switch statement in a predicate, which should be addressed
in future work.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Implements #14733. Allows to set IDs of external mons to
Cluster CRD. Rook will not remove external mons from quorum
and will add external mon addresses to mon endpoints.
Use-case for external mon is to maintain quorum for 2-AZ
k8s cluster in case of zone outage.
Signed-off-by: Artem Torubarov <artem.torubarov@clyso.com>
Some ceph health errors should not block the reconcile
of the cluster. Mgr modules do not have cause to block
the reconcile, as the cluster can usually work even
if a module is failing.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>