Commit Graph
78 Commits
Author SHA1 Message Date
Travis Nielsen 955f7c5428 Merge pull request #17785 from beyondcloud-co/fix/17761-duplicate-mgr-ips-in-endpointslice
fix(ceph): deduplicate external MGR endpoints in EndpointSlice
2026-07-22 09:05:31 -06:00
Joshua Hoblitt 49612461a4 docs: fix function comments to match their declaration names
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-26 12:52:39 -07:00
beyondcloud-co 8134d683e3 external: deduplicate external MGR endpoints in EndpointSlice
When the active MGR IP already exists at a non-zero position in the
externalMgrEndpoints list, the rotation logic that replaces endpoints[0]
creates duplicate IPs. Also guards against mgr map failure which could
overwrite endpoints[0] with an empty string.

Fixes: #17761

Signed-off-by: beyondcloud-co <87993136+beyondcloud-co@users.noreply.github.com>
2026-06-18 17:16:46 -04:00
Joshua Hoblitt 9eec9c730e core: rm unused consts
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-05-14 09:34:52 -07:00
Santosh 925acb3e92 core: remove newlines from liveness probe scripts
Strip unnecessary leading and trailing newlines from osdLivenessProbeScript and livenessProbeScript raw string literals in spec.go

Signed-off-by: Santosh <sapillai@redhat.com>
2026-04-24 14:24:51 +05:30
parth-gr 0a3d9b8517 external: fix endpoint slice type for ipv6 clusters
in endpoint slice for the mgr and rgw service
it was harcoded to use the ipv4,
now dynamically assign the type based on the endpoint
ip adrress type

Signed-off-by: parth-gr <partharora1010@gmail.com>
2025-12-08 15:52:50 +05:30
subhamkrai 7004213f92 mgr: add required k8s label for endpointSlice
k8s requires the endpointSlice must have label
`kubernetes.io/service-name` to be associated with
service. Without this label, the k8s service controller
can't link the EndpointSlice to the service.

Signed-off-by: subhamkrai <srai@redhat.com>
2025-10-17 15:50:55 +05:30
parth-gr e1873cb2e7 external: fix ipv6 monitoring endpoint reconcile
currently the extraction ipv6 was not correct
used the net package to extract the host ip from
the ipadress

Signed-off-by: parth-gr <partharora1010@gmail.com>
2025-09-12 12:50:05 +05:30
subhamkrai 1bde614eb5 core: migrate from v1.Endpoints to discoveryv1.EndpointSlice
This change updates the codebase to replace the deprecated
v1.Endpoints API with discoveryv1.EndpointSlice.
The v1.Endpoints API has been deprecated in Kubernetes v1.33+

Signed-off-by: subhamkrai <srai@redhat.com>
2025-07-23 21:02:02 +05:30
Patryk Rostkowski 7f3e6bd7fa mon: allow running mon pods as root
This change addresses a permission issue where mon pods crashloop
on some Kubernetes setups with SELinux enabled, even when
ROOK_HOSTPATH_REQUIRES_PRIVILEGED is set.

To resolve this, a new function makeMonSecurityContext() was introduced
to explicitly set runAsUser: 0 at the pod level for mon pods
when the environment variable ROOK_CEPH_MON_RUN_AS_ROOT is set to true.

Additionally, the function previously named PodSecurityContext() was renamed
to DefaultContainerSecurityContext() to avoid confusion between
container-level and pod-level security context configuration. All container
SecurityContext usages across Ceph daemons were updated to reflect this change.

This ensures the root user configuration is applied
automatically and consistently in environments where it is required
for mon pod startup, while preserving clear separation of container
and pod-level security logic.

Signed-off-by: Patryk Rostkowski <patrostkowski@gmail.com>
2025-06-16 18:29:31 +02:00
parth-gr 79f44a4660 nfs: fix the skip reconcile call
currently nfs is used as the unique selector
for daemon id, which returns  nfs: ocs-storagecluster-cephnfs-a

Updating the has() to function to check the same label value
(n.Name + "-" + id)

Signed-off-by: parth-gr <partharora1010@gmail.com>
2025-05-30 09:11:49 +05:30
Carlos Barria c63ebbe0e3 core: fix golangci-lint check ST1019
This change cleans up and standardizes import statements across the Ceph operator code. It removes redundant or duplicate imports and reorganizes alias names for improved clarity and consistency. Additionally, the ST1019 exception was removed from .golangci.yaml now that the code complies with the rule.

Signed-off-by: Carlos Barria <cbarria@yahoo.com>
2025-05-27 17:38:22 -04:00
Artem Torubarov c1fd2f2ee8 rgw: use pod name in ops log filename
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2025-04-29 09:43:35 +02:00
Joshua Hoblitt 3cb343f62a core: run gofumpt on all files
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2025-03-26 10:41:48 -07:00
Deepika Upadhyay c5d27467d6 object: add rgw ops sidecar for op logs
the rgw operations for s3 can now be accessible using sidecar
rgw-ops-log availabe in json form that can be further filtered logging
for observability, this will set the rgw_enable_ops_log setting

Signed-off-by: Deepika Upadhyay <deepika.upadhyay@clyso.com>
2024-12-17 22:36:03 +05:30
df511fb58f ci: update golangci-lint to the latest version (v1.62)
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.

Additionally, it  silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
 and string format errors found by golangci-lint, while at it.

Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
2024-12-14 14:47:30 +01:00
Peter Razumovsky 97a13628e5 core: add capabilities to securityContext to fix CIS 5.2.8
Resolves CIS benchmark 5.2.8 rule by adding capabilities with
explicitly defined empty "add" list (where it needed) and with
NET_RAW capability in "drop" list.

5.2.8 Minimize the admission of containers with added capabilities

Containers must drop the `NET_RAW` capability and are not permitted
to add back any capabilities.

Signed-off-by: Peter Razumovsky <prazumovsky@mirantis.com>
2024-11-05 16:27:43 +04:00
parth-gr d3742119b9 csi: add log rotation for csi rbd pod containers
1) Make the csi rbd container logs persisted in a file
   (csi plugin, csi provisioner, csi addons sidecar)

2) Use the cephcluster api specs to configure the log rotate

3) Add log rotation to rotate the log file and
   Add a sidecar log collector container

part-of: https://github.com/rook/rook/issues/12809

Signed-off-by: parth-gr <partharora1010@gmail.com>
2024-06-27 14:20:18 +05:30
gauravsitlani 3d0049c547 core: operator to skip reconcile of mgr, rgw, mds and rbd-mirror daemons in debug
During certain maintenance tasks the admin will own running
operations on the ceph mgr, rgw, mds and rbd-mirror daemons
and the operator should not interfere with those operations.

Co-authored-by: gauravsitlani <gaurav.sitlani@live.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2023-11-21 20:14:16 +05:30
travisn 03d077aa6b core: remove support for ceph pacific
Pacific is end of life and no longer necessary to
support in Rook with v1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-14 17:07:03 -07:00
Shachar Sharon e05184d0a7 nfs: allow livness-probe for nfs-ganesha container
Use K8s LivenessProbe mechanism to check OK-status of nfs-ganesha
container. A user may define his own lineness-probe, or a default one
which expects NFS TCP-port 2049 to be active; that is, willing to accept
new connections: for Ceph>=18.2.1 issue 'rpcinfo' call on local pod;
otherwise use standard K8s TCP-socket liveness probe mechanism.
Define permissive values to liveness-probe to ensure that the NFS
service is defined in failed-state only when it has non-recoverable
error.

The current default definition of LivenessProbe is expected to guard the
nfs pod from at least the following two cases:

  - Deadlocks: where an nfs-ganesha server is running, but unable serve
    new connections due to internal bad-state.

  - Resource exhaustion on the host node (e.g. OOM) which prevents the
    server from accepting new connections and reply to NULL RPC request.

In both cases we expect K8s to reschedule the nfs pod, most likely on
different host node.

Refs rook issue #12719

Signed-off-by: Shachar Sharon <ssharon@redhat.com>
2023-11-13 16:21:24 +02:00
sp98 4eb9f62205 osd: make osd pod to sleep when osds are flapping
When OSDs flap, ceph stops the OSD daemon if its marked down greater than
5 times in 600 seconds. But OSD pod restarts and marks the OSD `up` again.
This causes the PGs mapped to these OSDs to peer. While the PGs are peering,
IO to these PGs are blocked.

So we need to ensure that if ceph is marking OSD `down` due to flapping, OSD pod
should not restart to mark the OSDs `up` again.

This PR adds a sleep to the OSD pod if the container returned with a 0 exit code
Default behavior is to sleep for 6 hrs. But user can configure it from the
ceph cluster spec.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-09-15 22:59:46 +05:30
subhamkrai 29d2b6a071 core: restart ceph daemons when network updated
We need to restart all the ceph daemons whenever
cephCluster network settings are modified like
requiremsgr2, encryption and compression. This
required for Ceph to consider the new settings
it require new ceph daemons all over.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-01 11:30:48 +05:30
subhamkrai a25071ac29 ci: fix golangCI lint remove k8s.io/utils/pointer
golangci linter was throughing error `k8s.io/utils/pointer`
package is deprecated. So, I have removed that.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-16 21:48:42 +05:30
avanthakkar 541d091f9c core: use ROOK_CEPH_MON_HOST from config store in OSD pods too
Signed-off-by: avanthakkar <avanjohn@gmail.com>

Volume "ceph-daemons-sock-dir" is coming empty in case if dataDirHostPath,
which is the case for osd onPVC. Fix the volume creation by using the
ceph cluster spec dataDirHostPath, which allows to run socket commands
on osd containers.
2023-06-01 21:09:58 +05:30
Redouane Kachach 0ae7867dd1 Revert "mgr: use k8s readiness probe to implement mgr HA"
This reverts commit dc76f81fea.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-03-07 13:12:50 +01:00
Redouane Kachach 0c652d28e2 Revert "ci: fix MultiClusterDeploySuite CI"
This reverts commit b464428978.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-03-07 13:12:50 +01:00
Travis Nielsen 0a80789367 Merge pull request #11690 from rkachach/fix_issue_11685
ci: fix MultiClusterDeploySuite CI
2023-02-16 11:19:04 -07:00
Redouane Kachach b464428978 ci: fix MultiClusterDeploySuite CI
This fixes the the MultiClusterDeploySuite CI failure. The operator
is timing out waiting for the mgr deployments to be
ready, according to the WaitForDeploymentToStart() method. After
the reconcile times out after about five minutes, the next
reconcile succeeds since the wait is only done for new mgr
deployments.

Closes: https://github.com/rook/rook/issues/11685

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-02-16 18:43:38 +01:00
Travis Nielsen 3dabc6dcb6 Merge pull request #11317 from avanthakkar/introduce-ceph-exporter
core: introduce ceph-exporter
2023-02-15 11:59:05 -07:00
Travis Nielsen b008c5753f Merge pull request #11643 from rkachach/fix_issue_11640
mgr: use k8s readiness probe to implement mgr HA
2023-02-15 09:51:48 -07:00
Avan Thakkar 460900756c core: add service monitor for ceph-exporter service
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2023-02-15 15:44:56 +05:30
Redouane Kachach dc76f81fea mgr: use k8s readiness probe to implement mgr HA
The idea behind this change is to use the readiness probe to implement
the mgr HA mechanism. In the current ceph mgr implementation only the
active instance offers the command 'mgr_status' through the admin
socket. We use this command combined with a Readiness Exec Probe to
detect which manager is active. Kubernetes will automatically mark
it as ready and redirect any service traffic to the active instance.

Closes: https://github.com/rook/rook/issues/11640
Closes: https://github.com/rook/rook/issues/11638
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-02-15 10:12:16 +01:00
subhamkrai 036715c3b2 rbdmirror: set log rotation from 7 to 4x i.e 28
increasing the rotation from default 7 to 28 as
in case of rbdmirror logs seems not enough in some cases
with maxLogSize 500 so it's better to increase the rotation
for rbdmirror specific.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-02-13 21:37:57 +05:30
Avan Thakkar 2f8ee60374 core: introduce ceph-exporter
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2023-02-06 02:47:09 -05:00
Travis Nielsen df6d7af355 security: run the crash collector as ceph user
The crash collector does not have the command line arguments
to run as ceph user id 167, so we set the security context
to run as the ceph user in the main crash collector
container.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-10-27 13:12:21 -06:00
Travis Nielsen 7f0c83ad18 Merge pull request #10986 from randymtz/increase-liveness-timeout
core: increase liveness probe timeout to 5s
2022-10-17 11:49:18 -06:00
Avan Thakkar 934aa91056 operator: make imagePullPolicy customizable for csi driver and ceph pods
Introduce a new env variable ROOK_CSI_IMAGE_PULL_POLICY in rook operator configmap which should be used to
customize the imagePullPolicy for the csi driver and imagePullPolicy property in cephVersionSpec for ceph pods.

Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2022-09-27 11:57:43 +05:30
Randy J. Martinez ac9df66b76 core: increase liveness probe timeout to 5s
stability issues have been observed with 1s.
socket latency is expected whenever CPUs are
under minor pressure. Increasing value to
5s should cover most small-medium scale envs.

Resolves BZ: 2126566

Signed-off-by: Randy J. Martinez <randy@cephtips.com>
2022-09-13 17:32:43 -05:00
motorailgun ff0951738f operator: improve ProbeHandler error message
This commit implements more diagnostic error for unsccessful
Liveness- and Readiness-Probe of Pods.

Added codes are expected to catch failures of
`ceph status` and `ceph mon_status`, and report it.

Closes: https://github.com/rook/rook/issues/9846
Closes: https://github.com/rook/rook/issues/9852

Signed-off-by: motorailgun <motoi_public@mail.aria-on-the-planet.es>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-09-07 07:35:33 +00:00
subhamkrai 9d4d5bc702 core: fix logrotate bash check and periodicity logic
we need use `!=` for string comparision in bash instead of
`-ne`. Also, need to correct periodicity if condition to
make it work as expected.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-08-22 17:14:51 +05:30
subhamkrai 62f73dcd98 core: add support to rotate log based on logfile size
this commits add new field `MaxLogSize` inside `LogCollectorSpec` of
cephCluster cr which will take max size of log after which we want to
rotate the log.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-07-26 13:06:39 +00:00
Travis Nielsen dad97f3425 core: remove support for ceph octopus
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-07 15:03:26 -06:00
subhamkrai 62f0fb2b1d core: increase liveness probe timeout to 2s
we have noticed multiple failures because of the probe
failing, most of the time it's due to fewer resources.
But increasing timeout fixes that, so increasing the
probe timeout to 2s from default 1s so that it will
give more time to probe before failing.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-06-17 19:06:40 +05:30
Travis Nielsen 28e721d877 osd: allow the osd to take a long time to start
The startup probe for the OSD has been too aggressive to kill the OSD
in case the OSD is taking some time to start. The OSD may be self-optimizing,
scrubbing, or some other internal operation before it is ready to start.
Rather than disable the startup probe completely, the default is now
to retry for two hours in case the OSD is performing those operations.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-05-11 14:29:25 -06:00
subhamkrai bf7daccf60 core: make code changes to support latest cntrl runtime
making necessary code changes to support controller
runtime version.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-20 22:30:15 +05:30
Sébastien Han 5404ec13a2 core: dereference pointer before trying to compare with deepequal
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.

Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-27 17:41:42 +01:00
Blaine Gardner c07d89d9ea core: rgw: allow specifying daemon startup probes
Allow specifying daemon startup probes where we also allow configuring
liveness probes. Startup probes allow Rook to tolerate when Ceph daemons
occasionally take a long time to start up while not also making
Kubernetes liveness probes slower to detect runtime failures of daemons.

Startup probes are beta in Kubernetes 1.18, so we should not enable
probes by default for earlier Kubernetes versions.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-21 15:12:41 -07:00
Sébastien Han 5c1e459a6a mgr: run active-watch as root and privileged
For now, we must run the container with UID 0 and privileged for
multiple reasons:

* the rook binary writes ceph config to /var/lib/rook which is owned by
  root
* it's difficult to use /etc/ceph since it will conflict with the
  rook-ceph-override configmap AND is also owned by root since it's a
  mounted configmap.
* using /etc/ceph might be possible but has other issues with rook's
  exec package since the ceph config is built from /var/lib/rook

Closes: https://github.com/rook/rook/issues/9385
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-13 14:02:33 +01:00
Yuichiro Ueno d1d252c5c2 core: add context parameter to k8sutil endpoint
This commit adds context parameter to k8sutil endpoint functions. By
this, we can handle cancellation during API call of endpoint resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-12-11 11:26:21 +09:00