Commit Graph
3161 Commits
Author SHA1 Message Date
sp98 42ea2897cf mon: allow changing hostNetwork settings
Allow changing spec.network.hostNetwork settings on a running cluster
by failing over the mons

Signed-off-by: sp98 <sapillai@redhat.com>
2023-12-18 10:30:48 +05:30
Travis Nielsen ca810f5c2c Merge pull request #13360 from sp98/mon-failover-hostnetwork
mon: fix mon failover on path change
2023-12-12 11:02:43 -07:00
sp98 4b6fd89365 mon: fix mon failover on path change
This PR fails over mon when the mon path is changed from hostPath to PVC or vice versa

Signed-off-by: sp98 <sapillai@redhat.com>
2023-12-12 10:32:05 +05:30
parth-gr c9fd382c13 subvolumegroup: add pinning spec in subvolumegroup CRD
subvolumegroup can be pinned by pintype and pinsetting
So adding the spec to enhance its functionality

Closes: https://github.com/rook/rook/issues/12607

Signed-off-by: parth-gr <paarora@redhat.com>
2023-12-11 21:54:25 +05:30
Rakshith R a3220d827d csi: update default cephcsi version to 3.10.0
This commit updates default cephcsi driver version
to v3.10.0 and filesystem reconciler now creates
csi subvolumegroup by default.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-12-06 19:57:40 +05:30
travisn 6f8e42422d core: fix golang linter issues with variables in loops
Loop variables cannot be reliably uses since they will
change with each iteration. Update these loop variable
uses to be safe by indexing the slice rather than
using the loop variable directly.

Also suppress the linter issues for passwords used
in tests.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-12-05 14:14:58 -07:00
Travis Nielsen 2edaf3a1f7 Merge pull request #13270 from parth-gr/radosnamespace-name
namespace: add name spec in cephBlockPoolRadosNamespace CRD
2023-12-05 10:46:55 -07:00
Alexander Trost 3097455d78 Merge pull request #13246 from koor-tech/ceph_config_via_cluster_crd_impl
operator: allow setting ceph config options via ceph cluster crd
2023-12-02 11:27:20 +01:00
Alexander Trost 4ed35d6bd6 operator: allow setting ceph config options via ceph cluster crd
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-12-02 00:05:48 +01:00
Blaine Gardner f4d67fd4db Merge pull request #13261 from subhamkrai/remove-controller-runtime
core: remove webhook & controller-runtime from apis
2023-12-01 09:55:34 -07:00
subhamkrai 28cc1ebc55 core: remove webhook & controller-runtime from apis
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-01 14:15:40 +05:30
Travis Nielsen 1efb694050 Merge pull request #13266 from parth-gr/svg-name
subvolumegroup: add name spec in subvolumegroup CRD
2023-11-30 10:16:40 -07:00
parth-gr 939a1b7d20 subvolumegroup: add name spec in subvolumegroup
Originally we create it using this cmd
ceph fs subvolume create <vol_name> <subvol_name>
So we can have 2 variables filesystem and subvolume name,
Currently the CR doesn't allow us to make subvolume-name
as constant as needed to "csi" because of k8s limitations

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-30 20:39:11 +05:30
Michael Adam 93bf6540b2 mgr: honor the ContinueUpgradeAfterChecksEvenIfNotHealthy flag
Fixes: #13167

Previously, the mgr did not honor the flag
ContinueUpgradeAfterChecksEvenIfNotHealthy
from the cluster spec. Only osd, mds, and rgw did.

To render the update behavior correct and complete across the daemons, this
change implements the honoring of the flag for the mgr.

Signed-off-by: Michael Adam <obnox@samba.org>
2023-11-29 20:23:25 +01:00
parth-gr 5bf3847b13 namespace: add name spec in cephBlockPoolRadosNamespace CRD
instead of use metadata name as the backend resource name
we need to add a new name in the spec which can be the
actual backend name

Closes: https://github.com/rook/rook/issues/13220

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-29 22:35:21 +05:30
Rakshith R 59cb0dd4bf csi: add CSIDriverOptions section in cephCluster CR
This commit adds new CSIDriverOptions section in
cephCluster CR. This section contains settings
for read affinity and kernel+fuse Mount options
These settings will be injected directly into
rook-ceph-csi-config cm to be applicable per
ceph cluster.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-11-29 19:34:28 +05:30
Travis Nielsen d1c4667c8b Merge pull request #13225 from ushitora-anqou/use-regex-match-for-pdb-reset
core: add pgHealthyRegex to DisruptionManagementSpec
2023-11-27 14:23:25 -07:00
Travis Nielsen 49c4163b32 Merge pull request #13256 from rkachach/fix_issue_radosgw_admin
mgr: adding CEPH_ARGS to the mgr pod so radosgw-admin can use it
2023-11-27 10:01:12 -07:00
Divyansh Kamboj f9d6cd7f3a exporter: change deployment strategy to Recreate
Restarting the exporter using RollingRelease causes a race condition,
that results in exporter crashing and the ceph health to show a warning.

Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2023-11-27 15:20:25 +05:30
Redouane Kachach afc485fa03 mgr: adding CEPH_ARGS to the mgr pod so radosgw-admin can use it
ceph dashboard uses radosgw-admin for certain tasks that
aren't accessible via the rgw REST API. Due to the absence of
a valid ceph.conf file at /etc/ceph/ceph.conf within the mgr pod,
radosgw-admin fails to operate, resulting in 500 errors across
various 'Object Gateway' views on the dashboard. This change
adds CEPH_ARGS environment variable to the mgr pod enabling
its propagation and utilization by the dashboard/radosgw-admin
for executing rgw commands.

closes: https://github.com/rook/rook/issues/13255

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-23 10:48:55 +01:00
Ryotaro Banno 0ad88e9d68 core: add pgHealthyRegex to DisruptionManagementSpec
This patch adds `pgHealthyRegex` field to DisruptionManagementSpec.
`pgHealthyRegex` is a regular expression that is used to determine which
PG states should be considered healthy. The default value of
`pgHealthyRegex` is:

        ^(active\+clean|active\+clean\+scrubbing|active\+clean\+scrubbing\+deep)$

which is effectively the same as before.

Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
2023-11-22 00:19:04 +00:00
Travis Nielsen 510d2721b4 Merge pull request #13247 from riya-singhal31/master
csi: add csi-addons sidecar to cephfs deployment
2023-11-21 16:13:31 -07:00
Travis Nielsen da2513e484 Merge pull request #13248 from rkachach/fix_issue_exporter_interval
mgr: get servicemonitor exporter's interval from MonitoringSpec
2023-11-21 16:02:38 -07:00
Riya Singhal 401ce4697a csi: add csi-addons sidecar to cephfs deployment
this commit adds csi-addon sidecar to
cephfsplugin provisioner deployment

Signed-off-by: Riya Singhal <rsinghal@redhat.com>
2023-11-22 02:50:45 +05:30
Redouane Kachach 6b71325dbe mgr: get servicemonitor exporter's interval from MonitoringSpec
this change updates the serviceMonitor interval of the the rook-ceph-exporter
with the value from the MonitoringSpec.

closes: https://github.com/rook/rook/issues/13159

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-21 21:47:28 +01:00
Travis Nielsen 38ceb90ab2 Merge pull request #13211 from subhamkrai/pr/gauravsitlani/12957
core: operator to skip reconcile of mgr, rgw, mds and rbd-mirror
2023-11-21 12:15:34 -07:00
Travis Nielsen 346746c43b Merge pull request #13244 from iPraveenParihar/upgrade-sidecar-versions
csi: update csi sidecars' image version
2023-11-21 12:10:49 -07:00
avanthakkar 8fa561484a exporter: run exporter with specific keyring
Similar to the ceph crash collector daemon that generates a keyring with more
restrictive privileges, the exporter should also generate and use a more limited keyring.

Signed-off-by: avanthakkar <avanjohn@gmail.com>
2023-11-21 23:52:50 +05:30
gauravsitlani 3d0049c547 core: operator to skip reconcile of mgr, rgw, mds and rbd-mirror daemons in debug
During certain maintenance tasks the admin will own running
operations on the ceph mgr, rgw, mds and rbd-mirror daemons
and the operator should not interfere with those operations.

Co-authored-by: gauravsitlani <gaurav.sitlani@live.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2023-11-21 20:14:16 +05:30
Praveen M ddf03bce6f csi: update csi sidecars' image version
csi-node-driver-registrar: v2.9.1
csi-resizer: v1.9.2
csi-provisioner: v3.6.2
csi-attacher: v4.4.2
csi-snapshotter: v6.3.2
Signed-off-by: Praveen M <m.praveen@ibm.com>
2023-11-21 14:41:57 +05:30
Travis Nielsen 297e8400e5 Merge pull request #12850 from parth-gr/node-telemetry
core: report node metrics using ceph telemetry
2023-11-16 14:55:44 -07:00
parth-gr f007f2aca1 core: report node metrics using ceph telemetry
Add this reporting with the cephcluster reconcile,
Similar way we reported other telemetry's

Closes: https://github.com/rook/rook/issues/12344

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-16 19:13:22 +05:30
travisn 03d077aa6b core: remove support for ceph pacific
Pacific is end of life and no longer necessary to
support in Rook with v1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-14 17:07:03 -07:00
Blaine Gardner d01029a4b0 Merge pull request #13206 from BlaineEXE/multus-fix-network-canary-all-placement
multus: fix placement error for net addr detect job
2023-11-14 17:58:19 -06:00
Blaine Gardner 0a538bfc37 multus: fix placement error for net addr detect job
Fix an issue in the network address detection job where placement was
only retreived from osd and not merged with all.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-11-14 12:36:58 -07:00
Travis Nielsen 1f157df699 Merge pull request #12817 from parth-gr/objectstore-improvements
object: improve the error handling for multisite objs
2023-11-14 12:33:38 -07:00
parth-gr 46c241433d object: improve the error handling for multisite objs
here is the https://go.dev/play/p/SS9Q-dAiIx3 example which says the
error handling was wrongly implemented

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-14 15:12:15 +05:30
Shachar Sharon e05184d0a7 nfs: allow livness-probe for nfs-ganesha container
Use K8s LivenessProbe mechanism to check OK-status of nfs-ganesha
container. A user may define his own lineness-probe, or a default one
which expects NFS TCP-port 2049 to be active; that is, willing to accept
new connections: for Ceph>=18.2.1 issue 'rpcinfo' call on local pod;
otherwise use standard K8s TCP-socket liveness probe mechanism.
Define permissive values to liveness-probe to ensure that the NFS
service is defined in failed-state only when it has non-recoverable
error.

The current default definition of LivenessProbe is expected to guard the
nfs pod from at least the following two cases:

  - Deadlocks: where an nfs-ganesha server is running, but unable serve
    new connections due to internal bad-state.

  - Resource exhaustion on the host node (e.g. OOM) which prevents the
    server from accepting new connections and reply to NULL RPC request.

In both cases we expect K8s to reschedule the nfs pod, most likely on
different host node.

Refs rook issue #12719

Signed-off-by: Shachar Sharon <ssharon@redhat.com>
2023-11-13 16:21:24 +02:00
Blaine Gardner a7fce61f68 Merge pull request #13129 from BlaineEXE/multus-use-rook-image-for-ip-detect
multus: use rook image for ip range detection
2023-11-09 11:29:24 -06:00
Travis Nielsen 635570a168 Merge pull request #13179 from rkachach/fix_issue_13159
mgr: set interval of serviceMonitor to the value from MonitoringSpec
2023-11-09 07:17:41 -06:00
Redouane Kachach 06bc976959 mgr: set interval of serviceMonitor to the value from MonitoringSpec
This change updates the serviceMonitor interval field with the
value from the MonitoringSpec.

closes: https://github.com/rook/rook/issues/13159

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-09 06:35:20 +01:00
Blaine Gardner db1ca8c93e multus: use rook image for ip range detection
Use the Rook image (defined by the operator pod) to detect the Multus
network address ranges. It is reasonable for users to want to have a
minimal Ceph image that does not have the `ip` utility installed, which
is used for detecting the address ranges of multus interfaces. Instead,
use the Rook image, which Rook can ensure has the `ip` tool if Ceph ever
removes it from their image.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-11-08 10:11:04 -07:00
Madhu Rajanna 5658e026cb csi: remove deprecated grpc metrics code
GRPC metrics got deprecated in cephcsi
3.7.0 and the deprecated flags will get
removed in the next release. This PR
removes the deprecated metrics code which
allow us to run with older cephcsi as well.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2023-11-07 11:24:17 +01:00
Travis Nielsen 9738a0f19d Merge pull request #13071 from travisn/reef-base-image
core: Update operator base image to reef
2023-10-31 11:02:12 -06:00
Travis Nielsen 674407db7b Merge pull request #12952 from sp98/hostpath-to-pvs
mon: failover mons from hostpath to persistent volumes
2023-10-30 11:31:21 -06:00
sp98 09356636fa mon: failover mon from hostpath to pv
failover mons from  hostpath to pv and vice versa

Signed-off-by: sp98 <sapillai@redhat.com>
2023-10-30 11:46:00 +05:30
subhamkrai ef4dd76df1 pool: rbd cmd shouldn't use admin in external mode
when creating networkFence, rbd command was loading
admin config and hence running rbd command use client.admin
in case of external cluster also. With this commit instead
of client.admin user it will use what is being passed to
config.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-10-27 14:45:01 +05:30
travisn 80244fa6ba object: change is_master from string to bool
In Reef the is_master changed from a string to a bool
so we must update the type for proper json
serialization.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-10-25 15:26:17 -06:00
travisn a421143ddd Revert "core: use crash profile in crash daemon keyring"
This reverts commit fab23d3407.
The mgr requires rw access for the cron job that collects
the crashes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-10-20 15:46:36 -06:00
Bin WangandTravis Nielsen 4a705b46b0 osd: print warning message if no matching node found for osd
As seen in https://github.com/rook/rook/issues/1988, it's a common
mistake to configure rook nodes with names that don't match Kubernetes's
node label. This PR prints more detailed message to help debugging
problem.

Signed-off-by: Bin Wang <bin.wang@mail.binwang.me>

Add break according to review comment

Co-authored-by: Travis Nielsen <tnielsen@redhat.com>

Update log message according to review comment

Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2023-10-16 18:37:25 -04:00