Commit Graph
1394 Commits
Author SHA1 Message Date
Travis Nielsen 591ddf2c3d Merge pull request #13420 from iPraveenParihar/api/csi-config-map
csi: use ceph-csi exposed csi config map API
2024-01-16 12:34:11 -07:00
Praveen M dfd3d625a8 csi: use ceph-csi exposed csi config map API
This commit uses an csi config map API exposed by Ceph-CSI.
This will reduce the efforts of keeping the configuration in
sync with Ceph-CSI.

Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-01-12 15:04:20 +05:30
Satoru Takeuchi 1935033a1e Merge pull request #13479 from cybozu-go/core-fix-error-handling-on-setting-watcher
core: fix error handling on setting watcher
2024-01-09 22:17:57 +09:00
dependabot[bot] f56f06f955 build(deps): bump the k8s-dependencies group with 4 updates
Bumps the k8s-dependencies group with 4 updates: [k8s.io/api](https://github.com/kubernetes/api), [k8s.io/apiextensions-apiserver](https://github.com/kubernetes/apiextensions-apiserver), [k8s.io/cli-runtime](https://github.com/kubernetes/cli-runtime) and [k8s.io/cloud-provider](https://github.com/kubernetes/cloud-provider).

Updates `k8s.io/api` from 0.28.4 to 0.29.0
- [Commits](https://github.com/kubernetes/api/compare/v0.28.4...v0.29.0)

Updates `k8s.io/apiextensions-apiserver` from 0.28.4 to 0.29.0
- [Release notes](https://github.com/kubernetes/apiextensions-apiserver/releases)
- [Commits](https://github.com/kubernetes/apiextensions-apiserver/compare/v0.28.4...v0.29.0)

Updates `k8s.io/cli-runtime` from 0.28.4 to 0.29.0
- [Commits](https://github.com/kubernetes/cli-runtime/compare/v0.28.4...v0.29.0)

Updates `k8s.io/cloud-provider` from 0.28.4 to 0.29.0
- [Commits](https://github.com/kubernetes/cloud-provider/compare/v0.28.4...v0.29.0)

---
updated-dependencies:
- dependency-name: k8s.io/api
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: k8s-dependencies
- dependency-name: k8s.io/apiextensions-apiserver
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: k8s-dependencies
- dependency-name: k8s.io/cli-runtime
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: k8s-dependencies
- dependency-name: k8s.io/cloud-provider
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: k8s-dependencies
...

Signed-off-by: dependabot[bot] <support@github.com>
2024-01-05 09:08:53 -07:00
travisn 56a3370d83 docs: update all examples to ceph v18.2.1
Since v18.2.1 is the default recommendation for the ceph version,
update all the docs and examples to that version.

Signed-off-by: travisn <tnielsen@redhat.com>
2024-01-03 14:37:11 -07:00
Satoru Takeuchi b1a81509a5 core: core: fix error handling on setting watcher
If `cmClient.Watch()` failed, `watcher.Stop()` causes null pointer
dereference.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2023-12-28 02:18:31 +00:00
Travis Nielsen c8d2955f98 Merge pull request #13348 from riya-singhal31/fencing
csi: implement network fencing for cephFS
2023-12-19 12:22:14 -07:00
Riya Singhal 9202e20829 csi: implement network fencing for cephFS
Signed-off-by: Riya Singhal <rsinghal@redhat.com>
2023-12-18 21:44:47 +05:30
sp98 42ea2897cf mon: allow changing hostNetwork settings
Allow changing spec.network.hostNetwork settings on a running cluster
by failing over the mons

Signed-off-by: sp98 <sapillai@redhat.com>
2023-12-18 10:30:48 +05:30
sp98 4b6fd89365 mon: fix mon failover on path change
This PR fails over mon when the mon path is changed from hostPath to PVC or vice versa

Signed-off-by: sp98 <sapillai@redhat.com>
2023-12-12 10:32:05 +05:30
Alexander Trost 3097455d78 Merge pull request #13246 from koor-tech/ceph_config_via_cluster_crd_impl
operator: allow setting ceph config options via ceph cluster crd
2023-12-02 11:27:20 +01:00
Alexander Trost 4ed35d6bd6 operator: allow setting ceph config options via ceph cluster crd
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-12-02 00:05:48 +01:00
Michael Adam 93bf6540b2 mgr: honor the ContinueUpgradeAfterChecksEvenIfNotHealthy flag
Fixes: #13167

Previously, the mgr did not honor the flag
ContinueUpgradeAfterChecksEvenIfNotHealthy
from the cluster spec. Only osd, mds, and rgw did.

To render the update behavior correct and complete across the daemons, this
change implements the honoring of the flag for the mgr.

Signed-off-by: Michael Adam <obnox@samba.org>
2023-11-29 20:23:25 +01:00
Rakshith R 59cb0dd4bf csi: add CSIDriverOptions section in cephCluster CR
This commit adds new CSIDriverOptions section in
cephCluster CR. This section contains settings
for read affinity and kernel+fuse Mount options
These settings will be injected directly into
rook-ceph-csi-config cm to be applicable per
ceph cluster.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-11-29 19:34:28 +05:30
Travis Nielsen d1c4667c8b Merge pull request #13225 from ushitora-anqou/use-regex-match-for-pdb-reset
core: add pgHealthyRegex to DisruptionManagementSpec
2023-11-27 14:23:25 -07:00
Travis Nielsen 49c4163b32 Merge pull request #13256 from rkachach/fix_issue_radosgw_admin
mgr: adding CEPH_ARGS to the mgr pod so radosgw-admin can use it
2023-11-27 10:01:12 -07:00
Divyansh Kamboj f9d6cd7f3a exporter: change deployment strategy to Recreate
Restarting the exporter using RollingRelease causes a race condition,
that results in exporter crashing and the ceph health to show a warning.

Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2023-11-27 15:20:25 +05:30
Redouane Kachach afc485fa03 mgr: adding CEPH_ARGS to the mgr pod so radosgw-admin can use it
ceph dashboard uses radosgw-admin for certain tasks that
aren't accessible via the rgw REST API. Due to the absence of
a valid ceph.conf file at /etc/ceph/ceph.conf within the mgr pod,
radosgw-admin fails to operate, resulting in 500 errors across
various 'Object Gateway' views on the dashboard. This change
adds CEPH_ARGS environment variable to the mgr pod enabling
its propagation and utilization by the dashboard/radosgw-admin
for executing rgw commands.

closes: https://github.com/rook/rook/issues/13255

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-23 10:48:55 +01:00
Ryotaro Banno 0ad88e9d68 core: add pgHealthyRegex to DisruptionManagementSpec
This patch adds `pgHealthyRegex` field to DisruptionManagementSpec.
`pgHealthyRegex` is a regular expression that is used to determine which
PG states should be considered healthy. The default value of
`pgHealthyRegex` is:

        ^(active\+clean|active\+clean\+scrubbing|active\+clean\+scrubbing\+deep)$

which is effectively the same as before.

Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
2023-11-22 00:19:04 +00:00
Travis Nielsen da2513e484 Merge pull request #13248 from rkachach/fix_issue_exporter_interval
mgr: get servicemonitor exporter's interval from MonitoringSpec
2023-11-21 16:02:38 -07:00
Redouane Kachach 6b71325dbe mgr: get servicemonitor exporter's interval from MonitoringSpec
this change updates the serviceMonitor interval of the the rook-ceph-exporter
with the value from the MonitoringSpec.

closes: https://github.com/rook/rook/issues/13159

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-21 21:47:28 +01:00
Travis Nielsen 38ceb90ab2 Merge pull request #13211 from subhamkrai/pr/gauravsitlani/12957
core: operator to skip reconcile of mgr, rgw, mds and rbd-mirror
2023-11-21 12:15:34 -07:00
avanthakkar 8fa561484a exporter: run exporter with specific keyring
Similar to the ceph crash collector daemon that generates a keyring with more
restrictive privileges, the exporter should also generate and use a more limited keyring.

Signed-off-by: avanthakkar <avanjohn@gmail.com>
2023-11-21 23:52:50 +05:30
gauravsitlani 3d0049c547 core: operator to skip reconcile of mgr, rgw, mds and rbd-mirror daemons in debug
During certain maintenance tasks the admin will own running
operations on the ceph mgr, rgw, mds and rbd-mirror daemons
and the operator should not interfere with those operations.

Co-authored-by: gauravsitlani <gaurav.sitlani@live.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2023-11-21 20:14:16 +05:30
Travis Nielsen 297e8400e5 Merge pull request #12850 from parth-gr/node-telemetry
core: report node metrics using ceph telemetry
2023-11-16 14:55:44 -07:00
parth-gr f007f2aca1 core: report node metrics using ceph telemetry
Add this reporting with the cephcluster reconcile,
Similar way we reported other telemetry's

Closes: https://github.com/rook/rook/issues/12344

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-16 19:13:22 +05:30
travisn 03d077aa6b core: remove support for ceph pacific
Pacific is end of life and no longer necessary to
support in Rook with v1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-14 17:07:03 -07:00
Redouane Kachach 06bc976959 mgr: set interval of serviceMonitor to the value from MonitoringSpec
This change updates the serviceMonitor interval field with the
value from the MonitoringSpec.

closes: https://github.com/rook/rook/issues/13159

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-09 06:35:20 +01:00
Travis Nielsen 674407db7b Merge pull request #12952 from sp98/hostpath-to-pvs
mon: failover mons from hostpath to persistent volumes
2023-10-30 11:31:21 -06:00
sp98 09356636fa mon: failover mon from hostpath to pv
failover mons from  hostpath to pv and vice versa

Signed-off-by: sp98 <sapillai@redhat.com>
2023-10-30 11:46:00 +05:30
subhamkrai ef4dd76df1 pool: rbd cmd shouldn't use admin in external mode
when creating networkFence, rbd command was loading
admin config and hence running rbd command use client.admin
in case of external cluster also. With this commit instead
of client.admin user it will use what is being passed to
config.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-10-27 14:45:01 +05:30
travisn a421143ddd Revert "core: use crash profile in crash daemon keyring"
This reverts commit fab23d3407.
The mgr requires rw access for the cron job that collects
the crashes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-10-20 15:46:36 -06:00
sp98 ffd69291f1 osd: use bash script if restart interval is > 0
If `FlappingRestartIntervalHours` for OSD is > 0 in the cephCluster CR,
then start OSD inside a bash script.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-10-11 11:38:26 +05:30
Blaine Gardner 29fa7771f8 Merge pull request #12998 from travisn/exporter-logging
exporter: Change verbose info to debug logs
2023-10-05 11:28:45 -06:00
Eng Zer Jun 77bff6a07c core: remove redundant len check
From the Go specification [1]:

  "1. For a nil slice, the number of iterations is 0."
  "3. If the map is nil, the number of iterations is 0."

`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.

[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len

Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
2023-10-05 20:16:46 +08:00
travisn 29fb879675 exporter: change verbose info to debug logs
Several messages were showing up in the operator log
more often than necessary, so let's change them to
debug level.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-10-04 12:45:39 -06:00
travisn fab23d3407 core: use crash profile in crash daemon keyring
The ceph documentation indicates to use the crash profile
instead of specifying rw for the mgr access.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-09-28 11:43:01 -06:00
Travis Nielsen 33fb7b801b Merge pull request #12926 from rkachach/fix_issue_12876
mgr: adding support for prometheus endpoint configuration
2023-09-27 08:09:13 -06:00
Redouane Kachach c763935436 mgr: adding support for prometheus endpoint configuration
Using the new configuration parameters prometheusEndpoint and
prometheusEndpointSSLVerify users can configure dashboard to
point to their prometheus instance.

closes: #12876

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-09-27 11:07:18 +02:00
Travis Nielsen cb9ffacce5 Merge pull request #12909 from testwill/pkg-import
core: import packages only once
2023-09-21 15:30:04 -06:00
guoguangwu 235ac293ff core: import packages only once
Signed-off-by: guoguangwu <guoguangwu@magic-shield.com>
2023-09-16 13:40:17 +08:00
Travis Nielsen a3dafbe9cf Merge pull request #12715 from sp98/osd-sleep
osd: make osd pod to sleep when osds are flapping
2023-09-15 12:24:08 -06:00
sp98 4eb9f62205 osd: make osd pod to sleep when osds are flapping
When OSDs flap, ceph stops the OSD daemon if its marked down greater than
5 times in 600 seconds. But OSD pod restarts and marks the OSD `up` again.
This causes the PGs mapped to these OSDs to peer. While the PGs are peering,
IO to these PGs are blocked.

So we need to ensure that if ceph is marking OSD `down` due to flapping, OSD pod
should not restart to mark the OSDs `up` again.

This PR adds a sleep to the OSD pod if the container returned with a 0 exit code
Default behavior is to sleep for 6 hrs. But user can configure it from the
ceph cluster spec.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-09-15 22:59:46 +05:30
travisn 66547d27be mgr: allow more than two mgrs
While two mgrs is considered sufficient, more mgr daemons are
now allowed in case the admin wants even more fault tolerance
to the active mgr going down, in case multiple mgrs go down.
Up to five mgr pods will be allowed. All mgrs will be in
standby mode except one active mgr. All mgr pods have a sidecar
that will update their respective pod specs with the active
or passive label.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-09-14 13:29:11 -06:00
Matthew Penner d09a43df37 exporter: bind to all interfaces if IPv6 is enabled
Previously ceph-exporter would only bind to IPv4 interfaces. Now if the
CephCluster is configured with `dualStack: true` and/or `ipFamily: IPv6`
an additional flag (`--addrs ::`) will be added to the ceph-exporter
container to make it listen both IPv6 and IPv4 interfaces.

Signed-off-by: Matthew Penner <me@matthewp.io>
2023-09-12 11:39:46 -06:00
Blaine Gardner 17f0072d9d Merge pull request #12778 from BlaineEXE/multus-allow-cidr-spec
multus: allow using NADs without inspectable CIDRs
2023-09-07 13:42:41 -06:00
Blaine Gardner 3c43268d0a multus: detect network CIDRs via canary
Change how Rook detects network CIDRs for Multus networks. The IPAM
configuration is only defined as an arbitrary string JSON blob with a
"type" field and nothing more. Rook's detection of CIDRs for whereabouts
had already grown out of date since the initial implementation.
Additionally, Rook did not support DHCP IPAM, which is a reasonable
choice for users. And more, Rook did not support CNI plugin chaining,
which further complicates NADs. Based on the CNI spec, network chaning
can result in any changes to network CIDRs from the first-given plugin.

All these problems make it more and more difficult for Rook to support
Multus by inspecting the NAD itself to predict network CIDRs. Instead,
it is better for Rook to treat the CNI process as a black box. To
preserve legacy functionality of auto-detecting networks and to make
that as robust as possible, change to a canary-style architecture like
that used for Ceph mons, from which Rook will detect the network CIDRs
if possible.

Also allow users to specify overrides for CIDR ranges. This allows Rook
to still support esoteric and unexpected NAD or network configurations
where a CIDR range is not detectable or where the range detected would
be incomplete. Because it may be impossible for Rook to understand the
network CIDRs wholistically while residing only on a portion of the
network, this feature should have been present from Multus's inception.

Improving CIDR auto-detection and allowing users to specify overrides
for auto-detected CIDRs rounds out Rook's Multus support for CephCluster
(core/RADOS) installations. No further architectural changes should be
needed for CephClusters as regards application of public/cluster network
CIDRs for Multus networks.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-09-07 10:12:55 -06:00
Travis Nielsen 0ae3800929 Merge pull request #12770 from sp98/osd-migration-part2
osd: replace existing OSDs to use new store
2023-09-05 08:57:20 -06:00
Travis Nielsen ded675e758 Merge pull request #12825 from weirdwiz/service-monitor-port
monitoring: set port for servicemonitor for ceph-exporter
2023-09-01 11:58:41 -06:00
sp98 11b8d10a5a osd: replace existing OSDs to use new store
This follow up the #12507
- fixes replacing of encrypted OSDs.
- Updates OSD status at the end of reconcile

Signed-off-by: sp98 <sapillai@redhat.com>
2023-09-01 21:07:38 +05:30