Commit Graph
720 Commits
Author SHA1 Message Date
subhamkrai d8766ff871 ci: remove ceph SVGs in helm test
Both the helm tests are failing because,
```
2023-12-19 08:47:06.136640 E | ceph-file-controller: failed to reconcile CephFilesystem "helm-ns/ceph-filesystem-test". CephFilesystem "helm-ns/ceph-filesystem-test" will not be deleted until all dependents are removed: CephFilesystemSubVolumeGroups: [ceph-filesystem-test-csi]
```
So, let's remove the ceph SVGs before removing Ceph filesystem.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-19 17:05:21 +05:30
parth-gr 60e879050a ci: delete svg in helm test
The filesystem is not being deleted because of the existing svg.
Add a cliet call to delete the deafult csi svg,
for helm test

Signed-off-by: parth-gr <partharora1010@gmail.com>
2023-12-12 00:52:19 +05:30
Anthony D'Atri 59d0240676 doc: improve ceph-csi-drivers.md and lintrolling
Signed-off-by: Anthony D'Atri <anthonyeleven@users.noreply.github.com>
2023-12-01 16:10:11 -07:00
subhamkrai 28cc1ebc55 core: remove webhook & controller-runtime from apis
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-01 14:15:40 +05:30
travisn 6c16c0eb83 tests: upgrade from 1.12 to master
The upgrade test should always upgrade from the previous
minor release to the latest master. With 1.13 releasing
soon, now we uprade from 1.12 to master, to confirm
if there are any upgrade issues to 1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-15 16:11:14 -07:00
travisn 03d077aa6b core: remove support for ceph pacific
Pacific is end of life and no longer necessary to
support in Rook with v1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-14 17:07:03 -07:00
parth-gr 46c241433d object: improve the error handling for multisite objs
here is the https://go.dev/play/p/SS9Q-dAiIx3 example which says the
error handling was wrongly implemented

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-14 15:12:15 +05:30
parth-gr 40295a989c ci: fix objectsuite flakiness
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-08 15:33:03 +05:30
Eng Zer Jun 77bff6a07c core: remove redundant len check
From the Go specification [1]:

  "1. For a nil slice, the number of iterations is 0."
  "3. If the map is nil, the number of iterations is 0."

`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.

[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len

Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
2023-10-05 20:16:46 +08:00
Redouane Kachach b3dd74ea20 docs: fixing some spelling issues
closes: https://github.com/rook/rook/issues/12987

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-10-03 13:50:17 +02:00
travisn 90d06a862b tests: start toolbox earlier in tests
The toolbox sometimes times out when the tests are waiting
for it to start. Now we create the toolbox spec sooner in
the tests so we won't need to wait so long for it to start
or increase the wait timeout in the tests.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-08-16 15:30:12 -06:00
subhamkrai 1bd4ab9d81 ci: skip mgr pod restart count upgrade 1.22.x suite
for now, let's skip the mgr pod restart count
for upgrade suite 1.22.x to get the CI green
and so that we don't skip any other error in name
of mgr restart count.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:03:08 +05:30
travisn c17724b536 tests: multi cluster suite toolbox version
The toolbox image tag was always being set from #12625
even when the image was not set in the test. If the image
is not set, skip replacing the image name to use the default
test version.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-07-31 16:18:24 -06:00
travisn 600068ea21 tests: toolbox uses same ceph image as the test
The test is currently only using the ceph version that is
included with toolbox.yaml, which may be different
from the version of the ceph cluster being tested.
Now the version will be replaced to match the desired
ceph test version.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-07-31 13:18:59 -06:00
Blaine Gardner c86f22e4f7 Merge pull request #12578 from subhamkrai/fix-upgrade-suite
test: fix upgrade suite for 1.27.x version
2023-07-26 09:47:43 -06:00
subhamkrai 0ed91458dc test: collect kube-system logs for debugging
added kube-system namespace also to collect
logs from to debug the smoke suite issue

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-26 14:09:15 +05:30
subhamkrai e8e8a4673a test: fix upgrade suite for 1.27.x version
the upgrade suite for version 1.27.x is failing
due
```
debug 2023-07-24T19:11:20.264+0000 7f3395a82c80 -1 error: monitor data filesystem reached concerning levels of available storage space (available: 5% 4.4 GiB)
you may adjust 'mon data avail crit' to a lower value to make this go away (default: 5%)
```
so adding mon setting to start on `compact` also adding
other setting present in cluster-test.yaml.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-26 12:14:34 +05:30
Travis Nielsen 91fea5395e Merge pull request #12415 from thotz/ceph-cosi-driver
object: adding ceph cosi driver
2023-07-18 12:12:12 -06:00
Jiffin Tony Thottan b48dc8a335 object: intial cosi driver controller design
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-07-18 22:49:41 +05:30
Javier d17109a5f5 test: update reef as default version
update ceph to reef as default version

Signed-off-by: Javier <sjavierlopez@gmail.com>
2023-07-17 12:45:28 -06:00
Javier 46cb9435a6 test: support Ceph Reef v18 with updated integration tests
update test files to support the next version of ceph, ceph reef v18

Signed-off-by: Javier <sjavierlopez@gmail.com>
2023-07-10 15:24:36 -06:00
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
Sheetal Pamecha a6d57c164d test: upgrade tests to update from v1.11
With the upcoming v1.12 release, Rook upgrade tests
should upgrade from v1.11.x to master instead of
from v1.10.x to master.

Signed-off-by: Sheetal Pamecha <spamecha@redhat.com>
2023-06-13 03:48:45 +05:30
travisn db04b127b0 core: remove obsolete journal size config value
The journal size was only applicable to the filestore OSD
format which has not been supported by rook since v1.2.
Remove the remaining obsolete setting from the examples and
code.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-04-24 15:54:53 -06:00
Rakshith R cc72697f2a ci: increase wait time for external cluster to be ready
Signed-off-by: Rakshith R <rar@redhat.com>
2023-03-10 12:26:53 +05:30
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
Travis Nielsen 88f66ff1f2 test: fix upgrade test for msgr2 setting
The requireMsgr2 setting is not available in 1.10, therefore
the setting must be removed from the cephcluster CR
when the 1.10 cluster is created before the upgrade.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-14 17:23:38 -07:00
Travis Nielsen a4f71baf8a core: option to require msgrv2 even without encryption
Enabling msgr v2 and disabling msgr v1 currently requires enabling
either encryption on the wire or compression on the wire.
As more clients are running on the latest kernel, allow
the clients to run on v2 even when encryption and compression
are not enabled. Clusters that are fully running on v2
will more easily be able to change configuration between
enabling or disabling msgr v2 features.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-10 09:45:57 -07:00
Travis Nielsen 230e635611 test: upgrade tests to update from v1.10
With the v1.11 release approaching, the upgrade tests in
master are not updated to upgrade starting from v1.10
instead of from v1.9.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-06 16:50:07 -07:00
subhamkrai e9e2313126 webhook: disable webhook by default
sadly, we need to disable webhook by default until
we find solution to https://github.com/rook/rook/issues/10719

Signed-off-by: subhamkrai <srai@redhat.com>
2023-01-12 22:08:56 +05:30
Madhu Rajanna 41e4082730 csi: use krbd for encryption
use krbd for mapping the rbd image when
ceph cluster encryption is enabled.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2023-01-12 10:02:15 +01:00
Loïc Antoine-GombeaudandTravis Nielsen 59f6cd00ff helm: fix toolbox command, default to cluster image
Signed-off-by: Loïc Antoine-Gombeaud <lantoinegombeaud@gestform.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2023-01-04 14:23:37 -07:00
Travis Nielsen 25ac236b5c core: remove support for machine disruption budgets
The machine disruption budgets for handling openshift
machines and machinesets are now removed since they
have been unused and unmaintained since implemented.
This feature is expected to be handled with the more
common Pod Disruption Budgets. A workaround is for the
cluster admin to set up their machine sets so they
match the zone topology. See the original design
doc from the feature here:
https://github.com/rook/rook/blob/master/design/ceph/ceph-openshift-fencing-mitigation.md

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-16 11:15:12 -07:00
parth-gr 9319eaa86d ci: watch for pods that restart unexpectedly
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives

Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
2022-12-14 21:50:15 +05:30
Travis Nielsen b693a5cca5 core: refactor crash collector for more node daemons
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md

Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-09 16:57:05 -07:00
Blaine Gardner fc56108865 object: allow status endpoints to be null
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.

This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-08 13:36:15 -07:00
Blaine Gardner 0937eaee38 Merge pull request #11124 from BlaineEXE/object-revise-health-check
object: remove health checker
2022-10-24 11:31:23 -06:00
Travis Nielsen d695a0f930 Merge pull request #11163 from gauravsitlani/gdb-coredump-tests
core: enabling logCollector by default for coredump collection
2022-10-19 12:34:54 -06:00
gauravsitlani cc02399c4c core: enabling logCollector by default for coredump
enabling logCollector by default will enable the coredump
generation in case a process terminates with a segmentation fault.

Signed-off-by: gauravsitlani <gauravsitlani@riseup.net>
2022-10-19 23:04:22 +05:30
Blaine Gardner a7c0c7ee93 object: remove health checker
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.

Full discussion: https://github.com/rook/rook/issues/11031

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-18 13:41:46 -06:00
Rakshith R db56ba22d2 ci: add nfs snap & restore e2e testcases
Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-18 15:06:58 +05:30
Rakshith R be10dac98c ci: enable and fixes for nfs ci
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
  ci is more resource intensive.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-06 11:53:46 +00:00
Travis Nielsen 25ee33f029 ci: ceph master images renamed to main
The ceph master images were recently renamed to main
so we need to pick up the new tag.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-10-05 16:46:04 -06:00
Liang Zheng 9a78f6af07 osd: update ceph status parse
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 12:48:53 +08:00
Liang Zheng f9c360e609 osd: optimize device probe
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 10:41:51 +08:00
yati1998 9b4361b379 rbdmirror: remove volume replication sidecar
The volume replication operator is being moved to
kubernetes-csi-addons. This commit hence, removes
the volume replication sidecar and updates the
related documentation.
It will also update the csi-addons sidecar version
to the latest one.

Closes: #10655

Signed-off-by: yati1998 <ypadia@redhat.com>
2022-09-07 14:49:19 +05:30
Travis Nielsen 5c478e5e04 test: wait for resource deletion during helm uninstall
The helm integration test was immediately removing both the cluster
and the operator charts, without waiting for any resource deletion
between the deletions. This causes intermittent failures in the
test since the cluster CR and mon secret and configmap to never
be deleted because the operator will be stopped before their
finalizers can be removed. Now we wait for their finalizers
to be removed before removing the operator chart.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-09-02 09:37:47 -06:00
Travis Nielsen a6dfb77827 build: update min version to k8s 1.19
Rook will support the most recent seven K8s releases,
so we update the min version to 1.19 for Rook v1.10.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-30 15:15:20 -06:00
Travis Nielsen 4028b50de5 test: remove obsolete gateway type setting
Long ago the gateway type for s3 was removed from the object
store CRD, and k8s 1.25 is now complaining about the obsolete
setting still included in the test.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-29 16:44:09 -06:00
Travis Nielsen 0cd9361b3e test: upgrade from v1.9.10 to pick up psp changes
K8s 1.25 is not capable of creating PSPs, so we need to pick
up v1.9.10 as the base for the upgrade test before upgrading
to master.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-29 15:50:42 -06:00