Both the helm tests are failing because,
```
2023-12-19 08:47:06.136640 E | ceph-file-controller: failed to reconcile CephFilesystem "helm-ns/ceph-filesystem-test". CephFilesystem "helm-ns/ceph-filesystem-test" will not be deleted until all dependents are removed: CephFilesystemSubVolumeGroups: [ceph-filesystem-test-csi]
```
So, let's remove the ceph SVGs before removing Ceph filesystem.
Signed-off-by: subhamkrai <srai@redhat.com>
The filesystem is not being deleted because of the existing svg.
Add a cliet call to delete the deafult csi svg,
for helm test
Signed-off-by: parth-gr <partharora1010@gmail.com>
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.
Signed-off-by: subhamkrai <srai@redhat.com>
The upgrade test should always upgrade from the previous
minor release to the latest master. With 1.13 releasing
soon, now we uprade from 1.12 to master, to confirm
if there are any upgrade issues to 1.13.
Signed-off-by: travisn <tnielsen@redhat.com>
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition
Signed-off-by: parth-gr <paarora@redhat.com>
From the Go specification [1]:
"1. For a nil slice, the number of iterations is 0."
"3. If the map is nil, the number of iterations is 0."
`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.
[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len
Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
The toolbox sometimes times out when the tests are waiting
for it to start. Now we create the toolbox spec sooner in
the tests so we won't need to wait so long for it to start
or increase the wait timeout in the tests.
Signed-off-by: travisn <tnielsen@redhat.com>
for now, let's skip the mgr pod restart count
for upgrade suite 1.22.x to get the CI green
and so that we don't skip any other error in name
of mgr restart count.
Signed-off-by: subhamkrai <srai@redhat.com>
The toolbox image tag was always being set from #12625
even when the image was not set in the test. If the image
is not set, skip replacing the image name to use the default
test version.
Signed-off-by: travisn <tnielsen@redhat.com>
The test is currently only using the ceph version that is
included with toolbox.yaml, which may be different
from the version of the ceph cluster being tested.
Now the version will be replaced to match the desired
ceph test version.
Signed-off-by: travisn <tnielsen@redhat.com>
the upgrade suite for version 1.27.x is failing
due
```
debug 2023-07-24T19:11:20.264+0000 7f3395a82c80 -1 error: monitor data filesystem reached concerning levels of available storage space (available: 5% 4.4 GiB)
you may adjust 'mon data avail crit' to a lower value to make this go away (default: 5%)
```
so adding mon setting to start on `compact` also adding
other setting present in cluster-test.yaml.
Signed-off-by: subhamkrai <srai@redhat.com>
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
With the upcoming v1.12 release, Rook upgrade tests
should upgrade from v1.11.x to master instead of
from v1.10.x to master.
Signed-off-by: Sheetal Pamecha <spamecha@redhat.com>
The journal size was only applicable to the filestore OSD
format which has not been supported by rook since v1.2.
Remove the remaining obsolete setting from the examples and
code.
Signed-off-by: travisn <tnielsen@redhat.com>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
The requireMsgr2 setting is not available in 1.10, therefore
the setting must be removed from the cephcluster CR
when the 1.10 cluster is created before the upgrade.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Enabling msgr v2 and disabling msgr v1 currently requires enabling
either encryption on the wire or compression on the wire.
As more clients are running on the latest kernel, allow
the clients to run on v2 even when encryption and compression
are not enabled. Clusters that are fully running on v2
will more easily be able to change configuration between
enabling or disabling msgr v2 features.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the v1.11 release approaching, the upgrade tests in
master are not updated to upgrade starting from v1.10
instead of from v1.9.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The machine disruption budgets for handling openshift
machines and machinesets are now removed since they
have been unused and unmaintained since implemented.
This feature is expected to be handled with the more
common Pod Disruption Budgets. A workaround is for the
cluster admin to set up their machine sets so they
match the zone topology. See the original design
doc from the feature here:
https://github.com/rook/rook/blob/master/design/ceph/ceph-openshift-fencing-mitigation.md
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives
Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md
Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.
This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
enabling logCollector by default will enable the coredump
generation in case a process terminates with a segmentation fault.
Signed-off-by: gauravsitlani <gauravsitlani@riseup.net>
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.
Full discussion: https://github.com/rook/rook/issues/11031
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
ci is more resource intensive.
Signed-off-by: Rakshith R <rar@redhat.com>
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
The volume replication operator is being moved to
kubernetes-csi-addons. This commit hence, removes
the volume replication sidecar and updates the
related documentation.
It will also update the csi-addons sidecar version
to the latest one.
Closes: #10655
Signed-off-by: yati1998 <ypadia@redhat.com>
The helm integration test was immediately removing both the cluster
and the operator charts, without waiting for any resource deletion
between the deletions. This causes intermittent failures in the
test since the cluster CR and mon secret and configmap to never
be deleted because the operator will be stopped before their
finalizers can be removed. Now we wait for their finalizers
to be removed before removing the operator chart.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook will support the most recent seven K8s releases,
so we update the min version to 1.19 for Rook v1.10.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Long ago the gateway type for s3 was removed from the object
store CRD, and k8s 1.25 is now complaining about the obsolete
setting still included in the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
K8s 1.25 is not capable of creating PSPs, so we need to pick
up v1.9.10 as the base for the upgrade test before upgrading
to master.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>