Commit Graph
696 Commits
Author SHA1 Message Date
Rakshith R cc72697f2a ci: increase wait time for external cluster to be ready
Signed-off-by: Rakshith R <rar@redhat.com>
2023-03-10 12:26:53 +05:30
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
Travis Nielsen 88f66ff1f2 test: fix upgrade test for msgr2 setting
The requireMsgr2 setting is not available in 1.10, therefore
the setting must be removed from the cephcluster CR
when the 1.10 cluster is created before the upgrade.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-14 17:23:38 -07:00
Travis Nielsen a4f71baf8a core: option to require msgrv2 even without encryption
Enabling msgr v2 and disabling msgr v1 currently requires enabling
either encryption on the wire or compression on the wire.
As more clients are running on the latest kernel, allow
the clients to run on v2 even when encryption and compression
are not enabled. Clusters that are fully running on v2
will more easily be able to change configuration between
enabling or disabling msgr v2 features.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-10 09:45:57 -07:00
Travis Nielsen 230e635611 test: upgrade tests to update from v1.10
With the v1.11 release approaching, the upgrade tests in
master are not updated to upgrade starting from v1.10
instead of from v1.9.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-06 16:50:07 -07:00
subhamkrai e9e2313126 webhook: disable webhook by default
sadly, we need to disable webhook by default until
we find solution to https://github.com/rook/rook/issues/10719

Signed-off-by: subhamkrai <srai@redhat.com>
2023-01-12 22:08:56 +05:30
Madhu Rajanna 41e4082730 csi: use krbd for encryption
use krbd for mapping the rbd image when
ceph cluster encryption is enabled.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2023-01-12 10:02:15 +01:00
Loïc Antoine-GombeaudandTravis Nielsen 59f6cd00ff helm: fix toolbox command, default to cluster image
Signed-off-by: Loïc Antoine-Gombeaud <lantoinegombeaud@gestform.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2023-01-04 14:23:37 -07:00
Travis Nielsen 25ac236b5c core: remove support for machine disruption budgets
The machine disruption budgets for handling openshift
machines and machinesets are now removed since they
have been unused and unmaintained since implemented.
This feature is expected to be handled with the more
common Pod Disruption Budgets. A workaround is for the
cluster admin to set up their machine sets so they
match the zone topology. See the original design
doc from the feature here:
https://github.com/rook/rook/blob/master/design/ceph/ceph-openshift-fencing-mitigation.md

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-16 11:15:12 -07:00
parth-gr 9319eaa86d ci: watch for pods that restart unexpectedly
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives

Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
2022-12-14 21:50:15 +05:30
Travis Nielsen b693a5cca5 core: refactor crash collector for more node daemons
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md

Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-09 16:57:05 -07:00
Blaine Gardner fc56108865 object: allow status endpoints to be null
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.

This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-08 13:36:15 -07:00
Blaine Gardner 0937eaee38 Merge pull request #11124 from BlaineEXE/object-revise-health-check
object: remove health checker
2022-10-24 11:31:23 -06:00
Travis Nielsen d695a0f930 Merge pull request #11163 from gauravsitlani/gdb-coredump-tests
core: enabling logCollector by default for coredump collection
2022-10-19 12:34:54 -06:00
gauravsitlani cc02399c4c core: enabling logCollector by default for coredump
enabling logCollector by default will enable the coredump
generation in case a process terminates with a segmentation fault.

Signed-off-by: gauravsitlani <gauravsitlani@riseup.net>
2022-10-19 23:04:22 +05:30
Blaine Gardner a7c0c7ee93 object: remove health checker
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.

Full discussion: https://github.com/rook/rook/issues/11031

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-18 13:41:46 -06:00
Rakshith R db56ba22d2 ci: add nfs snap & restore e2e testcases
Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-18 15:06:58 +05:30
Rakshith R be10dac98c ci: enable and fixes for nfs ci
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
  ci is more resource intensive.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-06 11:53:46 +00:00
Travis Nielsen 25ee33f029 ci: ceph master images renamed to main
The ceph master images were recently renamed to main
so we need to pick up the new tag.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-10-05 16:46:04 -06:00
Liang Zheng 9a78f6af07 osd: update ceph status parse
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 12:48:53 +08:00
Liang Zheng f9c360e609 osd: optimize device probe
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 10:41:51 +08:00
yati1998 9b4361b379 rbdmirror: remove volume replication sidecar
The volume replication operator is being moved to
kubernetes-csi-addons. This commit hence, removes
the volume replication sidecar and updates the
related documentation.
It will also update the csi-addons sidecar version
to the latest one.

Closes: #10655

Signed-off-by: yati1998 <ypadia@redhat.com>
2022-09-07 14:49:19 +05:30
Travis Nielsen 5c478e5e04 test: wait for resource deletion during helm uninstall
The helm integration test was immediately removing both the cluster
and the operator charts, without waiting for any resource deletion
between the deletions. This causes intermittent failures in the
test since the cluster CR and mon secret and configmap to never
be deleted because the operator will be stopped before their
finalizers can be removed. Now we wait for their finalizers
to be removed before removing the operator chart.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-09-02 09:37:47 -06:00
Travis Nielsen a6dfb77827 build: update min version to k8s 1.19
Rook will support the most recent seven K8s releases,
so we update the min version to 1.19 for Rook v1.10.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-30 15:15:20 -06:00
Travis Nielsen 4028b50de5 test: remove obsolete gateway type setting
Long ago the gateway type for s3 was removed from the object
store CRD, and k8s 1.25 is now complaining about the obsolete
setting still included in the test.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-29 16:44:09 -06:00
Travis Nielsen 0cd9361b3e test: upgrade from v1.9.10 to pick up psp changes
K8s 1.25 is not capable of creating PSPs, so we need to pick
up v1.9.10 as the base for the upgrade test before upgrading
to master.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-29 15:50:42 -06:00
Travis Nielsen 3696a560d8 build: test upgrade from 1.9 to master
With the pending release of 1.10, the upgrade tests
in master should be testing the upgrade from 1.9
to the master branch.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-26 10:04:37 -06:00
subhamkrai d1e3023439 ci: collect operator logs in debug mode
collect operator logs in debug mode for both
canary tests and integration tests.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-08-26 19:32:59 +05:30
Alexander Trost 49440e8b20 test: fix test pod log collector
When a test has a slash in it's name the log collection can fail due to
the "directory" not existing, this makes sure the slashes are replaced
by underscores.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-08-12 23:44:13 +02:00
subhamkrai c4742278d1 Revert "ci: use ceph v17.2.1 instead of latest v17"
This reverts commit 3279cb91870e59eccabb9b5eff903d71beb31594.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-08-01 21:05:48 +05:30
subhamkrai 4501cc25c6 ci: use ceph v17.2.1 instead of latest v17
there could be something breaking with latest ceph.
So, let's stick to v17.2.1.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-07-25 20:14:52 +05:30
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Travis Nielsen dad97f3425 core: remove support for ceph octopus
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-07 15:03:26 -06:00
Rakshith R 034c4756d3 ci: make sure PVC is deleted before proceeding for file test
Signed-off-by: Rakshith R <rar@redhat.com>
2022-07-04 11:36:37 +05:30
Rakshith R 621dd2a159 ci: explain why "nolock" mountOption is needed for nfs
"nolock" mountOptions has been added to prevent the following error in ci:
Output: mount.nfs: rpc.statd is not running but is required for remote locking.
mount.nfs: Either use '-o nolock' to keep locks local, or start statd.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-06-23 16:53:08 +05:30
Rakshith R 928dd46593 ci: rename TestCephSettings.EnableCSINFS to TestNFSCSI
Signed-off-by: Rakshith R <rar@redhat.com>
2022-06-23 16:48:13 +05:30
Rakshith R 5b017cfd67 ci: add tests for nfs csi pvc
This commit adds nfs csi pvc test into
ceph smoke suite.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-06-20 15:01:22 +05:30
Blaine Gardner 800d3e5050 Merge pull request #9915 from BlaineEXE/dependents-cephfilesystem
file: block deletion on more dependents
2022-06-17 14:07:07 -06:00
Blaine Gardner a63844f8bf file: block deletion on more dependents
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.

One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.

Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.

A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.

Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-06-14 13:02:35 -06:00
Madhu Rajanna 767cfcbf09 csi: update sidecar to latest release
updating node-driver-registrar to v2.5.1 release
, nfsplugin to 4.0.0 and csi-snapshotter to v6.0.1

Note:- with snapshotter v6.0.1 v1beta1 snapshot CRD
is not supported. If you want betav1 snapshot support
please switch to v5.x.x snapshotter.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-06-07 19:27:13 +05:30
Travis Nielsen 8b334685a9 Merge pull request #10123 from travisn/min-version-1.17
build: Update min version to k8s 1.17
2022-04-26 12:24:04 -06:00
Travis Nielsen d26b6cf403 build: update min version to k8s 1.17
Support is removed for k8s for various limitations such as
priority classes not working and csi driver feature
incompleteness and missing snapshots. Documentation is updated
with the new min version of k8s 1.17 and also the tests are
updated to run on the min version of 1.17.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-25 15:32:09 -06:00
Travis Nielsen 2db3910052 test: remove dead test helper code
The K8sHelper has a number of methods that are no longer
in use that we can remove.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-25 15:28:29 -06:00
Travis Nielsen 23c29b4e28 test: skip upgrade checks during integration tests
The upgrade checks are not necessary during the CI.
The flag had previously been set, but was changed
unintentionally with https://github.com/rook/rook/pull/10096.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-20 14:12:27 -06:00
Travis Nielsen fb86955f01 core: examples set default priority class names
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-19 16:09:34 -06:00
subhamkrai f6f03d272b core: start admission controller without any script
finally, admission controller will be enabled default
without any script/manual step. But it still requires cert-manager
to be installed which I believe is already installed in clusters.

**Note**
Code doesn't return error it just logs the error since
we don't want to stop reconciling if the admission controller fails.
We can work on this once the admission controller is stable.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-11 19:44:54 +05:30
subhamkrai 24802c559e core: fix golangci linter
fix golangci linter

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-04 20:59:31 +05:30
Jiffin Tony Thottan 5e72b26948 object: add service account for RGW pod
For supporting features like service account authentication for vault
KMS , a service account account need to attach with pod.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-04-04 11:24:50 +05:30
Travis Nielsen 6c50530737 Merge pull request #9685 from weirdwiz/context-remove
Remove context.TODO and add context parameter
2022-03-30 12:20:30 -05:00
Jiffin Tony Thottan fc2b8012c6 object: use us-east-1 for aws go lang sdk
The aws go lang sdk needs value for region, it is set differently in
various part of current code. With PR the value is always `us-east-1` so
that it will work RGW server without any issues.

This reverts commit 280c29f330.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-03-24 14:37:39 +05:30