Commit Graph
428 Commits
Author SHA1 Message Date
subhamkrai 28cc1ebc55 core: remove webhook & controller-runtime from apis
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-01 14:15:40 +05:30
travisn 6c16c0eb83 tests: upgrade from 1.12 to master
The upgrade test should always upgrade from the previous
minor release to the latest master. With 1.13 releasing
soon, now we uprade from 1.12 to master, to confirm
if there are any upgrade issues to 1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-15 16:11:14 -07:00
travisn 03d077aa6b core: remove support for ceph pacific
Pacific is end of life and no longer necessary to
support in Rook with v1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-14 17:07:03 -07:00
parth-gr 46c241433d object: improve the error handling for multisite objs
here is the https://go.dev/play/p/SS9Q-dAiIx3 example which says the
error handling was wrongly implemented

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-14 15:12:15 +05:30
parth-gr 40295a989c ci: fix objectsuite flakiness
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-08 15:33:03 +05:30
Jiffin Tony Thottan a941b3c33f object: create cosi user for each object store
Create each cosi user for each object store and secret which holds
credentials.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-09-19 13:48:58 +05:30
travisn 0bd7d4ac84 tests: use same ceph version for toolbox in multicluster
The multicluster test was using a different ceph image for the
toolbox than from the original cluster. For efficiency of
pulling images in the test, the toolbox of the external cluster
will now use the same ceph image as the internal cluster.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-08-16 15:50:57 -06:00
parth-gr 715bd31a0a ci: fix testobjectsuite ci
Current reason for failure
clients: "cephbuckettopic" "my-topic" exist, but ARN was not set

Added some more re-try as it might take some more time to
update the ARN in status

OBC bound status was failing and going into
race condition adding a seprate timeout
makes recover it from that state

Signed-off-by: parth-gr <paarora@redhat.com>
2023-08-08 23:46:02 +05:30
Jiffin Tony Thottan b48dc8a335 object: intial cosi driver controller design
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-07-18 22:49:41 +05:30
Javier 46cb9435a6 test: support Ceph Reef v18 with updated integration tests
update test files to support the next version of ceph, ceph reef v18

Signed-off-by: Javier <sjavierlopez@gmail.com>
2023-07-10 15:24:36 -06:00
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
Sheetal Pamecha a6d57c164d test: upgrade tests to update from v1.11
With the upcoming v1.12 release, Rook upgrade tests
should upgrade from v1.11.x to master instead of
from v1.10.x to master.

Signed-off-by: Sheetal Pamecha <spamecha@redhat.com>
2023-06-13 03:48:45 +05:30
travisn 0484a3973f csi: update port to 3300 if msgr2 is required
When msgr2 is required, ensure the mon endpoints passed to
the csi configmap are on port 3300. The mons will be
listening on both 6789 and 3300. Existing volumes can
continue using port 6789 while new volumes will use
port 3300.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-03-10 10:34:56 -07:00
parth-grandBlaine Gardner 69ae9569cd build: update k8s version to 1.26.1
Update the Kubernetes API version used to 1.26.1, and start testing
against Kubernetes version 1.26.1 in CI.

Co-authored-by: parth-gr <paarora@redhat.com>
Co-authored-by: Blaine Gardner <blaine.gardner@redhat.com>

Signed-off-by: parth-gr <paarora@redhat.com>
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2023-02-23 20:18:47 -07:00
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
Travis Nielsen a4f71baf8a core: option to require msgrv2 even without encryption
Enabling msgr v2 and disabling msgr v1 currently requires enabling
either encryption on the wire or compression on the wire.
As more clients are running on the latest kernel, allow
the clients to run on v2 even when encryption and compression
are not enabled. Clusters that are fully running on v2
will more easily be able to change configuration between
enabling or disabling msgr v2 features.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-10 09:45:57 -07:00
Travis Nielsen 230e635611 test: upgrade tests to update from v1.10
With the v1.11 release approaching, the upgrade tests in
master are not updated to upgrade starting from v1.10
instead of from v1.9.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-06 16:50:07 -07:00
Blaine Gardner a777b1d7d1 object: do not create service for external object stores
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.

This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.

The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-04 17:30:30 -06:00
Blaine Gardner 0937eaee38 Merge pull request #11124 from BlaineEXE/object-revise-health-check
object: remove health checker
2022-10-24 11:31:23 -06:00
Blaine Gardner a7c0c7ee93 object: remove health checker
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.

Full discussion: https://github.com/rook/rook/issues/11031

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-18 13:41:46 -06:00
Rakshith R 914e01e77d ci: add nfs clone e2e testcase
Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-18 15:07:01 +05:30
Rakshith R db56ba22d2 ci: add nfs snap & restore e2e testcases
Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-18 15:06:58 +05:30
Rakshith R be10dac98c ci: enable and fixes for nfs ci
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
  ci is more resource intensive.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-06 11:53:46 +00:00
Travis Nielsen 25ee33f029 ci: ceph master images renamed to main
The ceph master images were recently renamed to main
so we need to pick up the new tag.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-10-05 16:46:04 -06:00
Liang Zheng f9c360e609 osd: optimize device probe
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 10:41:51 +08:00
Travis Nielsen a6dfb77827 build: update min version to k8s 1.19
Rook will support the most recent seven K8s releases,
so we update the min version to 1.19 for Rook v1.10.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-30 15:15:20 -06:00
Travis Nielsen 3696a560d8 build: test upgrade from 1.9 to master
With the pending release of 1.10, the upgrade tests
in master should be testing the upgrade from 1.9
to the master branch.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-26 10:04:37 -06:00
Blaine Gardner b98efbb320 nfs: add support for running SSSD as a sidecar
Implement the SSSD sidecar design for NFS.

Because NFS documents are getting long, create an NFS
`Storage-Configuration` section with separated topics to keep the `NFS
Overview` document sane.

Adjust the SSSD design to allow any VolumeSource, not just ConfigMaps.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-08-11 14:59:23 -06:00
Travis Nielsen a596471757 test: pass test suite by value
The latest go modules come with a sync mutex that requires
us to pass the test suite by reference for proper use
of the mutex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-11 12:31:49 -06:00
Travis Nielsen 13d527b590 test: increase timeout waiting for file deletion status
The file deletion status has been intermittently failing for several
months. Investigation shows that the timeout is just missing by
a few seconds. Now we increase the timeout from 15 to 45 seconds
to stabilize the test.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-02 13:14:05 -06:00
Travis Nielsen dad97f3425 core: remove support for ceph octopus
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-07 15:03:26 -06:00
Rakshith R 034c4756d3 ci: make sure PVC is deleted before proceeding for file test
Signed-off-by: Rakshith R <rar@redhat.com>
2022-07-04 11:36:37 +05:30
Rakshith R 16d34e1425 ci: skip nfs csi e2e to avoid flaky e2e
Skip nfs csi e2e until https://github.com/rook/rook/issues/10518
is resolved to avoid flakiness.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-07-04 11:36:37 +05:30
Travis Nielsen e30a068e68 test: skip cluster cleanup in upgrade test
The cluster cleanup is tested in all of the test suites. Now that the
correct manifest test installer is being used for cleanup, the
upgrade tests are failing to delete the namespace due to some
finalizer that is holding on for some resource that wasn't cleaned
up properly. Let's just disable the cleanup for these tests since
it's tested in other test suites and not worth the time to
troubleshoot every resource getting cleaned up.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-06-30 16:19:04 -06:00
Travis Nielsen e86b2f8b02 test: upgrade test to use correct manifests for uninstall
The uninstall after the upgrade test was attempting to read
the crds from github, instead of the local repo as expected.
Now the correct manifests interface is updated during
the upgrade so it will be used during the uninstall.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-06-30 15:33:31 -06:00
Travis Nielsen 8e65d69c7a Merge pull request #10495 from Rakshith-R/nfs-ci-cleanup
ci: rename EnableCSINFS to TestNFSCSI & explain "nolock" mountOption
2022-06-29 07:59:40 -06:00
Markus Schmidleitner 00689476df test: delete object sc before upgrade tests (#10153)
Signed-off-by: Markus Schmidleitner <markus.schmidleitner@ocilion.com>
2022-06-28 10:39:58 +02:00
Rakshith R 928dd46593 ci: rename TestCephSettings.EnableCSINFS to TestNFSCSI
Signed-off-by: Rakshith R <rar@redhat.com>
2022-06-23 16:48:13 +05:30
Rakshith R 5b017cfd67 ci: add tests for nfs csi pvc
This commit adds nfs csi pvc test into
ceph smoke suite.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-06-20 15:01:22 +05:30
Blaine Gardner 800d3e5050 Merge pull request #9915 from BlaineEXE/dependents-cephfilesystem
file: block deletion on more dependents
2022-06-17 14:07:07 -06:00
Blaine Gardner a63844f8bf file: block deletion on more dependents
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.

One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.

Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.

A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.

Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-06-14 13:02:35 -06:00
subhamkrai 4d92c2770a test: add e2e test for webhook validation
adding e2e test for webhook validation to test invalid
cluster, pool,objectStore, radosNamespace, subVolumeGroup CR
that will each be rejected by their admission controller.

Closes: https://github.com/rook/rook/issues/10024
Signed-off-by: subhamkrai <srai@redhat.com>
2022-06-07 14:01:02 +05:30
Alexander Trost 3005a6fc68 Merge pull request #10137 from koor-tech/fix_9099
rgw: fix dashboard admin creation for multiple object stores
2022-04-27 16:13:38 +00:00
Alexander Trost 5d03061d47 rgw: fix dashboard admin creation for multiple object stores
Check the radosgw-admin realm user list per object store instead of relying
on the ceph dashboard get-rgw-api- command.

Resolves #9099

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-27 14:57:17 +02:00
Travis Nielsen d26b6cf403 build: update min version to k8s 1.17
Support is removed for k8s for various limitations such as
priority classes not working and csi driver feature
incompleteness and missing snapshots. Documentation is updated
with the new min version of k8s 1.17 and also the tests are
updated to run on the min version of 1.17.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-25 15:32:09 -06:00
Alexander Trost 8686296e17 core: remove double imported packages
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-25 13:51:45 +02:00
Travis Nielsen f7cefbbbde test: improve reliability of mon failover test
The mon failover test was failing if the cluster reconcile
was being triggered in the middle of the failover and if
the ordering of the mons happened to attempt to start the
mons that were not failed over. The other mons couldn't be updated
due to the upgrade checks, so the test would timeout waiting
for the mons that would never update in time.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-20 13:45:09 -06:00
Travis Nielsen fb86955f01 core: examples set default priority class names
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-19 16:09:34 -06:00
subhamkrai f6f03d272b core: start admission controller without any script
finally, admission controller will be enabled default
without any script/manual step. But it still requires cert-manager
to be installed which I believe is already installed in clusters.

**Note**
Code doesn't return error it just logs the error since
we don't want to stop reconciling if the admission controller fails.
We can work on this once the admission controller is stable.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-11 19:44:54 +05:30
subhamkrai 24802c559e core: fix golangci linter
fix golangci linter

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-04 20:59:31 +05:30