This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.
Signed-off-by: subhamkrai <srai@redhat.com>
The upgrade test should always upgrade from the previous
minor release to the latest master. With 1.13 releasing
soon, now we uprade from 1.12 to master, to confirm
if there are any upgrade issues to 1.13.
Signed-off-by: travisn <tnielsen@redhat.com>
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition
Signed-off-by: parth-gr <paarora@redhat.com>
The multicluster test was using a different ceph image for the
toolbox than from the original cluster. For efficiency of
pulling images in the test, the toolbox of the external cluster
will now use the same ceph image as the internal cluster.
Signed-off-by: travisn <tnielsen@redhat.com>
Current reason for failure
clients: "cephbuckettopic" "my-topic" exist, but ARN was not set
Added some more re-try as it might take some more time to
update the ARN in status
OBC bound status was failing and going into
race condition adding a seprate timeout
makes recover it from that state
Signed-off-by: parth-gr <paarora@redhat.com>
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
With the upcoming v1.12 release, Rook upgrade tests
should upgrade from v1.11.x to master instead of
from v1.10.x to master.
Signed-off-by: Sheetal Pamecha <spamecha@redhat.com>
When msgr2 is required, ensure the mon endpoints passed to
the csi configmap are on port 3300. The mons will be
listening on both 6789 and 3300. Existing volumes can
continue using port 6789 while new volumes will use
port 3300.
Signed-off-by: travisn <tnielsen@redhat.com>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
Enabling msgr v2 and disabling msgr v1 currently requires enabling
either encryption on the wire or compression on the wire.
As more clients are running on the latest kernel, allow
the clients to run on v2 even when encryption and compression
are not enabled. Clusters that are fully running on v2
will more easily be able to change configuration between
enabling or disabling msgr v2 features.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the v1.11 release approaching, the upgrade tests in
master are not updated to upgrade starting from v1.10
instead of from v1.9.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.
This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.
The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.
Full discussion: https://github.com/rook/rook/issues/11031
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
ci is more resource intensive.
Signed-off-by: Rakshith R <rar@redhat.com>
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
Rook will support the most recent seven K8s releases,
so we update the min version to 1.19 for Rook v1.10.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the pending release of 1.10, the upgrade tests
in master should be testing the upgrade from 1.9
to the master branch.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Implement the SSSD sidecar design for NFS.
Because NFS documents are getting long, create an NFS
`Storage-Configuration` section with separated topics to keep the `NFS
Overview` document sane.
Adjust the SSSD design to allow any VolumeSource, not just ConfigMaps.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The latest go modules come with a sync mutex that requires
us to pass the test suite by reference for proper use
of the mutex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The file deletion status has been intermittently failing for several
months. Investigation shows that the timeout is just missing by
a few seconds. Now we increase the timeout from 15 to 45 seconds
to stabilize the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The cluster cleanup is tested in all of the test suites. Now that the
correct manifest test installer is being used for cleanup, the
upgrade tests are failing to delete the namespace due to some
finalizer that is holding on for some resource that wasn't cleaned
up properly. Let's just disable the cleanup for these tests since
it's tested in other test suites and not worth the time to
troubleshoot every resource getting cleaned up.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The uninstall after the upgrade test was attempting to read
the crds from github, instead of the local repo as expected.
Now the correct manifests interface is updated during
the upgrade so it will be used during the uninstall.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.
One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.
Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.
A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.
Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
adding e2e test for webhook validation to test invalid
cluster, pool,objectStore, radosNamespace, subVolumeGroup CR
that will each be rejected by their admission controller.
Closes: https://github.com/rook/rook/issues/10024
Signed-off-by: subhamkrai <srai@redhat.com>
Check the radosgw-admin realm user list per object store instead of relying
on the ceph dashboard get-rgw-api- command.
Resolves#9099
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Support is removed for k8s for various limitations such as
priority classes not working and csi driver feature
incompleteness and missing snapshots. Documentation is updated
with the new min version of k8s 1.17 and also the tests are
updated to run on the min version of 1.17.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
The mon failover test was failing if the cluster reconcile
was being triggered in the middle of the failover and if
the ordering of the mons happened to attempt to start the
mons that were not failed over. The other mons couldn't be updated
due to the upgrade checks, so the test would timeout waiting
for the mons that would never update in time.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
finally, admission controller will be enabled default
without any script/manual step. But it still requires cert-manager
to be installed which I believe is already installed in clusters.
**Note**
Code doesn't return error it just logs the error since
we don't want to stop reconciling if the admission controller fails.
We can work on this once the admission controller is stable.
Signed-off-by: subhamkrai <srai@redhat.com>