Implement an allow list mechanism that disables potentially unsafe OBC
fields by default. OBC fields beyond `maxObjects` and `maxSize` don't
neatly fit into the OBC framework as it was originally envisioned and
implemented.
Some of the newly added configs could allow users to cause confusion for
themselves. Others might allow users to hijack others buckets. Some
might allow bricking the entire S3 store.
Out of an abundance of safety, allow-list the known-safe options by
default, and require administrators to enable potentially troublesome
options via the new operator-level config
`ROOK_OBC_ALLOW_ADDITIONAL_CONFIG_FIELDS`.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Ceph has changed the image build process to only require a
dockerfile and stop using the ceph-container repo. The
daily images are pushed to the quay.io/ceph-ci/ceph repo,
so the Rook CI will now start using those images.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The upgrade tests in master have been upgrading from 1.14 to
master. In anticipation of the v1.16 release, we change
the upgrade tests to start from v1.15 to test if there
are any regressions in the supported upgrades.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The devel images have been fairly stable for Rook to test
against, but on occasion there are regressions from
Ceph development that affect the Rook CI. For stability during
Rook development, use the latest stable version of ceph for
PRs, master, and release tests. The daily CI will still use
the devel images from Ceph.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The mon canaries may be created even when the mon daemons
are not created thereafter during the integration tests.
Therefore, the integration tests need to also query a label
specific to the mon daemon so the canaries are not a distraction
to the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Given that Ceph Quincy (v17) is past end of life,
remove Quincy from the supported Ceph versions,
examples, and documentation.
Supported versions now include only Reef and Squid.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ROOK_ENFORCE_HOST_NETWORK option was implemented recently
and now we add the helm setting to expose this new setting
in the rook chart.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The docker.io image prefix is expected to be prepended
to the image names in the test images. This was missed
in 14550 related to some CI tests, which was now causing
the CI failures in the 1.15 branch where the search and
replace was missing the new docker.io prefix.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 3045076db8)
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>
* extend the API object specs for swift and keystone integration
* adapt rgw to the new go-ceph version
- The parameter lists of the API call have changes, as parameters
ignored by the RGW Admin Ops API are no longer serialized, therefore
the mock has to be adapted.
- There is now validation for the user keys that are passed to the
User get API, therefore things failed when we had empty keys in our
User proxy object.
* expand the reconcile loop for the swift and keystone integration
* fix minor mistakes in design document
* add env var to pass extra args to minikube
Minikube decides CPU cores and memory automatically based on the
available resources on the machine which may be insufficient to
run rook. This commit adds an environment variable to add arbitrary
arguments to the minikube command, so both can be specified if
desired.
* integration tests for swift and keystone
The new integration of swift or s3 and keystone support by rook
does not have any integration tests yet.
This commit introduces integration tests for swift and keystone. The
tests are done against a minimal keystone setup (keystone container
image from Yaook-project (https://yaook.cloud), sqlite as database
backend, cert-manager and trust-manager for test certificate setup).
To prevent hardcoded credentials, passwords are generated
by the tests. The integration tests use the openstack client
(keystone- and swift-functionality) (https://docs.openstack.org/
python-openstackclient/ latest/). This was a concious design decision
to use client tooling as close as possible to the end user instead of
using other go-libraries (such as gophercloud).
* add documentation on swift and keystone
Currently there is no documentation on the use of Swift to access
an object store as well as the use of OpenStack keystone for
authentication.
This commit adds documentation on the use of Swift and OpenStack
keystone, as well as CRD-related documentation and an example setup.
* add integration tests for S3 via keystone
This commit introduces integration tests for s3 and keystone. The
tests are run against the same minimal keystone setup that the tests
for swift and keystone use.
The integration tests use the aws s3 client to use client tooling as
close as possible to the end user instead of using other go-libraries.
Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
The upgrade tests in master have been upgrading from 1.13 to
master. In anticipation of the v1.15 release, we change
the upgrade tests to start from v1.14 to test if there
are any regressions in the supported upgrades.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The daily ceph upgrade tests were running on Rook v1.13.
This was causing the squid upgrade tests to fail since
1.13 does not support Squid. The purpose of the upgrade
tests is to test the ceph upgrades, therefore the daily
tests will just test them based on rook master.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the release of the first squid RC, we add squid
to the supported versions and add tests to run
Rook against the squid release.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the clusters reach full, nearfull, or backfill full thresholds
ceph will raise health warnings and stop allowing IO or backfill
depending on the threshold. These settings require special ceph
commands instead of being generic ceph config. Allow these settings
to be set from the CephCluster CR in the spec.storage
section.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
cephfs subvolumegroup supports creating svg
with quota and the datapool, This PR adds the
support for the same.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The upgrade test should always upgrade from the
previous minor release to the latest master. With 1.14
releasing soon, now we upgrade from 1.13 to master,
to confirm if there are any upgrade issues to 1.14.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The pg_autoscaler is always enabled by ceph and cannot be disabled.
In a previous PR the pg_autoscaler config was removed from the main example.
Now the remaining config for the pg_autoscaler is removed as well.
Signed-off-by: travisn <tnielsen@redhat.com>
For now we are using the operator namespace name
as the prefix for the csi driver, This PR provides
an option for the users if someone wants to have
their own prefix for the csi driver, if someone tries
to change the prefix for existing csi driver rook
operator will fail to reconcile the csi driver.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Both the helm tests are failing because,
```
2023-12-19 08:47:06.136640 E | ceph-file-controller: failed to reconcile CephFilesystem "helm-ns/ceph-filesystem-test". CephFilesystem "helm-ns/ceph-filesystem-test" will not be deleted until all dependents are removed: CephFilesystemSubVolumeGroups: [ceph-filesystem-test-csi]
```
So, let's remove the ceph SVGs before removing Ceph filesystem.
Signed-off-by: subhamkrai <srai@redhat.com>
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.
Signed-off-by: subhamkrai <srai@redhat.com>
The upgrade test should always upgrade from the previous
minor release to the latest master. With 1.13 releasing
soon, now we uprade from 1.12 to master, to confirm
if there are any upgrade issues to 1.13.
Signed-off-by: travisn <tnielsen@redhat.com>
The toolbox sometimes times out when the tests are waiting
for it to start. Now we create the toolbox spec sooner in
the tests so we won't need to wait so long for it to start
or increase the wait timeout in the tests.
Signed-off-by: travisn <tnielsen@redhat.com>
The toolbox image tag was always being set from #12625
even when the image was not set in the test. If the image
is not set, skip replacing the image name to use the default
test version.
Signed-off-by: travisn <tnielsen@redhat.com>
The test is currently only using the ceph version that is
included with toolbox.yaml, which may be different
from the version of the ceph cluster being tested.
Now the version will be replaced to match the desired
ceph test version.
Signed-off-by: travisn <tnielsen@redhat.com>
the upgrade suite for version 1.27.x is failing
due
```
debug 2023-07-24T19:11:20.264+0000 7f3395a82c80 -1 error: monitor data filesystem reached concerning levels of available storage space (available: 5% 4.4 GiB)
you may adjust 'mon data avail crit' to a lower value to make this go away (default: 5%)
```
so adding mon setting to start on `compact` also adding
other setting present in cluster-test.yaml.
Signed-off-by: subhamkrai <srai@redhat.com>
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
With the upcoming v1.12 release, Rook upgrade tests
should upgrade from v1.11.x to master instead of
from v1.10.x to master.
Signed-off-by: Sheetal Pamecha <spamecha@redhat.com>
The journal size was only applicable to the filestore OSD
format which has not been supported by rook since v1.2.
Remove the remaining obsolete setting from the examples and
code.
Signed-off-by: travisn <tnielsen@redhat.com>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
The requireMsgr2 setting is not available in 1.10, therefore
the setting must be removed from the cephcluster CR
when the 1.10 cluster is created before the upgrade.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Enabling msgr v2 and disabling msgr v1 currently requires enabling
either encryption on the wire or compression on the wire.
As more clients are running on the latest kernel, allow
the clients to run on v2 even when encryption and compression
are not enabled. Clusters that are fully running on v2
will more easily be able to change configuration between
enabling or disabling msgr v2 features.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the v1.11 release approaching, the upgrade tests in
master are not updated to upgrade starting from v1.10
instead of from v1.9.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The machine disruption budgets for handling openshift
machines and machinesets are now removed since they
have been unused and unmaintained since implemented.
This feature is expected to be handled with the more
common Pod Disruption Budgets. A workaround is for the
cluster admin to set up their machine sets so they
match the zone topology. See the original design
doc from the feature here:
https://github.com/rook/rook/blob/master/design/ceph/ceph-openshift-fencing-mitigation.md
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives
Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>