With the tentacle release approaching, let's start
running the Rook tests against the tentacle devel
images to catch if any issues.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Refactored multiple conditional if-else blocks into switch statements to improve readability and maintainability. Also removed outdated staticcheck rule comments (QF1002, QF1003) from .golangci.yaml.
Signed-off-by: Carlos Barria <cbarria@yahoo.com>
The Kubernetes CSI sidecars have had several releases that were not
included in deployments by Rook yet, update them to the versions that
are available today:
- csi-attacher:v4.8.1
- csi-provisioner:v5.2.0
- csi-resizer:v1.13.2
- csi-snapshotter:v8.2.1
This change is important, because Ceph-CSI will implement the new
Controller.GetSnapshot CSI procedure. A bug in csi-lib-utils causes a
panic when a ControllerCapability is provided, but not (yet) known to
the CSI sidecars. The updated sidecars consume a version of
csi-lib-utils with a fix for that panic.
See-also: kubernetes-csi/csi-lib-utils#188
Signed-off-by: Niels de Vos <ndevos@ibm.com>
When upgrading from one Ceph version to another, the new image to
upgrade to can take a long time to pull in some cases before the upgrade
can even begin. For example, some ceph-ci images regularly take 6+
minutes to pull, which exceeds the timeout waiting for mons to be ready.
Extend the timeout for mons to account for these cases.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
this PR updates the Prometheus Operator URL references from
version v0.71.1 to the latest release v0.81.0 in documentation
and integration test scripts. This ensures we are aligned with
the latest features and improvements from
the Prometheus Operator project.
Signed-off-by: Oded Viner <oviner@redhat.com>
userSecretRef and passwordSecretRef fields are added to allow the Kafka
endpoint username and password to be supplied from a Kubernetes Secret,
rather than exposed as plaintext as part of the endpoint URI. If the
endpoint URI has HTTP basic auth user-id and user-pass components, they
are overridden by userSecretRef and passwordSecretRef.
Squid added bucket topic attributes for configuring the user-name and
password for pushing notifications to Kafka as an alternative to
encoding credentials into the URI. However, these attributes are not
supported under Reef. Thus, this initial implementation relies on URI
mangling for compatible with both Reef and Squid. Future work could
using Ceph version detection and switch to the new attributes to be used
with Squid and/or the implementation could be converted exclusively to
use the new attributes once Rook has dropped support for Reef.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The tests sometimes see that the mons run out of space
in the CI, therefore, we reduce the space required
from 30% down to 10%.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The upgrade tests in master have been upgrading from 1.15 to
master. In anticipation of the v1.17 release, we change
the upgrade tests to start from v1.16 to test if there
are any regressions in the supported upgrades.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
One ore more OSDs can be down and PGs can still be active+clean (data
was rebalanced to other available OSDs). This PR sets
maxunavailable=1+downOSDs to allow one healthy to be drained.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Create the CSI operator in the go integration test suites
to test the new CSI operator scenarios. For the upgrade
test suite, enable the CSI operator after the upgrade
to verify the working cluster after it is enabled.
Signed-off-by: subhamkrai <srai@redhat.com>
Implement an allow list mechanism that disables potentially unsafe OBC
fields by default. OBC fields beyond `maxObjects` and `maxSize` don't
neatly fit into the OBC framework as it was originally envisioned and
implemented.
Some of the newly added configs could allow users to cause confusion for
themselves. Others might allow users to hijack others buckets. Some
might allow bricking the entire S3 store.
Out of an abundance of safety, allow-list the known-safe options by
default, and require administrators to enable potentially troublesome
options via the new operator-level config
`ROOK_OBC_ALLOW_ADDITIONAL_CONFIG_FIELDS`.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.
Additionally, it silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
and string format errors found by golangci-lint, while at it.
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
Ceph has changed the image build process to only require a
dockerfile and stop using the ceph-container repo. The
daily images are pushed to the quay.io/ceph-ci/ceph repo,
so the Rook CI will now start using those images.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The upgrade tests in master have been upgrading from 1.14 to
master. In anticipation of the v1.16 release, we change
the upgrade tests to start from v1.15 to test if there
are any regressions in the supported upgrades.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The devel images have been fairly stable for Rook to test
against, but on occasion there are regressions from
Ceph development that affect the Rook CI. For stability during
Rook development, use the latest stable version of ceph for
PRs, master, and release tests. The daily CI will still use
the devel images from Ceph.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The mon canaries may be created even when the mon daemons
are not created thereafter during the integration tests.
Therefore, the integration tests need to also query a label
specific to the mon daemon so the canaries are not a distraction
to the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The v1beta1 cron jobs have been obsolete since K8s 1.21,
and Rook has not supported that version of K8s
for many moons, so we can remove the obsolete code
for the handling of v1beta1 cron jobs for crash pruning.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Given that Ceph Quincy (v17) is past end of life,
remove Quincy from the supported Ceph versions,
examples, and documentation.
Supported versions now include only Reef and Squid.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ROOK_ENFORCE_HOST_NETWORK option was implemented recently
and now we add the helm setting to expose this new setting
in the rook chart.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The docker.io image prefix is expected to be prepended
to the image names in the test images. This was missed
in 14550 related to some CI tests, which was now causing
the CI failures in the 1.15 branch where the search and
replace was missing the new docker.io prefix.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 3045076db8)
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>
* extend the API object specs for swift and keystone integration
* adapt rgw to the new go-ceph version
- The parameter lists of the API call have changes, as parameters
ignored by the RGW Admin Ops API are no longer serialized, therefore
the mock has to be adapted.
- There is now validation for the user keys that are passed to the
User get API, therefore things failed when we had empty keys in our
User proxy object.
* expand the reconcile loop for the swift and keystone integration
* fix minor mistakes in design document
* add env var to pass extra args to minikube
Minikube decides CPU cores and memory automatically based on the
available resources on the machine which may be insufficient to
run rook. This commit adds an environment variable to add arbitrary
arguments to the minikube command, so both can be specified if
desired.
* integration tests for swift and keystone
The new integration of swift or s3 and keystone support by rook
does not have any integration tests yet.
This commit introduces integration tests for swift and keystone. The
tests are done against a minimal keystone setup (keystone container
image from Yaook-project (https://yaook.cloud), sqlite as database
backend, cert-manager and trust-manager for test certificate setup).
To prevent hardcoded credentials, passwords are generated
by the tests. The integration tests use the openstack client
(keystone- and swift-functionality) (https://docs.openstack.org/
python-openstackclient/ latest/). This was a concious design decision
to use client tooling as close as possible to the end user instead of
using other go-libraries (such as gophercloud).
* add documentation on swift and keystone
Currently there is no documentation on the use of Swift to access
an object store as well as the use of OpenStack keystone for
authentication.
This commit adds documentation on the use of Swift and OpenStack
keystone, as well as CRD-related documentation and an example setup.
* add integration tests for S3 via keystone
This commit introduces integration tests for s3 and keystone. The
tests are run against the same minimal keystone setup that the tests
for swift and keystone use.
The integration tests use the aws s3 client to use client tooling as
close as possible to the end user instead of using other go-libraries.
Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
The helm upgrade tests have been failing frequently, but not
always, on the oldest version of K8s that is tested in the CI
for the past few months. Add a retry to attempt to get
the CI passing more consistently.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The upgrade tests in master have been upgrading from 1.13 to
master. In anticipation of the v1.15 release, we change
the upgrade tests to start from v1.14 to test if there
are any regressions in the supported upgrades.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The daily ceph upgrade tests were running on Rook v1.13.
This was causing the squid upgrade tests to fail since
1.13 does not support Squid. The purpose of the upgrade
tests is to test the ceph upgrades, therefore the daily
tests will just test them based on rook master.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the release of the first squid RC, we add squid
to the supported versions and add tests to run
Rook against the squid release.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the clusters reach full, nearfull, or backfill full thresholds
ceph will raise health warnings and stop allowing IO or backfill
depending on the threshold. These settings require special ceph
commands instead of being generic ceph config. Allow these settings
to be set from the CephCluster CR in the spec.storage
section.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
cephfs subvolumegroup supports creating svg
with quota and the datapool, This PR adds the
support for the same.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The upgrade test should always upgrade from the
previous minor release to the latest master. With 1.14
releasing soon, now we upgrade from 1.13 to master,
to confirm if there are any upgrade issues to 1.14.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The pg_autoscaler is always enabled by ceph and cannot be disabled.
In a previous PR the pg_autoscaler config was removed from the main example.
Now the remaining config for the pg_autoscaler is removed as well.
Signed-off-by: travisn <tnielsen@redhat.com>
For now we are using the operator namespace name
as the prefix for the csi driver, This PR provides
an option for the users if someone wants to have
their own prefix for the csi driver, if someone tries
to change the prefix for existing csi driver rook
operator will fail to reconcile the csi driver.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Both the helm tests are failing because,
```
2023-12-19 08:47:06.136640 E | ceph-file-controller: failed to reconcile CephFilesystem "helm-ns/ceph-filesystem-test". CephFilesystem "helm-ns/ceph-filesystem-test" will not be deleted until all dependents are removed: CephFilesystemSubVolumeGroups: [ceph-filesystem-test-csi]
```
So, let's remove the ceph SVGs before removing Ceph filesystem.
Signed-off-by: subhamkrai <srai@redhat.com>