The generation of a long node name in the integration tests was
being done based on the k8s version. In the past, older K8s versions
did not support the changing name. Now it's more maintainable if
we generate the long name depending on the test suite.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The upgrade integration test was from rook v1.6 to the latest master.
This was necessary until we are ready for the v1.8 release, from which
time we want to focus the upgrade testing from v1.7 to the latest
master.
The duplication in the test CRs and other resources is now reduced
by the upgrade calling a thin wrapper to forward a call to the
master version of the resource. When a new feature is added that
needs to be differentiated from the previous version, the method
then can be implemented instead of wrapping the master implementation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ingress api version changed when it went to v1, and this has caused some upheaval
throughout the kubernetes ecosystem. This commit uses a common method of deciding which
ingress api to use, and allows the optional override of the kubernetes version
presented to helm using the helm build-in capabilities.
also add an ingress into the helm integration tests so any regressions to how ingresses
are handled in the future are caught easier.
Closes rook#9174
Signed-off-by: Tom Hellier <me@tomhellier.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Using a default value for CompressionMode to none effectively overrides
any values for Parameters. It is deprecated but still takes precedence.
Which means that in its previous form, Parameters was always ignored
since CompressionMode was always set to none when empty.
Signed-off-by: Sébastien Han <seb@redhat.com>
The volume replication CRDs are an external component, not owned by Rook.
Therefore, they should be installed as any other independent component
in case the admin will install other consumers of the volumereplication CRDs
in the future in addition to Rook and the CSI driver.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We have new jobs now:
* one that runs both smoke and object on the next Pacific version
* one that runs both smoke and object on Ceph master
* one that tests the upgrade from the current pacific stable to the
pacific devel
* one that tests the upgrade from the current octopus stable to the
octopus devel
Signed-off-by: Sébastien Han <seb@redhat.com>
In Rook v1.8 the min version of K8s supported is updated to 1.16.
Users running on older versions of K8s are recommended to update
to 1.16 or newer before updating to Rook v1.8.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests must always be run against the local
build of rook, and an image should never be pulled from dockerhub.
To prevent pulling a release or master tag, the local build
will use a tag specific to the build and not ever published
elsewhere.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit a8a40428b0)
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Let's force v16.2.5 since the CI is broken with 16.2.6. This gives us
time to continue to merge work and work on fixing deployments with
16.2.6 in parallel.
Signed-off-by: Sébastien Han <seb@redhat.com>
The release version needs to be the same in the example/test manifests
as it is in the local build image. The github actions are different in
this regard than the Jenkins builds were. The Jenkins builds always
locally used the master tag instead of a release-specific tag.
Now that the github actions use the release-specific tag, the test
framework no longer should be using the master tag in release branches.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Our admission webhooks will now run as part of the Operator container
and not an additional deployment. This has the advantage of consuming
fewer resources in the cluster and not having to manage affinities and
tolerations. This only drawback is that the Secret containing the
certificates is not mounted anymore and the content needs to be written
inside the Operator. This is not practical since we also need to watch
for the Secret content to change. Meaning that the certificates have
been renewed and the webhook server needs to use them.
A new approach is on its way to hopefully simplify this last issue and
implement a watcher for the Secret.
In the meantime, users need to use the cert-manager or renew
certificates manually. Additionally, they must update the
ValidatingWebhookConfiguration object with the new CA bundle.
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit adds an ingress resource to the rook-ceph-cluster helm chart, allowing
ingress to the ceph-dashboard service. It also adds the ability to define the
various storage types that you can run on ceph inside kubernetes.
Closes https://github.com/rook/rook/issues/8384
Signed-off-by: Tom Hellier <me@tomhellier.com>
Enable again the mgr test:
- Now is more reliable and robust the start of the test.
- Minor fixes to adapt the <service ls> to the new name of the crash daemon
deployed by rook
- The creation of OSDs is disabled in the orchestrator, so i have removed
this test until we will have the functionality ready again in the orchestrator
part
My plan is to provide in the orchestrator two different ways to create OSDs:
- Creation of OSD using specific devices if discovery daemon is running
- Creation of OSds using PVS (if we have LSO/other LS operator running)
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
Recently, the builds of `ceph/ceph` image moved to quay.io, see
https://github.com/ceph/ceph-build/pull/1883 for more details.
Current images will remain but new builds will happen on quay.io only.
This means that tags such as `v14.2`, `v15.2`,`v16.2` will need to
switch to quay.io to get updates.
Signed-off-by: Sébastien Han <seb@redhat.com>
With v1.7 approaching, the upgrade integration test will now test from
v1.6.x to the latest master, which will effectively become the v1.7
release soon.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRDs v1 requires the full schema for all settings, so we now
generate the CRDs for nfs for full fidelity of all settings.
Co-authored-by: Nicolaj Græsholt <figaw@hotmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRDs v1 requires the full schema for all settings, so we now
generate the CRDs for cassandra for full fidelity of all the
settings.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph provider had been refactored to read the crds
and operator manifests from a file instead of copying the
manifests into the test code. Now the helpers are refactored
to allow the other providers to also reduce the manifest
duplication in tests.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The upgrade test needs to also update the toolbox to ensure that it will not
be denied access to the cluster when the insecure connections are disabled
by the operator.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers
Signed-off-by: Henry Zhang <me@henry.dev>
`CephObjectStore->gateway->type` is not used. Rook has only supported s3-like
interface and hasn't had no code which handles `type` field.
In addition, this field was removed from CRD in the following commit.
ceph: auto-gen crds
31db03fece
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The multicluster test was never removing the finalizer of the external
cluster since the core cluster was being removed first. Now we remove
the external cluster first to ensure the finalizer will be removed
properly instead of forcefully by the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Because of the recent CRD changes made, the mon count had a minimum of
1, making the configuration of the external cluster impossible. The
operator would fail to add the finalizer:
```
2021-04-29 16:57:38.429451 E | ceph-cluster-controller: failed to reconcile. failed to add finalizer: failed to add finalizer "cephcluster.ceph.rook.io" on "test-external": CephCluster.ceph.rook.io "test-external" is invalid: spec.mon.count: Invalid value: 0: spec.mon.count in body should be greater than or equal to 1
```
We now fixed the CRD as well as adding the CR status.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have seen new cases where retrying to lock the device 3 times is not
enough. It's the same race we had experienced where Ceph tries to
acquire a lock on the device but systemd-udevd does the same too.
Retrying 20 times every 0.1sec seems to mitigate that issue in the CI.
Signed-off-by: Sébastien Han <seb@redhat.com>
Since Rook 1.6 is using raw mode for simple OSD scenarios we can use v14
again without having issue with ceph-volume.
We keep the upgraade test with v14.2.12 since Rook 1.5 does not have the
raw code to deploy OSDs.
Closes: https://github.com/rook/rook/issues/7669
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests pick up the manifests directly from the examples
folder, including with the release tag. In the release branch the Jenkins
build still uses the master build when building locally before tagging
it with the release tag. So the tests need to use the master tag.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 456e2f3a79)
With the Ceph Pacific release coming this week we add support
in Rook for Pacific with the Rook v1.6 release coming soon.
The integration tests will now run across nautilus, octopus,
and pacific to cover all supported Ceph versions. The default
examples still specify Octopus until there is more bake time
for Pacific.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If a CI test fails we don't want to cleanup anything and leave the
cluster in the state it is. This will help debugging the CI.
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests run independently in the github actions so there is
no need to cleanup from every test. The cleanup is still needed in the
Jenkins tests where all the suites run serially.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>