Commit Graph
617 Commits
Author SHA1 Message Date
Travis Nielsen 9d2aa1f6bd test: generate long node name depending on test suite
The generation of a long node name in the integration tests was
being done based on the k8s version. In the past, older K8s versions
did not support the changing name. Now it's more maintainable if
we generate the long name depending on the test suite.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-06 08:03:56 -07:00
Travis Nielsen d9ac8ce490 test: upgrade integration test from 1.7 to master
The upgrade integration test was from rook v1.6 to the latest master.
This was necessary until we are ready for the v1.8 release, from which
time we want to focus the upgrade testing from v1.7 to the latest
master.

The duplication in the test CRs and other resources is now reduced
by the upgrade calling a thin wrapper to forward a call to the
master version of the resource. When a new feature is added that
needs to be differentiated from the previous version, the method
then can be implemented instead of wrapping the master implementation.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-30 08:07:42 -07:00
Sébastien Han c890710b63 core: change directory layout
As per discussion, proposing a new layout for the charts/yaml/olm files.

./deploy
├── charts
│   ├── rook-ceph
│   │   └── templates
│   └── rook-ceph-cluster
│       └── templates
├── examples
│   ├── csi
│   │   ├── cephfs
│   │   └── rbd
│   ├── flex
│   ├── monitoring
│   ├── pre-k8s-1.16
└── olm
    └── assemble

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-30 09:12:53 +01:00
Travis Nielsen ba54567e8f Merge pull request #9176 from TomHellier/9174-ingress-support-more-k8s-versions
helm: Allow further configurability of the ingress version
2021-11-23 13:47:33 -07:00
Tom Hellier ba44602477 helm: allow further configurability of ingress version
The ingress api version changed when it went to v1, and this has caused some upheaval
throughout the kubernetes ecosystem. This commit uses a common method of deciding which
ingress api to use, and allows the optional override of the kubernetes version
presented to helm using the helm build-in capabilities.
also add an ingress into the helm integration tests so any regressions to how ingresses
are handled in the future are caught easier.

Closes rook#9174

Signed-off-by: Tom Hellier <me@tomhellier.com>
2021-11-22 10:03:02 +00:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Yuval Lifshitz 71ed45b69b rgw: implement bucket notifications for object storage
following the design from here:
https://github.com/rook/rook/blob/master/design/ceph/object/ceph-bucket-notification-crd.md

Closes: https://github.com/rook/rook/issues/5313
Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-11-04 11:20:40 +02:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Sébastien Han 28cc6f5514 ceph: remove default value for pool compression
Using a default value for CompressionMode to none effectively overrides
any values for Parameters. It is deprecated but still takes precedence.
Which means that in its previous form, Parameters was always ignored
since CompressionMode was always set to none when empty.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-13 17:06:22 +02:00
Travis Nielsen c420f2309c csi: no longer install the volumereplication crds from rook
The volume replication CRDs are an external component, not owned by Rook.
Therefore, they should be installed as any other independent component
in case the admin will install other consumers of the volumereplication CRDs
in the future in addition to Rook and the CSI driver.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-06 16:27:44 +00:00
Sébastien Han b4a36e9967 ci: wait longer for pod label to be deleted
I've seen cases were the CI needs a few more seconds to delete and
object store. When logging in the runner, the object store is gone and
the timing matches too with the runner's logs (comparing with the
operator's logs).

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-28 15:24:39 +02:00
Sébastien Han 33dbaba38b ci: add daily jobs
We have new jobs now:

* one that runs both smoke and object on the next Pacific version
* one that runs both smoke and object on Ceph master
* one that tests the upgrade from the current pacific stable to the
  pacific devel
* one that tests the upgrade from the current octopus stable to the
  octopus devel

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 11:59:04 +02:00
Travis Nielsen f5c1543ddf build: update min k8s version to 1.16
In Rook v1.8 the min version of K8s supported is updated to 1.16.
Users running on older versions of K8s are recommended to update
to 1.16 or newer before updating to Rook v1.8.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:56 -06:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Travis Nielsen 9ff0753514 test: run all integration tests against the local build
The integration tests must always be run against the local
build of rook, and an image should never be pulled from dockerhub.
To prevent pulling a release or master tag, the local build
will use a tag specific to the build and not ever published
elsewhere.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit a8a40428b0)
2021-09-22 07:39:42 -06:00
Sébastien Han 0c33493f27 ceph: bump manifests to ceph pacific 16.2.6
New version is out so let's use it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 09:16:32 +02:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han ae291afb2f ci: force a particular ceph version
Let's force v16.2.5 since the CI is broken with 16.2.6. This gives us
time to continue to merge work and work on fixing deployments with
16.2.6 in parallel.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 15:28:48 +02:00
Jiffin Tony Thottan ca43800119 ceph: add options for cephobjectstore user
Adding options for quota, bucket limit, caps for the
`cephobjectstoreuser`.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-07 22:43:09 +05:30
Blaine Gardner a1814af1d9 ceph: remove NFS and Cassandra operator code
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-08-31 14:07:02 -06:00
Travis Nielsen 23b772643b ceph: test against release version in the branch
The release version needs to be the same in the example/test manifests
as it is in the local build image. The github actions are different in
this regard than the Jenkins builds were. The Jenkins builds always
locally used the master tag instead of a release-specific tag.
Now that the github actions use the release-specific tag, the test
framework no longer should be using the master tag in release branches.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-26 12:10:05 -06:00
Sébastien Han 41e915d411 Merge pull request #8493 from leseb/admission-controller
ceph: move the admission webhook to the operator
2021-08-25 09:30:54 +02:00
Sébastien Han 656dd0f334 ceph: move the admission webhook to the operator
Our admission webhooks will now run as part of the Operator container
and not an additional deployment. This has the advantage of consuming
fewer resources in the cluster and not having to manage affinities and
tolerations. This only drawback is that the Secret containing the
certificates is not mounted anymore and the content needs to be written
inside the Operator. This is not practical since we also need to watch
for the Secret content to change. Meaning that the certificates have
been renewed and the webhook server needs to use them.
A new approach is on its way to hopefully simplify this last issue and
implement a watcher for the Secret.
In the meantime, users need to use the cert-manager or renew
certificates manually. Additionally, they must update the
ValidatingWebhookConfiguration object with the new CA bundle.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-24 19:07:04 +02:00
Jiffin Tony Thottan f4bb47e440 ceph: add support for update() from lib-bucket-provisioner
Recently lib-bucket-provisioner add support for update() API.
Include that on the obc implementation since it can be used to
update quota for OBC.

Fixes: #7146

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-08-19 23:57:11 +05:30
parth-gr b28455245d ci: fix for CephObjectStores flakiness
Integration test CephSmokeSuite fails frequently
A quick fix for it by reordering storeName,
running tlsteststore before teststore

Closes: https://github.com/rook/rook/issues/8309
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-09 14:08:32 +05:30
Travis Nielsen c2a551123f Merge pull request #8401 from TomHellier/add-additional-rook-ceph-cluster-helm-features
ceph: Add additional helm chart functionality for ingresses and defining the ceph storage crds
2021-07-28 16:14:48 -06:00
Tom Hellier 89ab4f90ec ceph: adds helm functionality for ingress, and ceph storage crds
This commit adds an ingress resource to the rook-ceph-cluster helm chart, allowing
ingress to the ceph-dashboard service. It also adds the ability to define the
various storage types that you can run on ceph inside kubernetes.

Closes https://github.com/rook/rook/issues/8384

Signed-off-by: Tom Hellier <me@tomhellier.com>
2021-07-28 21:51:56 +01:00
Juan Miguel Olmo Martínez 9edff582c3 ceph: enable again the Rook orchestrator mgr test
Enable again the mgr test:
- Now is more reliable and robust the start of the test.
- Minor fixes to adapt the <service ls>  to the new name of the crash daemon
deployed by rook

- The creation of OSDs is disabled in the orchestrator, so i have removed
this test until we will have the functionality ready again in the orchestrator
part

My plan is to provide in the orchestrator two different ways to create OSDs:
- Creation of OSD using specific devices if discovery daemon is running
- Creation of OSds using PVS (if we have LSO/other LS operator running)

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2021-07-28 16:27:42 +02:00
Sébastien Han c8ee5674dd Merge pull request #8103 from leseb/rm-jenkins
ci: remove jenkins from master branch
2021-07-23 18:20:10 +02:00
Sébastien Han 0811359d28 Merge pull request #8358 from leseb/move-to-quay
ceph: move all of our docker.io reference to quay.io
2021-07-23 18:03:53 +02:00
Travis Nielsen 07ddcda0ce ceph: update test to watch for v1 cronjob
The v1beta1 cronjob is deprecated and needs to use v1 for K8s 1.16 and newer

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-22 11:25:55 -06:00
Sébastien Han 6bce1ff3e9 ceph: move all of our docker.io reference to quay.io
Recently, the builds of `ceph/ceph` image moved to quay.io, see
https://github.com/ceph/ceph-build/pull/1883 for more details.
Current images will remain but new builds will happen on quay.io only.

This means that tags such as `v14.2`, `v15.2`,`v16.2` will need to
switch to quay.io to get updates.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-22 11:17:05 +02:00
Travis Nielsen d1f02f22f3 ceph: test upgrades from v1.6 to master
With v1.7 approaching, the upgrade integration test will now test from
v1.6.x to the latest master, which will effectively become the v1.7
release soon.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-20 11:03:40 -06:00
Travis NielsenandNicolaj Græsholt f114db952b nfs: generate the crds from types for v1
The CRDs v1 requires the full schema for all settings, so we now
generate the CRDs for nfs for full fidelity of all settings.

Co-authored-by: Nicolaj Græsholt <figaw@hotmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-20 08:18:05 -06:00
Travis Nielsen 093cf6dbbd cassandra: generate the crds from types for v1
The CRDs v1 requires the full schema for all settings, so we now
generate the CRDs for cassandra for full fidelity of all the
settings.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-20 08:18:05 -06:00
Travis Nielsen 6ad9190f55 ceph: factor out test manifest helpers for providers
The ceph provider had been refactored to read the crds
and operator manifests from a file instead of copying the
manifests into the test code. Now the helpers are refactored
to allow the other providers to also reduce the manifest
duplication in tests.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-20 08:17:29 -06:00
Sébastien Han eaa6e7732c Merge pull request #8272 from leseb/exec-in-pod
ceph: proxy ceph command when multus is configured
2021-07-07 21:36:54 +02:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Jiffin Tony Thottan 9e3cf68d04 test: ci test for TLS objectstore
Extend the object store smoke test to include TLS configurations.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-07-01 18:50:27 +05:30
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Andy Bursavich 11f0add706 build: update golangci-lint version and nolint comment format
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:50:04 -07:00
Andy Bursavich 7ab1baa794 test: extract utils dependency on k8s.io/kubernetes
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:36:08 -07:00
subhamkrai 0ff452fdd1 ceph: remove unnecessary file
this commit removes `mysql_helper.go`,
unnecessary file.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-06-15 15:28:13 +05:30
Travis Nielsen 44d31f148e ceph: update the toolbox in the upgrade test
The upgrade test needs to also update the toolbox to ensure that it will not
be denied access to the cluster when the insecure connections are disabled
by the operator.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-06-10 10:05:30 -06:00
Sébastien Han 18b1477352 ci: remove jenkins from master branch
Thank you Jenkins for all those years, but it's to retire.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-10 15:34:15 +02:00
Henry Zhang ecf6bdcfae ceph: add build and tests for rook-ceph-cluster Helm chart
Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers

Signed-off-by: Henry Zhang <me@henry.dev>
2021-05-27 01:08:27 -07:00
Travis Nielsen e46dc1cd17 Merge pull request #7928 from leseb/fix-7927
ci: retry when fetching online manifests
2021-05-18 10:52:47 -06:00
Sébastien Han f1f3416f49 ci: retry when fetching online manifests
We now retry up to 3 times if the manifest fails to be fetched online.

Closes: https://github.com/rook/rook/issues/7927
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-18 17:22:12 +02:00
Satoru Takeuchi c2adbdd8a8 ceph: remove an missing field in crd
`CephObjectStore->gateway->type` is not used. Rook has only supported s3-like
interface and hasn't had no code which handles `type` field.

In addition, this field was removed from CRD in the following commit.

ceph: auto-gen crds
31db03fece

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-05-17 07:51:13 +00:00
Travis Nielsen 2b2a1f66c4 ceph: multicluster test cleanup of external cluster first
The multicluster test was never removing the finalizer of the external
cluster since the core cluster was being removed first. Now we remove
the external cluster first to ensure the finalizer will be removed
properly instead of forcefully by the test.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-06 14:51:07 -06:00