Commit Graph
572 Commits
Author SHA1 Message Date
Henry Zhang ecf6bdcfae ceph: add build and tests for rook-ceph-cluster Helm chart
Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers

Signed-off-by: Henry Zhang <me@henry.dev>
2021-05-27 01:08:27 -07:00
Travis Nielsen e46dc1cd17 Merge pull request #7928 from leseb/fix-7927
ci: retry when fetching online manifests
2021-05-18 10:52:47 -06:00
Sébastien Han f1f3416f49 ci: retry when fetching online manifests
We now retry up to 3 times if the manifest fails to be fetched online.

Closes: https://github.com/rook/rook/issues/7927
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-18 17:22:12 +02:00
Satoru Takeuchi c2adbdd8a8 ceph: remove an missing field in crd
`CephObjectStore->gateway->type` is not used. Rook has only supported s3-like
interface and hasn't had no code which handles `type` field.

In addition, this field was removed from CRD in the following commit.

ceph: auto-gen crds
31db03fece

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-05-17 07:51:13 +00:00
Travis Nielsen 2b2a1f66c4 ceph: multicluster test cleanup of external cluster first
The multicluster test was never removing the finalizer of the external
cluster since the core cluster was being removed first. Now we remove
the external cluster first to ensure the finalizer will be removed
properly instead of forcefully by the test.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-06 14:51:07 -06:00
Rakshith R 7987e1b8a2 ceph: update snapshot APIs from v1beta1 to v1
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-05-05 21:04:10 +05:30
Sébastien Han 15a12577f5 ceph: fix external mode setup
Because of the recent CRD changes made, the mon count had a minimum of
1, making the configuration of the external cluster impossible. The
operator would fail to add the finalizer:

```
2021-04-29 16:57:38.429451 E | ceph-cluster-controller: failed to reconcile. failed to add finalizer: failed to add finalizer "cephcluster.ceph.rook.io" on "test-external": CephCluster.ceph.rook.io "test-external" is invalid: spec.mon.count: Invalid value: 0: spec.mon.count in body should be greater than or equal to 1
```

We now fixed the CRD as well as adding the CR status.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-03 12:07:40 +02:00
Sébastien Han d17fc6ac12 ci: increase the number of lock retries
We have seen new cases where retrying to lock the device 3 times is not
enough. It's the same race we had experienced where Ceph tries to
acquire a lock on the device but systemd-udevd does the same too.
Retrying 20 times every 0.1sec seems to mitigate that issue in the CI.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-04-22 11:30:37 +02:00
Sébastien Han d35ddcbecc ceph: use v14 again
Since Rook 1.6 is using raw mode for simple OSD scenarios we can use v14
again without having issue with ceph-volume.

We keep the upgraade test with v14.2.12 since Rook 1.5 does not have the
raw code to deploy OSDs.

Closes: https://github.com/rook/rook/issues/7669
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-04-20 16:51:46 +02:00
bipuladh 7d5cc8d06b ceph: adds a container to support volume replication controller
this commit adds a new container inside of rbd-provisioner.

Signed-off-by: bipuladh <badhikar@redhat.com>
2021-04-15 21:39:26 +05:30
Yannis Zarkadas c40a116110 cassandra: add missing sidecar permissions to testing manifests
Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
2021-04-14 21:40:50 +03:00
Yannis Zarkadas b71f1077b5 test: replace in-house config loader for controller-runtime's
Rook's testing utilities include a function for loading the default
kubeconfig for a cluster. This function requires maintainance effort and
also doesn't support authentication methods like Basic Authentication.
Instead of adding support for it, drop the config loader and use the one
provided by the controller-runtime library.

Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
2021-04-14 21:40:50 +03:00
Travis Nielsen 9546a2ef94 ceph: integration tests use master tag even in release branch
The integration tests pick up the manifests directly from the examples
folder, including with the release tag. In the release branch the Jenkins
build still uses the master build when building locally before tagging
it with the release tag. So the tests need to use the master tag.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 456e2f3a79)
2021-04-08 23:43:16 -06:00
Travis Nielsen 665856b90a ceph: enable pacific as a supported ceph version
With the Ceph Pacific release coming this week we add support
in Rook for Pacific with the Rook v1.6 release coming soon.
The integration tests will now run across nautilus, octopus,
and pacific to cover all supported Ceph versions. The default
examples still specify Octopus until there is more bake time
for Pacific.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-07 06:59:04 -06:00
Sébastien Han a196e8c0a1 ci: skip any cleanup if a test fails
If a CI test fails we don't want to cleanup anything and leave the
cluster in the state it is. This will help debugging the CI.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-24 19:03:09 +01:00
Travis Nielsen bdbf264e68 ceph: integration tests in github actions skip cleanup
The integration tests run independently in the github actions so there is
no need to cleanup from every test. The cleanup is still needed in the
Jenkins tests where all the suites run serially.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:11 -06:00
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
rohan47 f3a75216a8 ceph: updated rbac for multus in helm charts
The rbac required for multus was present in role which makes
only the network-attachment-definitions in the rook cluster namespace
accessible to the operator. Moved it to clusterrole for cluster wide
access to NAD.

Signed-off-by: rohan47 <rohgupta@redhat.com>
2021-03-16 19:40:17 +05:30
Satoru Takeuchi 756d4ecb89 Merge pull request #7259 from cybozu-go/ceph-improve-owner-reference-management-2
ceph: improve owner reference management
2021-03-16 20:22:57 +09:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Sébastien Han 31db03fece ceph: auto-gen crds
Finally! We can now stop editing manually our crds definition. Simply
run `make crds-gen`. These two files:

* `cluster/examples/kubernetes/ceph/crds.yaml`
* `cluster/charts/rook-ceph/templates/resources.yaml`

will be autogenerated for us.

We rely on the controller-gen tool, it reads our API definitions from
`pkg/apis/` and produces the CRD files accordingly.

It uses "markers" to add extra yaml fields to the CRD, for instance we
have a lot fields with:

```
nullable: true
x-kubernetes-preserve-unknown-fields: true
```

The corresponding API type markers are:

```
// +kubebuilder:pruning:PreserveUnknownFields
// +nullable
```

For more API convention see
https://github.com/kubernetes/community/blob/master/contributors/devel/sig-architecture/api-conventions.md#optional-vs-required
and for the markers see: https://book.kubebuilder.io/reference/markers/crd-validation.html

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-16 09:16:24 +01:00
Travis Nielsen 53cc160817 Merge pull request #7397 from fritchie/pool-quota-units
ceph: set pool quota with k8s quantity format
2021-03-15 18:09:48 -06:00
Frank Ritchie 9ae699d632 ceph: set pool quota with k8s quantity format
This is a follow up to:

https://github.com/rook/rook/pull/7264

which added pool quotas. It was stated that it would be nice to
be able to set pool max_bytes quotas using standard k8s quantity
formats rather than an integer representing the number of bytes.

I was too busy at the time to make the changes. This PR adds the
functionality.

Signed-off-by: Frank Ritchie <12985912+fritchie@users.noreply.github.com>
2021-03-15 19:26:41 -04:00
Satoru Takeuchi 4766954dfc ceph: suppress golanglint-ci complaints
Suppress the following complaints.

```
$ golangci-lint run -E gosec
pkg/operator/test/client.go:244:13: G404: Use of weak random number generator (math/rand instead of crypto/rand) (gosec)
        randIdx := rand.Intn(len(nodes.Items))
                   ^
tests/framework/installer/ceph_manifests_v1.5.go:44:19: G107: Potential HTTP request made with variable url (gosec)
        response, err := http.Get(url)
                         ^
pkg/daemon/ceph/client/pool.go:182:6: ineffectual assignment to stats (ineffassign)
        var stats = new(PoolStatistics)
            ^
pkg/operator/k8sutil/pod_test.go:34:2: ineffectual assignment to container (ineffassign)
        container, err := GetMatchingContainer([]v1.Container{}, expectedName)
        ^
```

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-12 15:34:27 +00:00
Travis Nielsen e7331b0191 ceph: multiple mgrs have antiaffinity across hosts or zones
If there are multiple mgrs running, the operator will automatically
add pod antiaffinity across hosts by default. For stretch clusters,
antiaffinity across the stretch failure domain (e.g. zones)
will be added. Required antiaffinity will be created unless the
test setting allowMultiplePerNode is set to true.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00
Travis Nielsen 388ff3e78b ceph: start sidecar to monitor active mgr
The mgr daemon may be failed over by ceph if the active mgr is not
responding and the standby mgr is available. If the active mgr changes
the services for the dashboard and metrics will be updated with a
label selector for the new active mgr. The services cannot direct
traffic to the standby mgr or else they will be incorrectly redirected
to the active mgr.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00
Madhu Rajanna 0a81ce2251 ceph: disable CSI GRPC metrics by default
The GRPC metrics exposed by both the provisioner pod
and the node plugin pod on some port. Provisioner pod
is running on the pod network and the daemonset pods
run on the host network. sometimes starting the GRPC
metrics by default can lead to node plugin pods
crashloopback state this is due to the port conflict.
Moreover, the GRPC metrics are not for the user it's
for the one which will help to debug the time taken
by cephcsi to serve each GRPC call. Enabling it by
default won't be a good idea. So the plan is to
disable it by default, If someone faces any issue it
can be enabled later at some point in time.

closes #7378

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-03-10 11:30:50 +05:30
Travis Nielsen c402e40f54 ceph: load test crds from files or urls
The integration tests will now read the manifests from the same source
that is used in the examples for installation. No longer is there a need
for a copy of the CRDs in the integration test code.

The tests will read the manifests directly from disk under a path
relative to the rook root. The upgrade test will read the manifests from
a github url where the base version has been published.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-08 21:34:59 -07:00
Travis Nielsen 3fd89dc355 ceph: upgrade test upgrades from 1.5 instead of 1.4
Since the 1.5 releaase we need the integration test to upgrade from
1.5 to master instead of from 1.4 to master.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-08 18:12:42 -07:00
Sébastien Han 2a192300e2 Merge pull request #4879 from leseb/raw-mode-non-pvc
ceph: add raw mode for non-pvc osd
2021-03-03 18:55:21 +01:00
Sébastien Han c1362ce482 ceph: revert "ceph: test latest ceph on raw device"
This reverts commit 114a949cfb since
14.2.14 is out and has a fix to use raw mode instead of LVM mode on
   partitions.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-03 14:06:19 +01:00
Blaine Gardner bc9c5f2057 ceph: start using latest Ceph v15 in integration
Do not use old Ceph v15.2.7 in integration that was working around the
issue of ceph-volume partition support being removed in v15.2.8.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
(cherry picked from commit 99cee7648cfaf72b02d62e281a2ac3a18bce4132)
2021-03-03 14:06:19 +01:00
Santosh Pillai 1159455345 ceph: update CRDs for healthcheck in crds.yaml
Even though CRs for bucket health check was added, CRDs were missing from
crds.yaml file. So the health check always enabled for RGW.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-03-03 10:42:35 +05:30
Blaine Gardner efa7d66d45 ceph: remove driveGroup support
Remove features and design that supports adding OSDs to Ceph clusters
via `spec:driveGroups`. Update the Ceph upgrade doc that informs users
who currently use Drive Groups (we believe there are none of these
users) how to migrate to using the `spec:storage` config.

Resolves https://github.com/rook/rook/issues/7275

Revert "ceph: fix drive group deployment failure"
This reverts commit 76f1d9944e.

Revert "ceph: osd: add drive groups spec to cluster CR"
This reverts commit 7117fc12b7.

Revert "design: ceph orchestrator module add/remove OSDs"
This reverts commit 178187d035.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-02-25 17:08:47 -07:00
Frank Ritchie a0488dac73 ceph: add ability to set pool quotas
This will allow setting ceph pool quotas in bytes and/or objects.

Ceph documentation: https://docs.ceph.com/en/latest/rados/operations/pools/#set-pool-quotas

Signed-off-by: Frank Ritchie <12985912+fritchie@users.noreply.github.com>
2021-02-19 13:49:25 -05:00
Sébastien Han fd0716a026 Merge pull request #7173 from leseb/fix-7168
ceph: add cephfilesystemmirrors CRD to 1.15
2021-02-09 16:31:25 +01:00
Sébastien Han df683a0c57 ceph: add cephfilesystemmirrors CRD to 1.15
Add the cephfilesystemmirrors definition for CI runs where the
Kubernetes version is 1.15 and extension v1 extensions are not
supported.

Closes: https://github.com/rook/rook/issues/7168
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-09 09:50:16 +01:00
Travis Nielsen 043f947321 build: remove yugabytedb operator
The YugabyteDB team has decided to move support of their operator
to a new repo at https://github.com/yugabyte/yugabyte-operator.
The operator in Rook is no longer needed. Further usage of
the YugabyteDB operator is recommended at the new location.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-02-06 00:03:41 -07:00
subhamkrai 9be70baf5f ci: intermittent integration test failures due to admission controller
integration test was failing frequently due to some
intermittent failures.
```
Error from server (InternalError): error when creating
"STDIN": Internal error occurred: failed calling webhook.
```
(This error is not coming from us most probably)
This is kind of workaround so we get green build
as we still don't know the root cause.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-02-04 09:42:29 +05:30
Madhu Rajanna 81248f1dce ceph: add configmap get RBAC for rbd provisioner
rbd provisioner need to access the configmap
in different namespaces.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-02-03 21:28:00 +05:30
Travis Nielsen 81b398f6b6 ceph: add deviceClass to the object store schema
The deviceClass was missing from the object store CRD schema,
disallowing the device class from being specified on
the object store pools.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-02-02 12:13:23 -07:00
Travis Nielsen bc5be15bb3 Merge pull request #7100 from subhamkrai/disable-adm-controller
ceph: disable admission controller for helm suite
2021-01-29 09:22:33 -07:00
subhamkrai 75030316ca ceph: disable admission controller for helm suite
disable admission controller for helm suite.
v1 version of admissionregistration.k8s.io/v1
require minimum v1.16.0 of k8s. All other suite
was disabled earlier this was left.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-01-29 12:53:03 +05:30
Sébastien Han c0123cf182 ceph: add cephfs mirroring support
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.

The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050

The same goes for integration tests, they will get added later once we
start testing on Pacific.

Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-28 19:21:18 +01:00
Travis Nielsen 7472cfd9fa ceph: enable pod disruption budgets by default
The PDBs have been stabilizing for a good length of time now since
they were created in v1.1, and later redesigned in v1.5. The time
has come to enable the feature by default so everyone can enjoy
the stable Rook storage even while draining K8s nodes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-28 10:25:33 -07:00
Travis Nielsen 53ed11f15b build: remove the edgefs operator from rook
The EdgeFS operator has been deprecated for some time in Rook.
If the replacement is added back to Rook it can be completed
according to the new guidelines in the documentation.
https://rook.io/docs/rook/master/storage-providers.html

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-25 17:51:24 -07:00
Sébastien Han 720380e3d3 Merge pull request #7024 from travisn/remove-cockroachdb
build: Remove the CockroachDB operator from Rook
2021-01-25 10:04:44 +01:00
subhamkrai f7d7c2979a ceph: disable admission controller for k8s older than v1.16.0
v1 version of admission controller require minimum v1.16.0
of k8s. So, this commits disable admission controller when
k8s version older than v1.16.0.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-01-22 14:21:03 +05:30
subhamkrai f17f908c31 ceph: disable admission controller for CephUpgradeSuite
disabling admission controller for the upgrade test.
Upgrade test was using older controller runtime version
and api v1 need latest version controller runtime(v0.7).

Signed-off-by: subhamkrai <srai@redhat.com>
2021-01-21 22:26:02 +05:30
Travis Nielsen 2294851ca3 build: remove the cockroachdb operator from rook
The cockroachDB operator has not had community support in Rook.
Therefore, the time has come to deprecate and remove it.
If the sources are still needed, there is always git history.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-20 17:28:58 -07:00