Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers
Signed-off-by: Henry Zhang <me@henry.dev>
`CephObjectStore->gateway->type` is not used. Rook has only supported s3-like
interface and hasn't had no code which handles `type` field.
In addition, this field was removed from CRD in the following commit.
ceph: auto-gen crds
31db03fece
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The multicluster test was never removing the finalizer of the external
cluster since the core cluster was being removed first. Now we remove
the external cluster first to ensure the finalizer will be removed
properly instead of forcefully by the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.
Signed-off-by: Rakshith R <rar@redhat.com>
Because of the recent CRD changes made, the mon count had a minimum of
1, making the configuration of the external cluster impossible. The
operator would fail to add the finalizer:
```
2021-04-29 16:57:38.429451 E | ceph-cluster-controller: failed to reconcile. failed to add finalizer: failed to add finalizer "cephcluster.ceph.rook.io" on "test-external": CephCluster.ceph.rook.io "test-external" is invalid: spec.mon.count: Invalid value: 0: spec.mon.count in body should be greater than or equal to 1
```
We now fixed the CRD as well as adding the CR status.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have seen new cases where retrying to lock the device 3 times is not
enough. It's the same race we had experienced where Ceph tries to
acquire a lock on the device but systemd-udevd does the same too.
Retrying 20 times every 0.1sec seems to mitigate that issue in the CI.
Signed-off-by: Sébastien Han <seb@redhat.com>
Since Rook 1.6 is using raw mode for simple OSD scenarios we can use v14
again without having issue with ceph-volume.
We keep the upgraade test with v14.2.12 since Rook 1.5 does not have the
raw code to deploy OSDs.
Closes: https://github.com/rook/rook/issues/7669
Signed-off-by: Sébastien Han <seb@redhat.com>
Rook's testing utilities include a function for loading the default
kubeconfig for a cluster. This function requires maintainance effort and
also doesn't support authentication methods like Basic Authentication.
Instead of adding support for it, drop the config loader and use the one
provided by the controller-runtime library.
Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
The integration tests pick up the manifests directly from the examples
folder, including with the release tag. In the release branch the Jenkins
build still uses the master build when building locally before tagging
it with the release tag. So the tests need to use the master tag.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 456e2f3a79)
With the Ceph Pacific release coming this week we add support
in Rook for Pacific with the Rook v1.6 release coming soon.
The integration tests will now run across nautilus, octopus,
and pacific to cover all supported Ceph versions. The default
examples still specify Octopus until there is more bake time
for Pacific.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If a CI test fails we don't want to cleanup anything and leave the
cluster in the state it is. This will help debugging the CI.
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests run independently in the github actions so there is
no need to cleanup from every test. The cleanup is still needed in the
Jenkins tests where all the suites run serially.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The rbac required for multus was present in role which makes
only the network-attachment-definitions in the rook cluster namespace
accessible to the operator. Moved it to clusterrole for cluster wide
access to NAD.
Signed-off-by: rohan47 <rohgupta@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Finally! We can now stop editing manually our crds definition. Simply
run `make crds-gen`. These two files:
* `cluster/examples/kubernetes/ceph/crds.yaml`
* `cluster/charts/rook-ceph/templates/resources.yaml`
will be autogenerated for us.
We rely on the controller-gen tool, it reads our API definitions from
`pkg/apis/` and produces the CRD files accordingly.
It uses "markers" to add extra yaml fields to the CRD, for instance we
have a lot fields with:
```
nullable: true
x-kubernetes-preserve-unknown-fields: true
```
The corresponding API type markers are:
```
// +kubebuilder:pruning:PreserveUnknownFields
// +nullable
```
For more API convention see
https://github.com/kubernetes/community/blob/master/contributors/devel/sig-architecture/api-conventions.md#optional-vs-required
and for the markers see: https://book.kubebuilder.io/reference/markers/crd-validation.html
Signed-off-by: Sébastien Han <seb@redhat.com>
This is a follow up to:
https://github.com/rook/rook/pull/7264
which added pool quotas. It was stated that it would be nice to
be able to set pool max_bytes quotas using standard k8s quantity
formats rather than an integer representing the number of bytes.
I was too busy at the time to make the changes. This PR adds the
functionality.
Signed-off-by: Frank Ritchie <12985912+fritchie@users.noreply.github.com>
If there are multiple mgrs running, the operator will automatically
add pod antiaffinity across hosts by default. For stretch clusters,
antiaffinity across the stretch failure domain (e.g. zones)
will be added. Required antiaffinity will be created unless the
test setting allowMultiplePerNode is set to true.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The mgr daemon may be failed over by ceph if the active mgr is not
responding and the standby mgr is available. If the active mgr changes
the services for the dashboard and metrics will be updated with a
label selector for the new active mgr. The services cannot direct
traffic to the standby mgr or else they will be incorrectly redirected
to the active mgr.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The GRPC metrics exposed by both the provisioner pod
and the node plugin pod on some port. Provisioner pod
is running on the pod network and the daemonset pods
run on the host network. sometimes starting the GRPC
metrics by default can lead to node plugin pods
crashloopback state this is due to the port conflict.
Moreover, the GRPC metrics are not for the user it's
for the one which will help to debug the time taken
by cephcsi to serve each GRPC call. Enabling it by
default won't be a good idea. So the plan is to
disable it by default, If someone faces any issue it
can be enabled later at some point in time.
closes#7378
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The integration tests will now read the manifests from the same source
that is used in the examples for installation. No longer is there a need
for a copy of the CRDs in the integration test code.
The tests will read the manifests directly from disk under a path
relative to the rook root. The upgrade test will read the manifests from
a github url where the base version has been published.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Since the 1.5 releaase we need the integration test to upgrade from
1.5 to master instead of from 1.4 to master.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This reverts commit 114a949cfb since
14.2.14 is out and has a fix to use raw mode instead of LVM mode on
partitions.
Signed-off-by: Sébastien Han <seb@redhat.com>
Do not use old Ceph v15.2.7 in integration that was working around the
issue of ceph-volume partition support being removed in v15.2.8.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
(cherry picked from commit 99cee7648cfaf72b02d62e281a2ac3a18bce4132)
Even though CRs for bucket health check was added, CRDs were missing from
crds.yaml file. So the health check always enabled for RGW.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Remove features and design that supports adding OSDs to Ceph clusters
via `spec:driveGroups`. Update the Ceph upgrade doc that informs users
who currently use Drive Groups (we believe there are none of these
users) how to migrate to using the `spec:storage` config.
Resolves https://github.com/rook/rook/issues/7275
Revert "ceph: fix drive group deployment failure"
This reverts commit 76f1d9944e.
Revert "ceph: osd: add drive groups spec to cluster CR"
This reverts commit 7117fc12b7.
Revert "design: ceph orchestrator module add/remove OSDs"
This reverts commit 178187d035.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Add the cephfilesystemmirrors definition for CI runs where the
Kubernetes version is 1.15 and extension v1 extensions are not
supported.
Closes: https://github.com/rook/rook/issues/7168
Signed-off-by: Sébastien Han <seb@redhat.com>
The YugabyteDB team has decided to move support of their operator
to a new repo at https://github.com/yugabyte/yugabyte-operator.
The operator in Rook is no longer needed. Further usage of
the YugabyteDB operator is recommended at the new location.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
integration test was failing frequently due to some
intermittent failures.
```
Error from server (InternalError): error when creating
"STDIN": Internal error occurred: failed calling webhook.
```
(This error is not coming from us most probably)
This is kind of workaround so we get green build
as we still don't know the root cause.
Signed-off-by: subhamkrai <srai@redhat.com>
The deviceClass was missing from the object store CRD schema,
disallowing the device class from being specified on
the object store pools.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
disable admission controller for helm suite.
v1 version of admissionregistration.k8s.io/v1
require minimum v1.16.0 of k8s. All other suite
was disabled earlier this was left.
Signed-off-by: subhamkrai <srai@redhat.com>
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.
The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050
The same goes for integration tests, they will get added later once we
start testing on Pacific.
Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
The PDBs have been stabilizing for a good length of time now since
they were created in v1.1, and later redesigned in v1.5. The time
has come to enable the feature by default so everyone can enjoy
the stable Rook storage even while draining K8s nodes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
v1 version of admission controller require minimum v1.16.0
of k8s. So, this commits disable admission controller when
k8s version older than v1.16.0.
Signed-off-by: subhamkrai <srai@redhat.com>
disabling admission controller for the upgrade test.
Upgrade test was using older controller runtime version
and api v1 need latest version controller runtime(v0.7).
Signed-off-by: subhamkrai <srai@redhat.com>
The cockroachDB operator has not had community support in Rook.
Therefore, the time has come to deprecate and remove it.
If the sources are still needed, there is always git history.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>