Commit Graph
148 Commits
Author SHA1 Message Date
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Travis Nielsen 07ddcda0ce ceph: update test to watch for v1 cronjob
The v1beta1 cronjob is deprecated and needs to use v1 for K8s 1.16 and newer

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-22 11:25:55 -06:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Andy Bursavich 11f0add706 build: update golangci-lint version and nolint comment format
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:50:04 -07:00
Andy Bursavich 7ab1baa794 test: extract utils dependency on k8s.io/kubernetes
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:36:08 -07:00
subhamkrai 0ff452fdd1 ceph: remove unnecessary file
this commit removes `mysql_helper.go`,
unnecessary file.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-06-15 15:28:13 +05:30
Henry Zhang ecf6bdcfae ceph: add build and tests for rook-ceph-cluster Helm chart
Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers

Signed-off-by: Henry Zhang <me@henry.dev>
2021-05-27 01:08:27 -07:00
Rakshith R 7987e1b8a2 ceph: update snapshot APIs from v1beta1 to v1
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-05-05 21:04:10 +05:30
Yannis Zarkadas b71f1077b5 test: replace in-house config loader for controller-runtime's
Rook's testing utilities include a function for loading the default
kubeconfig for a cluster. This function requires maintainance effort and
also doesn't support authentication methods like Basic Authentication.
Instead of adding support for it, drop the config loader and use the one
provided by the controller-runtime library.

Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
2021-04-14 21:40:50 +03:00
Travis Nielsen bdbf264e68 ceph: integration tests in github actions skip cleanup
The integration tests run independently in the github actions so there is
no need to cleanup from every test. The cleanup is still needed in the
Jenkins tests where all the suites run serially.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:11 -06:00
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
Blaine Gardner 9b0ba6ae8b ceph: add obc to upgrade test
Add object bucket claim to Ceph upgrade test.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-08 09:56:00 -07:00
Blaine Gardner 40fd80cf14 ceph: update smoke test to verify obc is bound
When verifying OBC creation, validating that secret and configmap exist
is good, but the definitive validation is to check that the OBC's phase
is "Bound".

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-07 13:58:22 -07:00
Travis Nielsen c15499ed7e ceph: use the correct snapshot controller version in the tests
The snapshot controller is specifying to use the canary image instead
of the expected version. Update the tests to set the correct version.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-16 10:25:00 -07:00
Renan Campos ad9dcb0c05 ceph: periodically prune crash entries older than user-provided days
Rook's crashcollector pod posts entries to the ceph cluster when a crash occurs.
Over time the number cluster may hold crash entries needlessly.
To clean up old crash entries, this PR adds a field to the ceph cluster CR for the user to specify the number of days a crash entry should be kept for.
Providing a value for the field keepXDays creates a cronjob that runs every day at midnight, calling "ceph crash prune <keepXDays>".

Closes: https://github.com/rook/rook/issues/6332
Signed-off-by: Renan Campos <rcampos@redhat.com>
2020-11-24 17:06:18 -05:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
subhamkrai 49a3511532 ci: manifests changes to run gh action
to run the integration test there needs
to be a few changes in manifests like using
`deviceFilter` and other related changes.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-18 09:31:07 +05:30
subhamkrai a2d3d1d8a4 ceph: update snapshotterVersion to v3.0.0
in the integration test we still use
snapshotteVersion v2.1.0 but it's expected to
use v3.0.0.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-16 12:10:50 +05:30
Pete Birley 152a05c85e ceph: update to helm 3 for the rook chart
This updates the chart to make use of helm3 which has been released
for some time, and also permits CRDs to be installed pror to other objects
allowing the chart to be deployed at the same time as CRs for rook objects.

Co-authored-by: Pete Birley <pete@port.direct>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-11 14:41:24 -07:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
subhamkrai cb0ca66a6b ci: enable gosec linter in golangci-lint
golangci-lint linter gosec showing more errors
than gosec gh. This commit resolve
new errors. And, removing
nosec comments from autogenerated files.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-10-19 11:50:06 +05:30
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 4c11e45155 ceph: integration test on OpenShift
this commit enables the rook integration
testing capability on openshift. currently
this commit enable CephSmokeSuite testing
only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 14:30:57 +05:30
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
Madhu Rajanna e864b8d42d ceph: add E2E testing for snapshot and clone
Added E2E testing to create,delete and restore
a snapshot, create a pvc-pvc clone, install
and uninstall snapshot controller and snapshot
CRD.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-09-10 22:05:15 +05:30
Alexander Trost ed3c0eeff2 test: operator use rbac.authorization.k8s.io/v1 consistently
This replaces any K8S clientset usages of the
`rbac.authorization.k8s.io/v1beta1` APIs with
`rbac.authorization.k8s.io/v1` client.
This is done as some RBAC objects used the `v1beta1` and others already
using `v1`, for consistency.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2020-08-20 12:08:32 +02:00
Takashi IIGUNI 39fae17fb1 ci: cleanup clusterrolebindings
`anon-user-access` clusterrolebinding is created when installing rook operator.
However each test suite does not delete the clusterrolebinding,
and cause errors.
This PR skip creating the clusterrolebinding in each suite if it already exists.

Signed-off-by: Takashi IIGUNI <iiguni.tks@gmail.com>
2020-08-13 03:02:14 +00:00
binoue 937b77f8e6 ceph: revised according to the review comments
revised according to the review comments

Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
2020-07-30 07:44:08 +09:00
binoue 9a44b95f68 ceph: create event log file
Currently, rook CI is not logging Kubernetes Events, but it sometimes gives hints
for bugs.
Therefore, this PR adds a function to collect and log Kubernetes Events.

Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
2020-07-29 16:02:05 +09:00
Sébastien Han e4eaa91ede ceph: add rgw endpoint healthcheck
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.

A good status will look like:

status:
  endpointStatus:
    lastChanged: "2020-06-25T13:47:45Z"
    lastChecked: "2020-06-25T13:48:46Z"
  phase: Connected

A failed status:

status:
  endpointStatus:
    details: |-
      error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
      caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
    health: ERROR

This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.

Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-02 16:34:10 +02:00
Vineet Badrinath ce1003aef8 ceph: adds scripts and components to support admission controllers
adds deploy.sh script to deploy validatingwebhookconfiguration and create secrets.
adds new command ceph admission-controller to start webhook servers.
adds validation for various rook custom resources

Signed-off-by: Vineet Badrinath <vbadrina@redhat.com>
2020-06-24 14:59:00 +05:30
Satoru Takeuchi abdac2b412 tests: remove the unused fields of a struct
There are the unused fields in `struct CommandArgs`. These can be removed.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-05-13 20:52:54 +09:00
Satoru Takeuchi c0c0dc4ed9 tests: return bool if the functions are prefixed by "Is"
There are several functions that are prefixed by "Is" and return
just error. It's straightfoward to return bool to make the meaning
of these functions clearer.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-05-05 22:55:01 +09:00
Sébastien Han a044b27650 ci: fix MockExecuteCommandWithOutput mock definition
With the recent exex update we missed this one.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Sébastien Han ca0a30f38d ceph: convert Filesystem controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 23:34:58 +01:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 84b8cdcf75 exec: simplify the exec package from unused methods and logging
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 607e900f2c ceph: increase timeout for file test pod start
The integration tests have failed intermittently due to needing just
a little longer to start the file test pod. This increases the
wait timeout.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-19 13:34:09 -07:00
Travis Nielsen a1db52c5a0 ceph: move integration test to csi driver
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.

A number of other test improvements are also made for code
readability, test stability, and removing unused options.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-17 16:24:28 -07:00
Blaine Gardner 3cab1cf4fd Merge pull request #4608 from SUSE/upgrade-test-v1-1-to-v1-3
upgrade v1.0->v1.1->v1.2->v1.3 in upgrade test
2020-01-15 11:36:46 -07:00
Travis Nielsen 2d892b0e88 tests: allow integration tests in minimal config to run on multiple versions
When the tests run in a PR, they can only run against a single version
of K8s by default. If more than five k8s versions are supported in hte
CI, we will need to run some of the suites on multiple versions.
The versions are comma-separated in the list.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-13 16:54:49 -07:00
Blaine Gardner c40763c5c6 integration: retry pod logs with kubectl on fail
Sometimes we fail to get logs for pods using this method, notably the
operator pod. It is unknown why this happens. Pod logs are VERY
important, so try again using kubectl.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-01-13 08:56:33 -07:00
Blaine Gardner 3cc5bf5191 integration: k8s helper, raise RetryLoop by 50%
To stabilize tests, raise the `RetryLoop` variable 50%.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-12-11 11:27:36 -07:00
Blaine Gardner 20e5639461 ceph: integration:upgrade: long wait for mgr module
Reset the integration test k8s helper's `RetryLoop` to its original
value, and instead only wait an extra long time to allow the mgr module
updates to take a long time after Ceph is updated from Mimic to Nautilus
as part of Ceph's upgrade integration test.

Updating mgr modules can hang for quite a while, which causes the tests
to time out waiting for the OSDs to be updated. Allow this to take a
long time so the tests aren't as flaky.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-12-11 11:27:36 -07:00
Blaine Gardner b9a5717356 integration: raise retry loop count
Raise the retry loop count to stabilize the integration tests.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-11-13 12:23:18 -07:00
Blaine Gardner 79160abd18 ceph: in upgrade test, ensure legacy osds run
During upgrade tests, Rook should verify that it can still run legacy
OSDs. This includes directory-based OSDs, filestore disk OSDs, and
bluestore disk OSDs installed without ceph-volume (i.e., before mimic
v13.2.2) can still be run after upgrade.

This necessitates running the upgrade test twice; once with filestore
and once with bluestore.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-11-13 12:23:17 -07:00
Sébastien Han eccea9fa05 ci: more debug
When we give up on waiting for the pod to be running which describe it
to see what's wrong.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-10-22 15:14:28 +02:00