Commit Graph
126 Commits
Author SHA1 Message Date
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 4c11e45155 ceph: integration test on OpenShift
this commit enables the rook integration
testing capability on openshift. currently
this commit enable CephSmokeSuite testing
only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 14:30:57 +05:30
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
Madhu Rajanna e864b8d42d ceph: add E2E testing for snapshot and clone
Added E2E testing to create,delete and restore
a snapshot, create a pvc-pvc clone, install
and uninstall snapshot controller and snapshot
CRD.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-09-10 22:05:15 +05:30
Alexander Trost ed3c0eeff2 test: operator use rbac.authorization.k8s.io/v1 consistently
This replaces any K8S clientset usages of the
`rbac.authorization.k8s.io/v1beta1` APIs with
`rbac.authorization.k8s.io/v1` client.
This is done as some RBAC objects used the `v1beta1` and others already
using `v1`, for consistency.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2020-08-20 12:08:32 +02:00
Takashi IIGUNI 39fae17fb1 ci: cleanup clusterrolebindings
`anon-user-access` clusterrolebinding is created when installing rook operator.
However each test suite does not delete the clusterrolebinding,
and cause errors.
This PR skip creating the clusterrolebinding in each suite if it already exists.

Signed-off-by: Takashi IIGUNI <iiguni.tks@gmail.com>
2020-08-13 03:02:14 +00:00
binoue 937b77f8e6 ceph: revised according to the review comments
revised according to the review comments

Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
2020-07-30 07:44:08 +09:00
binoue 9a44b95f68 ceph: create event log file
Currently, rook CI is not logging Kubernetes Events, but it sometimes gives hints
for bugs.
Therefore, this PR adds a function to collect and log Kubernetes Events.

Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
2020-07-29 16:02:05 +09:00
Sébastien Han e4eaa91ede ceph: add rgw endpoint healthcheck
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.

A good status will look like:

status:
  endpointStatus:
    lastChanged: "2020-06-25T13:47:45Z"
    lastChecked: "2020-06-25T13:48:46Z"
  phase: Connected

A failed status:

status:
  endpointStatus:
    details: |-
      error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
      caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
    health: ERROR

This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.

Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-02 16:34:10 +02:00
Vineet Badrinath ce1003aef8 ceph: adds scripts and components to support admission controllers
adds deploy.sh script to deploy validatingwebhookconfiguration and create secrets.
adds new command ceph admission-controller to start webhook servers.
adds validation for various rook custom resources

Signed-off-by: Vineet Badrinath <vbadrina@redhat.com>
2020-06-24 14:59:00 +05:30
Satoru Takeuchi abdac2b412 tests: remove the unused fields of a struct
There are the unused fields in `struct CommandArgs`. These can be removed.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-05-13 20:52:54 +09:00
Satoru Takeuchi c0c0dc4ed9 tests: return bool if the functions are prefixed by "Is"
There are several functions that are prefixed by "Is" and return
just error. It's straightfoward to return bool to make the meaning
of these functions clearer.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-05-05 22:55:01 +09:00
Sébastien Han a044b27650 ci: fix MockExecuteCommandWithOutput mock definition
With the recent exex update we missed this one.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Sébastien Han ca0a30f38d ceph: convert Filesystem controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 23:34:58 +01:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 84b8cdcf75 exec: simplify the exec package from unused methods and logging
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 607e900f2c ceph: increase timeout for file test pod start
The integration tests have failed intermittently due to needing just
a little longer to start the file test pod. This increases the
wait timeout.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-19 13:34:09 -07:00
Travis Nielsen a1db52c5a0 ceph: move integration test to csi driver
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.

A number of other test improvements are also made for code
readability, test stability, and removing unused options.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-17 16:24:28 -07:00
Blaine Gardner 3cab1cf4fd Merge pull request #4608 from SUSE/upgrade-test-v1-1-to-v1-3
upgrade v1.0->v1.1->v1.2->v1.3 in upgrade test
2020-01-15 11:36:46 -07:00
Travis Nielsen 2d892b0e88 tests: allow integration tests in minimal config to run on multiple versions
When the tests run in a PR, they can only run against a single version
of K8s by default. If more than five k8s versions are supported in hte
CI, we will need to run some of the suites on multiple versions.
The versions are comma-separated in the list.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-13 16:54:49 -07:00
Blaine Gardner c40763c5c6 integration: retry pod logs with kubectl on fail
Sometimes we fail to get logs for pods using this method, notably the
operator pod. It is unknown why this happens. Pod logs are VERY
important, so try again using kubectl.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-01-13 08:56:33 -07:00
Blaine Gardner 3cc5bf5191 integration: k8s helper, raise RetryLoop by 50%
To stabilize tests, raise the `RetryLoop` variable 50%.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-12-11 11:27:36 -07:00
Blaine Gardner 20e5639461 ceph: integration:upgrade: long wait for mgr module
Reset the integration test k8s helper's `RetryLoop` to its original
value, and instead only wait an extra long time to allow the mgr module
updates to take a long time after Ceph is updated from Mimic to Nautilus
as part of Ceph's upgrade integration test.

Updating mgr modules can hang for quite a while, which causes the tests
to time out waiting for the OSDs to be updated. Allow this to take a
long time so the tests aren't as flaky.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-12-11 11:27:36 -07:00
Blaine Gardner b9a5717356 integration: raise retry loop count
Raise the retry loop count to stabilize the integration tests.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-11-13 12:23:18 -07:00
Blaine Gardner 79160abd18 ceph: in upgrade test, ensure legacy osds run
During upgrade tests, Rook should verify that it can still run legacy
OSDs. This includes directory-based OSDs, filestore disk OSDs, and
bluestore disk OSDs installed without ceph-volume (i.e., before mimic
v13.2.2) can still be run after upgrade.

This necessitates running the upgrade test twice; once with filestore
and once with bluestore.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-11-13 12:23:17 -07:00
Sébastien Han eccea9fa05 ci: more debug
When we give up on waiting for the pod to be running which describe it
to see what's wrong.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-10-22 15:14:28 +02:00
travisn f5eb11d863 tests: add logging to track down file test cleanup issue
Adding more logging until we can track down the file-test pod
cleanup issue

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-15 17:13:25 -06:00
Blaine Gardner 70810dd55a ceph: osd: do not init ceph.conf for dir OSDs
Just as with the mon, mgr, mds, rgw, and rbd-mirror daemons, do not
generate a ceph.conf in an init container for directory-based OSDs.
Instead use the mon config database and the commandline to supply all
the needed arguments for running these OSDs.

Also generate a keyring secret to mount to OSD pods. Use this
secret for directory-based OSDs for now with the intention to
use this for other OSDs in the future.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-10-15 12:12:46 -06:00
travisn a92850b307 tests: simplify filesystem tests and client pod deletion
The filesystem test deletes the client test pod by deleting the entire yaml
and sometimes hangs the integration tests. This is the most common intermittent
error in the CI. Now the deletion will happen more directly with a pod delete
command. If it does still fail deletion at least the deletion should timeout
sooner.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-14 09:29:45 -06:00
travisn bd672de7db ceph: allow running minimal test matrix for ceph tests
There are six test suites and six k8s versions where we currently run
the tests. For efficiency we can restrict the testing to one suite
per k8s version. Bigger or riskier changes should still run the full
set of suites on all versions. To trigger the smaller test matrix, add
[test ceph min] to the PR description

Signed-off-by: travisn <tnielsen@redhat.com>
2019-09-18 23:06:49 -06:00
travisn 86968c5064 ceph: boolean settings not applied if set to false and default
THe boolean helm settings were only being applied if their value was true.
If the desired value was false, the value would be skipped in the chart
instead of adding it with the value of false.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-09-09 23:27:35 +03:00
Maksim Nabokikh 179e1a310d ceph: Add dynamic flexvolume expansion
Allow to dynamic resize of rook volumes

Signed-off-by: Maksim Nabokikh <maksim.nabokikh@flant.com>
2019-09-05 02:40:22 +04:00
Travis Nielsen 35adf26d79 Merge pull request #3611 from phlogistonjohn/jjm-support-non-docker-in-tests
tests: do not require actual docker cli
2019-08-13 21:23:48 -06:00
Blaine Gardner 1fce599d4c Merge pull request #3613 from SUSE/integration-use-apply-instead-of-create
Integration tests: use 'kubectl apply' vs create
2019-08-13 16:21:10 -06:00
Blaine Gardner 9009326ad7 Integration tests: use 'kubectl apply' vs create
Some integration tests fail due to resources from a previous integration
run not being cleaned up properly. Instead of using 'kubectl create' --
which fails with an error if the resource already exists -- to create
resources, use 'kubectl apply' -- which does not fail for pre-existing
resources. 'kubectl apply' will give a warning that it should be used to
apply changes to resources created with 'create' or 'apply', but there
is no error.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-08-13 15:03:39 -06:00
John Mulligan 28cb59943b tests: do not require actual docker cli
Other CLI tools, such as podman, and equivalent to docker so support
setting an environment var to use it.

Signed-off-by: John Mulligan <jmulligan@redhat.com>
2019-08-13 10:43:10 -04:00
travisn 7e670aa01f tests: collect operator log after each failed test
Signed-off-by: travisn <tnielsen@redhat.com>
2019-08-13 07:23:00 -06:00
travisn be391de73c resolve changes from merge conflict in upgrade changes
The upgrade changes for #2901 added the check for the correct version
of the ceph image before continuing with an upgrade. This commit is
to refactor that change to work with the new code path to validate
the ceph version.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-08-13 07:23:00 -06:00
travisn 1598a048a1 tests: collect all logs for pods in the test namespaces
Now we collect logs for all pods and all their init and main
containers during the integration tests. No longer will we be
missing logs from the integration tests as long as the pods
are available when they are collected at the end of the test.
The pod descriptions are also written to a log file instead
of being included inline with the test output.

Signed-off-by: travisn <tnielsen@redhat.com>
(cherry picked from commit 6e0bc338f90b3f36a89cf37b542ff4ff873b70bd)
2019-08-02 11:10:46 -06:00
travisn ef9dc0be74 tests: ensure cluster does not exist before starting a test
The CI instances are not always being properly cleaned up
between runs. This is an attempt to get the tests
to ensure a clean install before proceeding with the test.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-07-05 16:41:52 -06:00
Santosh Pillai d6f63166a0 Optionally collect all CI logs using [all logs] flag
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2019-05-03 12:46:18 +05:30
travisn bc87b440fb update to the k8s 1.14 client libraries
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-18 07:51:41 -06:00
travisn 91e9a5cd04 reference K8s1.13 and client-go 1.10 and the external provisioner for flex
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-18 07:51:41 -06:00
travisn b8ab35e6c5 tests: collect the operator log prior to a pod restart
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-17 22:21:53 -06:00
travisn eece96eea4 tests: retry the file mount integration test for stability
Signed-off-by: travisn <tnielsen@redhat.com>
2019-03-12 17:30:06 -06:00
Blaine Gardner 51cc7c2e85 Ceph integration tests: add more verbose logging
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-03-05 14:14:09 -07:00
travisn b24ebc8168 tests: remove unnecessary string returns from kubectl calls
Various kubectl helper methods return strings from calls to
create, delete, or apply resources. There is no need for this.
It is sufficient and complete to check the err from these calls to determine failure.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-03-01 16:16:41 -07:00
travisn 0230400e27 build: remove 1.8 and 1.9 and add 1.13 to integration tests
Signed-off-by: travisn <tnielsen@redhat.com>
2018-12-19 23:42:01 -07:00