Commit Graph
211 Commits
Author SHA1 Message Date
Travis Nielsen ba3e601a50 ceph: update k8s versions for ceph tests on 1.18
With 1.18 in the test matrix we need to update which versions
will run the different ceph test suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-30 16:51:02 -06:00
Santosh Pillai 2bfd7c42bc ceph: cleanup cluster.Spec.DataDirHostPath on cluster deletion
In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-03-30 19:36:01 +05:30
Umanga Chapagain ac9244fc1a Ceph: integration test creates ObjectStoreUser before ObjectStore
This fix is added to ensure that everything works as expected
when user tries to create ObjectStoreUser before creating
ObjectStore itself. ObjectStoreUser will wait for ObjectStore
to be up and running.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-27 18:16:28 +05:30
Satoru Takeuchi a5ea2858c9 manifests: remove unnecessary namespace fields
There are many namespace fields in ClusterRole{,Binding}. However,
ClusterRole{,Binding} are not namespaced. So we can remove these.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-03-18 15:58:31 +00:00
Travis Nielsen ce1fedf884 ci: only list the images in the needed pool during cleanup
The pool cleanup only needs to happen for an individual pool.
No need to query the block images in all pools. One of the rgw
pools is periodically causing a hang when it is queried,
but there is no need to query for it when we are cleaning
up the pool tests.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-17 15:28:17 -06:00
Sébastien Han a8159c62c1 ci: enable DEBUG logging
Activate DEBUG logging for the integration tests.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Sébastien Han 44a4c1826d ci: remove object store when done
Purge the object store properly, otherwise the cephobjectstore CRD won't
be deleted since a finalizer is in place for the CR.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Travis Nielsen 4ce9b45ffd ceph: remove dead test code for cluster purging
Cluster purging is no longer called, therefore can be removed.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-05 15:39:39 -07:00
Travis Nielsen 65f3f84ab0 tests: ensure pools are purged during integration tests
With a finalizer on the pools, the pools were not always being purged
during the integration tests. Now the multicluster suite will ensure
its pool is purged.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-03 19:44:23 -07:00
Satoru Takeuchi b1dd05a59a ceph: simplify verbose function names and a wrong var name
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-02-29 06:35:01 +09:00
Travis Nielsen a1db52c5a0 ceph: move integration test to csi driver
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.

A number of other test improvements are also made for code
readability, test stability, and removing unused options.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-17 16:24:28 -07:00
Sébastien Han 8c58cdbc78 ci: fix upgrade ordering
We now test the upgrade from M to N before the upgrade to master.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-13 09:17:34 +01:00
Travis Nielsen 9855db09bd ceph: describe file-test pod upon failure to terminate
The file-test consumer pod intermittently fails to stop during the
integration tests. When this happens, print the pod description
to see if there are events indicating the issue.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-04 12:17:27 -07:00
Travis Nielsen 8659cdbb3d ceph: refactor upgrade integration test checks for daemon versions
The upgrade test had duplicated code for checking that daemons
were upgraded after the Rook upgrade vs the Ceph upgrade.
This is now refactored for consistent checks after upgrade.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-04 12:17:27 -07:00
Travis Nielsen babd65d131 ceph: simplify most integration tests to single mds
The integration tests were always running two MDS daemons active,
with two standby. This is now parameterized so the test can request
how many MDS daemons to run. The smoke suite will run two active
and the rest will just run a single active MDS.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-04 12:17:27 -07:00
Travis Nielsen 97c4cf1631 ceph: reenable the upgrade integration test
The integration test was disabled due to removing support of pre-ceph-volume
OSDs as well as OSDs on directories. Now we re-enable the tests, with the
following approach:
- The base install in Rook v1.1 with the latest mimic release that
had c-v support
- The upgrade goes from v1.1 to Rook v1.2 then Rook master
- The final upgrade step is from mimic to nautilus

For efficiency the skipUpgradeChecks flag is added.
Logs are also collected between each upgrade step to improve
troubleshooting.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-04 12:17:24 -07:00
Ashish Ranjan 94d9de27d9 ceph: Adds integration tests for OSDs over PVC
This commit enables testing OSDs over PVCs by modifying the
multi-cluster test to use PVC for provisioning OSDs when `manual`
storageClass is present in the cluster.

Signed-off-by: Ashish Ranjan <aranjan@redhat.com>
2020-02-04 19:09:49 +05:30
Travis Nielsen 55b4cdabd8 Merge pull request #4742 from leseb/external-cluster-test
ci: add external cluster functional test
2020-01-29 16:16:25 -07:00
Sébastien Han 84fc10cdc1 ci: add external cluster functional test
Add coverage in the CI for the external cluster feature.

Closes: https://github.com/rook/rook/issues/3691
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-29 20:27:15 +01:00
Travis Nielsen 153f28344e ceph: skip the upgrade checks in tests
The integration tests only have a single node so the upgrade
checks do not provide any real safety for the cluster during
the upgrade. We will be able to improve the reliability and speed
of the tests by disabling these checks.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-27 15:23:57 -07:00
Sébastien Han 9e4c8092fc ci: re-enable multi-cluster test
We are about to add support for external cluster in the CI so let's
re-enable this scenario.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-27 22:56:01 +01:00
Sébastien Han 0e41b45e43 ci: disable multi-cluster test
Since filestore is not supported anymore, this scenario might not be
needed anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-24 12:18:21 +01:00
Sébastien Han 2245deb1a1 ci: temporarily disable upgrade suite
With the recent change, the upgrade suite needs to be adapted to only
test a c-v transition.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-24 12:18:21 +01:00
Sébastien Han 2367b65d4a ci: print more debug
Add more info on error to ease debugging.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-24 12:18:21 +01:00
Sébastien Han 227d2d527a ceph: osd store refactor
Multiple things:

1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.

2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.

3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.

Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-23 19:13:09 +01:00
Blaine Gardner 3cab1cf4fd Merge pull request #4608 from SUSE/upgrade-test-v1-1-to-v1-3
upgrade v1.0->v1.1->v1.2->v1.3 in upgrade test
2020-01-15 11:36:46 -07:00
Travis Nielsen 2d892b0e88 tests: allow integration tests in minimal config to run on multiple versions
When the tests run in a PR, they can only run against a single version
of K8s by default. If more than five k8s versions are supported in hte
CI, we will need to run some of the suites on multiple versions.
The versions are comma-separated in the list.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-13 16:54:49 -07:00
Blaine Gardner 2e242ca886 ceph: upgrade test versions v1.0->v1.1->v1.2->v1.3
In the Ceph upgrade integration test, upgrade from Rook v1.0 with Ceph
Mimic v13 to install legacy disk-based OSDs. Then upgrade Ceph to
Nautilus, then Rook to v1.1, then v1.2, then v1.3 (master currently) to
make sure that Rook is able to run legacy OSDs created without
ceph-volume.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-01-13 08:56:35 -07:00
Travis Nielsen a02c2c3089 tests: run integration tests on k8s 1.17 and remove 1.12
With the release of K8s 1.17 Rook needs to test on this new version.
The integration tests will now run the tests across K8s 1.13-1.17.
The pattern has been to run the tests across the most recent five
versions.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-08 07:18:53 -08:00
Mateusz Los cd713b11cb ceph: client CRD test fixes
wait for client to be updated before verifying caps

Signed-off-by: Mateusz Los <los.mateusz@gmail.com>
2019-12-16 23:06:33 +01:00
Blaine Gardner dceeff22e1 ceph: integration: kill process grp on file-test
The file-test pod still hangs sometimes when exiting. Instead of
trapping SIGTERM as before, use tini's `-g` option to kill the process
group. According to further reading on bash signal handling [1], Ctrl-C
sends a kill signal to the entire process group in order to kill `sleep`
commands.

[1] Link also noted in code comments:
http://mywiki.wooledge.org/SignalTrap#When_is_the_signal_handled.3F

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-12-11 11:27:37 -07:00
Blaine Gardner 20e5639461 ceph: integration:upgrade: long wait for mgr module
Reset the integration test k8s helper's `RetryLoop` to its original
value, and instead only wait an extra long time to allow the mgr module
updates to take a long time after Ceph is updated from Mimic to Nautilus
as part of Ceph's upgrade integration test.

Updating mgr modules can hang for quite a while, which causes the tests
to time out waiting for the OSDs to be updated. Allow this to take a
long time so the tests aren't as flaky.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-12-11 11:27:36 -07:00
Mateusz Los 50f3d29cf4 ceph: client controller tests and docs
ceph client CRD refactoring
Add info about client crd to documentation
Add tests for client controller CRD

Signed-off-by: Mateusz Los <los.mateusz@gmail.com>
2019-12-09 12:48:44 +01:00
Blaine Gardner 79160abd18 ceph: in upgrade test, ensure legacy osds run
During upgrade tests, Rook should verify that it can still run legacy
OSDs. This includes directory-based OSDs, filestore disk OSDs, and
bluestore disk OSDs installed without ceph-volume (i.e., before mimic
v13.2.2) can still be run after upgrade.

This necessitates running the upgrade test twice; once with filestore
and once with bluestore.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-11-13 12:23:17 -07:00
Sébastien HanandRohan CJ 90cd8224f4 ceph: add ceph-crash controller
A new controller to bootstrap ceph-crash pod on Ceph nodes running Ceph
pods only.

This implementation is unfortunately a mix of 'hacks' to make the daemon
working correclty. The ceph-crash script faced different issues:

* it's not a daemon
* it does not have any key in cephx
* it runs ceph commands under the hood which need an admin key and a
ceph.conf

On upstream Ceph, ceph-crash needs to grow better, once that happens we
will improve our implementation.

This also enhances osd provisionConfig struct with DataPathMap

Pod volumes and volume mount needs DataPathMap to perform the right
actions on log and crash dir. Exposing DataPathMap makes that possible.

Obviously, this is exposing ceph crash reports on the host as
well as pushing them into the mgr.
Basically, if a daemon fails, it'll put its core dump into
/var/lib/ceph/crash, bindmounting this dir on the host, ensures that the
crashes don't get lost when the pod dies.

Example:

```
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !] kubectl -n rook-ceph exec -ti rook-ceph-crashcollector-minikube-574858b99c-zvg4z bash
bash: warning: setlocale: LC_CTYPE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_COLLATE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_MESSAGES: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_NUMERIC: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_TIME: cannot change locale (en_US.UTF-8): No such file or directory
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ceph crash ls

[root@rook-ceph-osd-2-7644f99695-cljzh /]# pidof ceph-osd
13258 5415 5397
[root@rook-ceph-osd-2-7644f99695-cljzh /]# kill -SIGABRT 13258
[root@rook-ceph-osd-2-7644f99695-cljzh /]# ls /var/lib/ceph/crash/
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96     posted

... wait maximum 10 min (ceph-crash scraps every 10 minutes)
... the container will exit

[root@rook-ceph-osd-2-7644f99695-cljzh /]# exit
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !]
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !] kubectl -n rook-ceph exec -ti rook-ceph-crashcollector-minikube-574858b99c-zvg4z bash
bash: warning: setlocale: LC_CTYPE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_COLLATE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_MESSAGES: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_NUMERIC: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_TIME: cannot change locale (en_US.UTF-8): No such file or directory
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]#
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]#
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ceph crash ls
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96 osd.1
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ls /var/lib/ceph/crash/
posted
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ls /var/lib/ceph/crash/posted/
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# tail /var/lib/ceph/crash/posted/2019-11-12_12\:56\:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96/log
   1/ 5 mgr
   1/ 5 mgrc
   1/ 5 dpdk
   1/ 5 eventtrace
  -2/-2 (syslog threshold)
  -1/-1 (stderr threshold)
  max_recent     10000
  max_new         1000
  log_file /var/lib/ceph/crash/2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96/log
--- end dump of recent events ---
```

Signed-off-by: Sébastien Han <seb@redhat.com>
Co-authored-by: Rohan CJ <rohantmp@gmail.com>
2019-11-12 16:06:07 +01:00
Travis Nielsen 30f3535116 Merge pull request #4149 from umangachapagain/objectuser-retry
adds retry to ObjectUser creation
2019-11-07 17:45:26 -06:00
Umanga Chapagain 69c1c4c8ff Ceph: Create ObjectStore before ObjectStoreUser
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2019-11-06 14:04:52 +05:30
Umanga Chapagain c90c5f8711 ceph: fixes mon failover test
Real mons and canary both have the same rook-ceph-mon label.
It creates a confusion while retrieving list of mon deploymets
using label selector. This commit adds a check to remove canary
deployments from the deployment list.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2019-11-05 12:47:56 +05:30
Umanga Chapagain 6dd474d6f4 ceph: adds retry to ObjectUser creation
retry ObjectUser creation as long as valid input is provided
or it times out. Retry every 15sec for 5min.

integration test creates ObjectUser before ObjectStore to
ensure that ObjectUser waits for ObjectStore to be created
and available

Closes: https://github.com/rook/rook/issues/3937
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2019-11-05 12:45:45 +05:30
Travis Nielsen 3767d6e14b tests: initialize the cluster-specific rbac properly
The integration tests were creating several namespace-specific resources
in the system namespace instead of the namespace with the cluster. This
was only affecting the rook-ceph-osd and rook-ceph-mgr service accounts,
which hadn't been exercised until the nodes were queried by osds for
topology awareness.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2019-11-04 23:07:13 -07:00
Travis Nielsen e32729269f ceph: clean up block integration tests more reliably
The block integration tests are a source of intermittent failures
in the CI. The PVCs were not being confirmed as removed before the
pools were deleted. Then the pool would not be removed since there
might still be a block image from the PVC that wasn't deleted yet.
Now the tests will ensure the PVCs and their block images are
removed before the pool is deleted.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2019-10-30 15:37:35 -06:00
Travis Nielsen f5b032ad8e ceph: improve reliability of mon failover test
Depending on the state of the orchestration, the operator might trigger
re-creation of the deleted mon. In this case, consider the test successful
rather than wait for the failover which will never occur.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2019-10-24 18:55:01 -06:00
Sébastien Han fcc2e21faf ceph: rework csi keys and secrets
We know create 4 secrets containing 4 ceph keys where each of them have
limited permissions to access the cluster.

Closes: https://github.com/rook/rook/issues/4074
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-10-22 15:14:28 +02:00
travisn 4ef3ffb8e4 test: add k8s 1.16 to the integration test matrix
With the desire to run the integration tests on the five
most recent versions of K8s, 1.16 is added and 1.11
is removed from the integration test matrix.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-19 12:46:01 -06:00
travisn 64dce0c4e4 ceph: prevent object user test from crashing on failure
Require retrieving the user to succeed or else abort the
object integration test

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-19 12:46:00 -06:00
travisn 0cca2cdef4 tests: increase wait timeout for mon failover tests
The timeout of 80 seconds wasn't always sufficient for mon failover.
The timeout is now increased to 150 seconds.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-15 17:13:26 -06:00
travisn 378f209d63 tests: ceph multicluster tests to wait for one cluster at a time
The two clusters were being created in parallel by the tests
and sometimes causing the test to timeout. The operator handles
the cluster serially so we don't gain anything by starting
them asynchronously. With this change the tests will be started
serially to match the behavior of the operator and avoid
the issues with test timeouts.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-15 17:13:26 -06:00
travisn 98c21950aa tests: mark failed test for log collection
The logs are only collected by default if a test failed.
Ensure that we collect the logs if the test setup step
failed.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-15 17:13:25 -06:00
travisn a178cccd77 tests: consolidate block integration tests
The BlockCreateSuite and BlockMountUnmountSuite suites are almost identical
for running block tests with different types of mounts. We can
consolidate these into a single test suite and eliminate a couple of
the tests that are already covered by other tests.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-15 17:13:25 -06:00