With 1.18 in the test matrix we need to update which versions
will run the different ceph test suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
This fix is added to ensure that everything works as expected
when user tries to create ObjectStoreUser before creating
ObjectStore itself. ObjectStoreUser will wait for ObjectStore
to be up and running.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
There are many namespace fields in ClusterRole{,Binding}. However,
ClusterRole{,Binding} are not namespaced. So we can remove these.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The pool cleanup only needs to happen for an individual pool.
No need to query the block images in all pools. One of the rgw
pools is periodically causing a hang when it is queried,
but there is no need to query for it when we are cleaning
up the pool tests.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Purge the object store properly, otherwise the cephobjectstore CRD won't
be deleted since a finalizer is in place for the CR.
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
With a finalizer on the pools, the pools were not always being purged
during the integration tests. Now the multicluster suite will ensure
its pool is purged.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.
A number of other test improvements are also made for code
readability, test stability, and removing unused options.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The file-test consumer pod intermittently fails to stop during the
integration tests. When this happens, print the pod description
to see if there are events indicating the issue.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The upgrade test had duplicated code for checking that daemons
were upgraded after the Rook upgrade vs the Ceph upgrade.
This is now refactored for consistent checks after upgrade.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests were always running two MDS daemons active,
with two standby. This is now parameterized so the test can request
how many MDS daemons to run. The smoke suite will run two active
and the rest will just run a single active MDS.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration test was disabled due to removing support of pre-ceph-volume
OSDs as well as OSDs on directories. Now we re-enable the tests, with the
following approach:
- The base install in Rook v1.1 with the latest mimic release that
had c-v support
- The upgrade goes from v1.1 to Rook v1.2 then Rook master
- The final upgrade step is from mimic to nautilus
For efficiency the skipUpgradeChecks flag is added.
Logs are also collected between each upgrade step to improve
troubleshooting.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit enables testing OSDs over PVCs by modifying the
multi-cluster test to use PVC for provisioning OSDs when `manual`
storageClass is present in the cluster.
Signed-off-by: Ashish Ranjan <aranjan@redhat.com>
The integration tests only have a single node so the upgrade
checks do not provide any real safety for the cluster during
the upgrade. We will be able to improve the reliability and speed
of the tests by disabling these checks.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Multiple things:
1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.
2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.
3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.
Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
When the tests run in a PR, they can only run against a single version
of K8s by default. If more than five k8s versions are supported in hte
CI, we will need to run some of the suites on multiple versions.
The versions are comma-separated in the list.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In the Ceph upgrade integration test, upgrade from Rook v1.0 with Ceph
Mimic v13 to install legacy disk-based OSDs. Then upgrade Ceph to
Nautilus, then Rook to v1.1, then v1.2, then v1.3 (master currently) to
make sure that Rook is able to run legacy OSDs created without
ceph-volume.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
With the release of K8s 1.17 Rook needs to test on this new version.
The integration tests will now run the tests across K8s 1.13-1.17.
The pattern has been to run the tests across the most recent five
versions.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The file-test pod still hangs sometimes when exiting. Instead of
trapping SIGTERM as before, use tini's `-g` option to kill the process
group. According to further reading on bash signal handling [1], Ctrl-C
sends a kill signal to the entire process group in order to kill `sleep`
commands.
[1] Link also noted in code comments:
http://mywiki.wooledge.org/SignalTrap#When_is_the_signal_handled.3F
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Reset the integration test k8s helper's `RetryLoop` to its original
value, and instead only wait an extra long time to allow the mgr module
updates to take a long time after Ceph is updated from Mimic to Nautilus
as part of Ceph's upgrade integration test.
Updating mgr modules can hang for quite a while, which causes the tests
to time out waiting for the OSDs to be updated. Allow this to take a
long time so the tests aren't as flaky.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
ceph client CRD refactoring
Add info about client crd to documentation
Add tests for client controller CRD
Signed-off-by: Mateusz Los <los.mateusz@gmail.com>
During upgrade tests, Rook should verify that it can still run legacy
OSDs. This includes directory-based OSDs, filestore disk OSDs, and
bluestore disk OSDs installed without ceph-volume (i.e., before mimic
v13.2.2) can still be run after upgrade.
This necessitates running the upgrade test twice; once with filestore
and once with bluestore.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
A new controller to bootstrap ceph-crash pod on Ceph nodes running Ceph
pods only.
This implementation is unfortunately a mix of 'hacks' to make the daemon
working correclty. The ceph-crash script faced different issues:
* it's not a daemon
* it does not have any key in cephx
* it runs ceph commands under the hood which need an admin key and a
ceph.conf
On upstream Ceph, ceph-crash needs to grow better, once that happens we
will improve our implementation.
This also enhances osd provisionConfig struct with DataPathMap
Pod volumes and volume mount needs DataPathMap to perform the right
actions on log and crash dir. Exposing DataPathMap makes that possible.
Obviously, this is exposing ceph crash reports on the host as
well as pushing them into the mgr.
Basically, if a daemon fails, it'll put its core dump into
/var/lib/ceph/crash, bindmounting this dir on the host, ensures that the
crashes don't get lost when the pod dies.
Example:
```
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !] kubectl -n rook-ceph exec -ti rook-ceph-crashcollector-minikube-574858b99c-zvg4z bash
bash: warning: setlocale: LC_CTYPE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_COLLATE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_MESSAGES: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_NUMERIC: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_TIME: cannot change locale (en_US.UTF-8): No such file or directory
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ceph crash ls
[root@rook-ceph-osd-2-7644f99695-cljzh /]# pidof ceph-osd
13258 5415 5397
[root@rook-ceph-osd-2-7644f99695-cljzh /]# kill -SIGABRT 13258
[root@rook-ceph-osd-2-7644f99695-cljzh /]# ls /var/lib/ceph/crash/
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96 posted
... wait maximum 10 min (ceph-crash scraps every 10 minutes)
... the container will exit
[root@rook-ceph-osd-2-7644f99695-cljzh /]# exit
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !]
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !] kubectl -n rook-ceph exec -ti rook-ceph-crashcollector-minikube-574858b99c-zvg4z bash
bash: warning: setlocale: LC_CTYPE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_COLLATE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_MESSAGES: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_NUMERIC: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_TIME: cannot change locale (en_US.UTF-8): No such file or directory
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]#
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]#
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ceph crash ls
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96 osd.1
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ls /var/lib/ceph/crash/
posted
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ls /var/lib/ceph/crash/posted/
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# tail /var/lib/ceph/crash/posted/2019-11-12_12\:56\:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96/log
1/ 5 mgr
1/ 5 mgrc
1/ 5 dpdk
1/ 5 eventtrace
-2/-2 (syslog threshold)
-1/-1 (stderr threshold)
max_recent 10000
max_new 1000
log_file /var/lib/ceph/crash/2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96/log
--- end dump of recent events ---
```
Signed-off-by: Sébastien Han <seb@redhat.com>
Co-authored-by: Rohan CJ <rohantmp@gmail.com>
Real mons and canary both have the same rook-ceph-mon label.
It creates a confusion while retrieving list of mon deploymets
using label selector. This commit adds a check to remove canary
deployments from the deployment list.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
retry ObjectUser creation as long as valid input is provided
or it times out. Retry every 15sec for 5min.
integration test creates ObjectUser before ObjectStore to
ensure that ObjectUser waits for ObjectStore to be created
and available
Closes: https://github.com/rook/rook/issues/3937
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
The integration tests were creating several namespace-specific resources
in the system namespace instead of the namespace with the cluster. This
was only affecting the rook-ceph-osd and rook-ceph-mgr service accounts,
which hadn't been exercised until the nodes were queried by osds for
topology awareness.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The block integration tests are a source of intermittent failures
in the CI. The PVCs were not being confirmed as removed before the
pools were deleted. Then the pool would not be removed since there
might still be a block image from the PVC that wasn't deleted yet.
Now the tests will ensure the PVCs and their block images are
removed before the pool is deleted.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Depending on the state of the orchestration, the operator might trigger
re-creation of the deleted mon. In this case, consider the test successful
rather than wait for the failover which will never occur.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the desire to run the integration tests on the five
most recent versions of K8s, 1.16 is added and 1.11
is removed from the integration test matrix.
Signed-off-by: travisn <tnielsen@redhat.com>
The timeout of 80 seconds wasn't always sufficient for mon failover.
The timeout is now increased to 150 seconds.
Signed-off-by: travisn <tnielsen@redhat.com>
The two clusters were being created in parallel by the tests
and sometimes causing the test to timeout. The operator handles
the cluster serially so we don't gain anything by starting
them asynchronously. With this change the tests will be started
serially to match the behavior of the operator and avoid
the issues with test timeouts.
Signed-off-by: travisn <tnielsen@redhat.com>
The logs are only collected by default if a test failed.
Ensure that we collect the logs if the test setup step
failed.
Signed-off-by: travisn <tnielsen@redhat.com>
The BlockCreateSuite and BlockMountUnmountSuite suites are almost identical
for running block tests with different types of mounts. We can
consolidate these into a single test suite and eliminate a couple of
the tests that are already covered by other tests.
Signed-off-by: travisn <tnielsen@redhat.com>