Commit Graph
289 Commits
Author SHA1 Message Date
Travis Nielsen 53ed11f15b build: remove the edgefs operator from rook
The EdgeFS operator has been deprecated for some time in Rook.
If the replacement is added back to Rook it can be completed
according to the new guidelines in the documentation.
https://rook.io/docs/rook/master/storage-providers.html

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-25 17:51:24 -07:00
Travis Nielsen 2294851ca3 build: remove the cockroachdb operator from rook
The cockroachDB operator has not had community support in Rook.
Therefore, the time has come to deprecate and remove it.
If the sources are still needed, there is always git history.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-20 17:28:58 -07:00
Blaine Gardner 9b0ba6ae8b ceph: add obc to upgrade test
Add object bucket claim to Ceph upgrade test.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-08 09:56:00 -07:00
Blaine Gardner 40fd80cf14 ceph: update smoke test to verify obc is bound
When verifying OBC creation, validating that secret and configmap exist
is good, but the definitive validation is to check that the OBC's phase
is "Bound".

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-07 13:58:22 -07:00
Travis Nielsen ba0d7288d8 ceph: delete object store in test in a defer method
Always attempt to delete the object store in case some test
fails and aborts the remainder of the test.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-17 18:21:16 -07:00
Travis Nielsen 114a949cfb ceph: test latest ceph on raw devices and previous ceph on partitions
The latest ceph-volume does not support creating OSDs on partitions.
The github actions only have a partition available, so we will run
the github tests on v14.2.12 and v15.2.7 that still support partitions,
while the Jenkins environment has a raw device available where we
can run the latest versions of Ceph in the tests.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-17 18:21:16 -07:00
Renan Campos ad9dcb0c05 ceph: periodically prune crash entries older than user-provided days
Rook's crashcollector pod posts entries to the ceph cluster when a crash occurs.
Over time the number cluster may hold crash entries needlessly.
To clean up old crash entries, this PR adds a field to the ceph cluster CR for the user to specify the number of days a crash entry should be kept for.
Providing a value for the field keepXDays creates a cronjob that runs every day at midnight, calling "ceph crash prune <keepXDays>".

Closes: https://github.com/rook/rook/issues/6332
Signed-off-by: Renan Campos <rcampos@redhat.com>
2020-11-24 17:06:18 -05:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Pete Birley 152a05c85e ceph: update to helm 3 for the rook chart
This updates the chart to make use of helm3 which has been released
for some time, and also permits CRDs to be installed pror to other objects
allowing the chart to be deployed at the same time as CRs for rook objects.

Co-authored-by: Pete Birley <pete@port.direct>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-11 14:41:24 -07:00
Travis Nielsen 61539b6699 Merge pull request #6559 from BlaineEXE/upgrade-test-1.5
ceph: upgrade test Rook only from v1.4 to master
2020-11-11 10:13:39 -07:00
Blaine Gardner 2f034824cb ceph: upgrade test Rook only from v1.4 to master
Since Rook no longer supports legacy filestore devices, there is
no need to keep testing upgrades from v1.2 all the way to master.
We can now just test v1.4 to master.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-11-10 15:44:05 -07:00
Travis Nielsen f1287bdb19 ceph: skip csi tests on k8s 1.13
The csi tests require 1.14 or newer to run

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-06 11:16:32 -07:00
Travis Nielsen 556de488c0 nfs: skip running tests on older than k8s 1.14
NFS tests are not supported on older than K8s 1.14.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 4adc54e6f1)
2020-11-06 10:47:26 -07:00
Lalit Maganti c6aec79c4f ceph: add option to preserve filesystem on CRD deletion
Due to #6492, preservePoolsOnDelete is not useful at all for CephFS;
after the filesystem is deleted, the leftover pools cannot be
reassocaited with a newly created filesystem without wiping all
metadata. The only way we can actually preserve data is keeping around
the entire filesystem.

This commit implements a `preserveFilesystemOnDelete` option which work
similar to the existing pool preservation option but instead keeps the
whole CephFS while taking it down and removing all MDSes.

This commit also changes all documentation to refer to this new option
with the intent of essentially deprecating `preservePoolsOnDelete`. IMO,
keeping around `preservePoolsOnDelete` is actively harmful because it
lulls users into thinking their data will be safe but, in reality,
recovering from this situation is highly complex and has large potential
for data loss.

Signed-off-by: Lalit Maganti <lalitm@google.com>
2020-10-30 17:04:18 +00:00
subhamkrai cb0ca66a6b ci: enable gosec linter in golangci-lint
golangci-lint linter gosec showing more errors
than gosec gh. This commit resolve
new errors. And, removing
nosec comments from autogenerated files.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-10-19 11:50:06 +05:30
Satoru Takeuchi 90c03c814a ci: fix tests/README to use proper test command
In local environment, We should use `go test` than the test binary
under "_output/".

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-10-09 22:20:09 +00:00
Satoru Takeuchi 9cd42a17a4 ci: fix an intermittent failure of ceph flex suite
Sometimes CephFlexSuite fails with the following backtrace.

```
--- FAIL: TestCephFlexSuite (584.14s)
    --- PASS: TestCephFlexSuite/TestBlockStorageMountUnMountForDifferentAccessModes (163.66s)
    --- PASS: TestCephFlexSuite/TestBlockStorageMountUnMountForStatefulSets (83.21s)
    --- FAIL: TestCephFlexSuite/TestFileSystem (119.51s)
        ceph_base_file_test.go:487:
                Error Trace:    ceph_base_file_test.go:487
                                                        ceph_base_file_test.go:265
                                                        ceph_flex_test.go:113
                Error:          Expected nil, but got: &errors.errorString{s:"kubectl exec command bash failed on pod rook-ceph-tools in namespace flex-ns. Failed to run: kubectl [exec -n flex-ns rook-ceph-tools -- bash -c mount -t ceph -o mds_namespace=smoke-test-fs,name=admin,secret=$(grep key /etc/ceph/keyring | awk '{print $3}') $(grep mon_host /etc/ceph/ceph.conf | awk '{print $3}'):/ /tmp/testrook] : exit status 32"}
                Test:           TestCephFlexSuite/TestFileSystem
```

It's due to insufficient retry of mount.

There is only one difference between CI logs of both "PASS" cases and "FAIL" cases.

The log of PASS cases:
```
2020-09-30 21:47:03.369019 D | exec: Running command: kubectl exec -n flex-ns rook-ceph-tools -- bash -c mount -t ceph -o mds_namespace=smoke-test-fs,name=admin,secret=$(grep key /etc/ceph/keyring | awk '{print $3}') $(grep mon_host /etc/ceph/ceph.conf | awk '{print $3}'):/ /tmp/testrook
2020-09-30 21:47:03.679359 D | exec: Running command: kubectl exec -n flex-ns rook-ceph-tools -- mkdir -p /tmp/testrook/foo
```

The log of FAIL cases:
```
2020-09-30 21:49:21.591702 D | exec: Running command: kubectl exec -n flex-ns rook-ceph-tools -- bash -c mount -t ceph -o mds_namespace=smoke-test-fs,name=admin,secret=$(grep key /etc/ceph/keyring | awk '{print $3}') $(grep mon_host /etc/ceph/ceph.conf | awk '{print $3}'):/ /tmp/testrook
2020-09-30 21:49:36.592161 I | exec: timeout waiting for process kubectl to return. Sending interrupt signal to the process
2020-09-30 21:49:40.355420 E | utils: Failed to execute: kubectl [exec -n flex-ns rook-ceph-tools -- bash -c mount -t ceph -o mds_namespace=smoke-test-fs,name=admin,secret=$(grep key /etc/ceph/keyring | awk '{print $3}') $(grep mon_host /etc/ceph/ceph.conf | awk '{print $3}'):/ /tmp/testrook] : timeout waiting for the command kubectl to return.
```

Everything looks fine before showing the above-mentioned error in FAIL cases.
In addition, this problem hasn't happened in high-spec local machine
over 10 times.

Related issue: 6358

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-10-07 09:36:44 +00:00
subhamkrai 5be7f0ae07 ceph: flex test failing in CI
this commit will add error not found condition
check before assert delete on PVC and
statefulset due to whch flex test was failing.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-10-05 20:00:01 +05:30
Travis Nielsen c13d1edf22 ceph: skip smoke suite on k8s 1.11
The smoke suite relies on the csi driver for all the file and block
tests. In the release branch on k8s 1.11 we only need to run the
flex suite and can skip the smoke suite.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-10-01 15:57:00 -06:00
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
Sébastien Han d118bd38a6 Merge pull request #6278 from subhamkrai/golanci-lint-unused
ceph: handle golangci-lint linter unused
2020-09-18 12:44:37 +02:00
subhamkrai 4c11e45155 ceph: integration test on OpenShift
this commit enables the rook integration
testing capability on openshift. currently
this commit enable CephSmokeSuite testing
only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 14:30:57 +05:30
Sébastien Han 204ffd4065 Merge pull request #6250 from leseb/mirroring-config
ceph: add rbd-mirror configuration
2020-09-18 09:16:08 +02:00
subhamkrai f9fafe62d4 ceph: handle golangci-lint linter unused
this commit will enable one more linter
in golangci-lint.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 22:53:24 +05:30
Sébastien Han 451622a955 ceph: add rbd-mirror configuration
Rook is now capable of configuring mirroring between sites. The
implementation works at different levels:

* CephBlockPool: which introduces a new `mirroring` configuration as well
as `statusCheck`. When turned on, Rook will enable mirroring on the
pool. It will also create a bootstrap peer token and store it in a
Kubernetes Secret. The name of that Secret can be found in the Status
field of the CephBlockPool CRD. This token can be fetched and used by
other clusters to configure the site as a peer. Mirroring can be
configured either at the pool or the image level.

* CephRBDMirror: which introduces a new `peers` configuration allowing
Rook to connect to peers by passing a Secret name. The administrator will
create a Kubernetes Secret with 2 keys: 'token' for the bootstrap peer
token and 'pool' for the name of pool. Once detected the rbd-mirror
controller will go ahead and import the peer configuration.

Pool mirroring status example:

```
status:
  info:
    rbdMirrorBootstrapPeerSecretName: pool-peer-token-test
  mirroringInfo:
    lastChanged: "2020-09-17T14:47:27Z"
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        mode: image
        peers:
        - client_name: client.rbd-mirror-peer
          direction: rx-tx
          mirror_uuid: ""
          site_name: rhcs
          uuid: c50522a4-28a4-4bd3-ba68-e11780308882
        site_name: 91eae0dd-06b1-4d2c-91f3-1311c9df382b-rook-ceph
  mirroringStatus:
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        daemon_health: OK
        health: OK
        image_health: OK
        states:
          replaying: 1
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-17 19:05:04 +02:00
Travis Nielsen d55815d698 Merge pull request #6066 from Madhu-1/rbd-snap-clone-test
ceph: Add E2E testing for csi snapshot and clone
2020-09-14 11:24:58 -06:00
Travis Nielsen df47f58c05 ceph: run integration tests on k8s 1.19 for PRs
In K8s 1.19 the integration tests are not running in PRs since no test suites
were assigned to 1.19. Similarly, the flex suite was not being run since 1.14
was removed from the test matrix in master.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-09-10 11:04:40 -06:00
Madhu Rajanna e864b8d42d ceph: add E2E testing for snapshot and clone
Added E2E testing to create,delete and restore
a snapshot, create a pvc-pvc clone, install
and uninstall snapshot controller and snapshot
CRD.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-09-10 22:05:15 +05:30
Sébastien Han 7dea2b10ff ceph: disable ceph mgr test temporarily
The CI keeps failing intermittently due to
https://github.com/rook/rook/issues/5877 and the fix has not merged yet
so disabling until fixed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 11:24:09 +02:00
Sébastien Han 96b4a56a57 ceph: silence aws s3 sdk logs on healthcheck
We don't need to activate the debug logs on the objectstore
healthchecks.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Jared Watts 06ab07be8b Merge pull request #5863 from prksu/nfs-provisioner
nfs: nfs provisioner controlled by operator
2020-08-24 16:26:23 -07:00
Alexander Trost be16bc455f docs: operator use rbac.authorization.k8s.io/v1
This replaces any uses of `rbac.authorization.k8s.io/v1beta1` with
`rbac.authorization.k8s.io/v1` affecting RBAC objects like ClusterRole
and ClusterRoleBindings, etc. to make it consistent.
This is done as some RBAC objects used the `v1beta1` and others already
using `v1`.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2020-08-20 12:08:32 +02:00
binoue aa2af5a24b ci: clean up after SmokeTest
Currently, SmokeTest is not cleaned up after the test.
As this can affect other following tests, set rookCephCleanup as true.

Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-08-20 03:53:03 +00:00
binoue 5c897e4a21 ceph: fix error log
"Giving up waiting for Rook Toolbox to be ready" error always
logged even if Rook Toolbox is ready.
This commit fix the problem.

Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
2020-08-18 08:04:17 +09:00
Travis Nielsen 9168c22183 ceph: improve stability of object bucket tests
The object bucket tests will now always check if the conditions being waited
upon for creating users and buckets were completed, instead of allowing
the test to continue even when they didn't succeed.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-10 14:46:36 -06:00
Travis Nielsen 89284d9155 ceph: improve object bucket test stability
The object bucket check is sometimes not completed yet
that causes the integration test to fail intermittently.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-06 09:25:35 -06:00
tenzen-y 29087f8e41 ci: fix function comments
This commit will fix outdated function comments.

Signed-off-by: tenzen-y <toyonomajyutushi@yahoo.co.jp>
2020-08-05 08:08:53 +00:00
Travis Nielsen 772b7990db ceph: updated clusterroles applied in upgrade test
In the upgrade integration test we need to mirror what we expect
users to do during the upgrade by applying the new clusterroles.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-04 16:04:07 -06:00
Travis Nielsen e0ea87a353 Merge pull request #5674 from thotz/quotaobc
ceph: adding support for quota in object bucket claims
2020-07-30 22:51:58 -06:00
Sébastien Han 44355a921b Merge pull request #5920 from binoue/fix-log-output
ci: fix log output
2020-07-30 08:59:12 +02:00
binoue bb224733f6 ceph: fix log output
fix log output

Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
2020-07-30 07:26:20 +09:00
Madhu Rajanna 28f0471106 ceph: update clusterrole of operator for csidrivers
updated the clusterrole of the rook operator
to delete the csidrivers object when csi-drivers
are disabled, without the required access rook operator
cannot delete the csidrivers object.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-07-29 09:41:26 +05:30
Jiffin Tony Thottan f4240cb31d ceph: add quota support for obc
Closes: #5274
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-07-28 12:00:11 +05:30
Ahmad Nurus S 45f0debd00 nfs: nfs provisioner controlled by operator
Signed-off-by: Ahmad Nurus S <prksu.sh@gmail.com>
2020-07-27 12:24:32 +07:00
Sébastien Han e517ac96aa ceph: fail if the pool fails to be deleted
Let's assert if the pool is still present.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-22 17:25:15 +02:00
Sébastien Han 8729a206f3 ceph: let the healthcheck set the status phase
Instead of setting the Phase of the CR once the reconcile is done, let's
actually set in from the healthcheck so it is more accurate.

Closes: https://github.com/rook/rook/issues/5249
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-22 17:25:12 +02:00
Sébastien Han 698f46e978 ci: purge after flexsuite
If we don't the mgr tests cannot verify the osd creation and will fail
with:

```
--- FAIL: TestCephMgrSuite (169.32s)
    --- FAIL: TestCephMgrSuite/TestCreateOSD (1.65s)
        ceph_mgr_test.go:217:
            	Error Trace:	ceph_mgr_test.go:217
            	Error:      	Should not be: ""
            	Test:       	TestCephMgrSuite/TestCreateOSD
            	Messages:   	No devices available to create test OSD
        ceph_mgr_test.go:218:
            	Error Trace:	ceph_mgr_test.go:218
            	Error:      	Should not be: ""
            	Test:       	TestCephMgrSuite/TestCreateOSD
            	Messages:   	No nodes available to create test OSD
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-21 17:10:28 +02:00
Sébastien Han 94c8546d31 ceph: fix some mgr test
Fail if the orch status never gets ready.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-17 14:03:58 +02:00