rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
We have seen cases where the signed certificate used for the RGW does not
contain the internal DNS endpoint, resulting in the health check to fail
since the certificate is not valid for this domain.
People consuming the gateways by external clients and for specific
domains do not necessarily have the internal DNS configured in the
certificate.
So let's be a bit more flexible and simply ensure a connectivity check
and bypass the certificate validation.
Also, this is fixing the tls code in newS3Agent and adds unit tests.
Closes: #8663
Signed-off-by: Sébastien Han <seb@redhat.com>
In general, make the output from the object store test easier to follow
and debug.
- Make sure assertions are associated with the appropriate subtest.
- Remove `t.Helper()` from complex helpers.
- Add info to log messages.
- Move long parts of (sub-)test names to log messages.
- Move some reused functions back to ceph_base_object_test.go
- Refactor runObjectE2ETestLite() to be suitable for use as the second
store created in the "run a second object store" test.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The CephSmokeSuite is becoming quite large and long, and most of the
length is now related to the object e2e tests. Separate the full object
e2e test into CephObjectSuite, and only test the object 'lite' test in
the CephSmokeSuite.
Resolves#8714
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
We just need to wait longer when the status is not ready. We needed
another sleep otherwise the status was never nil and the loop went too
fast. See:
```
2021-09-02 16:29:11.372249 I | integrationTest:
2021-09-02 16:29:11.374427 I | integrationTest:
2021-09-02 16:29:11.377764 I | integrationTest:
2021-09-02 16:29:11.379950 I | integrationTest:
2021-09-02 16:29:11.382084 I | integrationTest:
2021-09-02 16:29:11.385383 I | integrationTest:
2021-09-02 16:29:11.388499 I | integrationTest:
2021-09-02 16:29:11.391301 I | integrationTest:
2021-09-02 16:29:11.393545 I | integrationTest:
2021-09-02 16:29:11.396249 I | integrationTest:
```
Signed-off-by: Sébastien Han <seb@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
we are adding new linter for shellcheck.
As we are writing more shell scripts this
will help maintain quality.
Also, doing all the changes required in
bash files to pass this shellcheck.
Closes: https://github.com/rook/rook/issues/8431
Signed-off-by: subhamkrai <srai@redhat.com>
Recently lib-bucket-provisioner add support for update() API.
Include that on the obc implementation since it can be used to
update quota for OBC.
Fixes: #7146
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Add an integration regression test for Ceph Object Storage Object Bucket
Claim (OBC) to verify that the OBC stays in "Bound" state after being
created. The claim should not fall back into "Pending" (or any other)
state soon after being created.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
In case ssl enabled for RGW, bucket healthcheck won't work since access is denied from the server.
In that case configure S3 agent using "InsecureSkipVerify: true" option.
Fixes: 7288
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Work around issue https://github.com/rook/rook/issues/7573
and make sure integration tests check for regressions.
Eventually we should use the RADOS Gateway admin REST API, but for now
we need to work around an issue where the built version of
'radosgw-admin' has incompatibilities with the RADOS Gateway version
running in the Ceph cluster.
Of note, Rook built on the Ceph Pacific image will not support
some 'radosgw-admin' commands to Ceph Nautilus (v14) or Octopus (v15)
clusters.
Further complicating matters, the flag used for the workaround changes
between Ceph v16.2.0 and v16.2.1 (both Pacific).
This bug needs to be treated a little differently than most of the ways
Rook handles different commands for different Ceph versions because this
is based on the Ceph version that is installed in the container with the
Rook operator primarily and not the version of Ceph running in the
cluster.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Sometimes the ci fails to find the object store user, this could be
because the user is created before the object store. So first it fails
since there is no object store and stay in a retry back-off.
Then the object store is created and always need a bit of time, at the
same time the object-user is still waiting and the next retry is longer.
Creating the object store first and then the user hopefully will help
resolve that so we don't have to increase the retry for the object-user
presence.
The error:
```
2021-03-26T08:39:34.7283769Z 2021-03-26 08:39:34.727946 I | testutil: All 2 pod(s) with label rook_object_store=teststore are running
2021-03-26T08:39:34.7284744Z 2021-03-26 08:39:34.727972 I | testutil: Running kubectl [apply -f -]
2021-03-26T08:39:35.1110681Z service/rgw-external-teststore created
2021-03-26T08:39:35.1142392Z 2021-03-26 08:39:35.112479 I | integrationTest: Check that RGW pods are Running
2021-03-26T08:39:35.1181683Z 2021-03-26 08:39:35.117763 I | testutil: 2 of 1 pods with label app=rook-ceph-rgw were found
2021-03-26T08:39:35.1251276Z 2021-03-26 08:39:35.124933 I | testutil: 2 of 1 pods with label app=rook-ceph-rgw were found
2021-03-26T08:39:35.1348603Z 2021-03-26 08:39:35.133276 I | integrationTest: RGW pods are running
2021-03-26T08:39:35.1350090Z 2021-03-26 08:39:35.133290 I | integrationTest: Object store created successfully
2021-03-26T08:39:35.1351373Z 2021-03-26 08:39:35.133295 I | integrationTest: Waiting 5 seconds for the object user to be created
2021-03-26T08:39:40.1340659Z 2021-03-26 08:39:40.133401 I | integrationTest: Checking to see if the user secret has been created
2021-03-26T08:39:40.1342715Z 2021-03-26 08:39:40.133440 D | exec: Running command: kubectl get -n smoke-ns secrets -l rook_object_store=teststore -l user=rook-user
2021-03-26T08:39:40.2618726Z 2021-03-26 08:39:40.261412 I | clients: Unable to find user secret
2021-03-26T08:39:40.2619705Z 2021-03-26 08:39:40.261433 I | integrationTest: (0) secret check sleeping for 5 seconds ...
2021-03-26T08:39:45.2621870Z 2021-03-26 08:39:45.261641 D | exec: Running command: kubectl get -n smoke-ns secrets -l rook_object_store=teststore -l user=rook-user
2021-03-26T08:39:45.3863424Z 2021-03-26 08:39:45.386048 I | clients: Unable to find user secret
2021-03-26T08:39:45.3864586Z 2021-03-26 08:39:45.386075 I | integrationTest: (1) secret check sleeping for 5 seconds ...
2021-03-26T08:39:50.3869496Z 2021-03-26 08:39:50.386292 D | exec: Running command: kubectl get -n smoke-ns secrets -l rook_object_store=teststore -l user=rook-user
2021-03-26T08:39:50.4991166Z 2021-03-26 08:39:50.498862 I | clients: Unable to find user secret
2021-03-26T08:39:50.4991987Z 2021-03-26 08:39:50.498881 I | integrationTest: (2) secret check sleeping for 5 seconds ...
2021-03-26T08:39:55.4998176Z 2021-03-26 08:39:55.499038 D | exec: Running command: kubectl get -n smoke-ns secrets -l rook_object_store=teststore -l user=rook-user
2021-03-26T08:39:55.6075640Z 2021-03-26 08:39:55.607337 I | clients: Unable to find user secret
2021-03-26T08:39:55.6076472Z 2021-03-26 08:39:55.607355 I | integrationTest: (3) secret check sleeping for 5 seconds ...
2021-03-26T08:40:00.6080687Z 2021-03-26 08:40:00.607466 D | exec: Running command: kubectl get -n smoke-ns secrets -l rook_object_store=teststore -l user=rook-user
2021-03-26T08:40:00.8415649Z 2021-03-26 08:40:00.841309 I | clients: Unable to find user secret
2021-03-26T08:40:00.8417179Z 2021-03-26 08:40:00.841331 I | integrationTest: (4) secret check sleeping for 5 seconds ...
2021-03-26T08:40:05.8419407Z 2021-03-26 08:40:05.841385 D | exec: Running command: kubectl get -n smoke-ns secrets -l rook_object_store=teststore -l user=rook-user
2021-03-26T08:40:05.9732221Z 2021-03-26 08:40:05.972910 I | clients: Unable to find user secret
2021-03-26T08:40:05.9733777Z 2021-03-26 08:40:05.972932 I | integrationTest: (5) secret check sleeping for 5 seconds ...
2021-03-26T08:40:10.9734134Z 2021-03-26 08:40:10.973067 D | exec: Running command: kubectl get -n smoke-ns secrets -l rook_object_store=teststore -l user=rook-user
2021-03-26T08:40:11.1686581Z 2021-03-26 08:40:11.168373 I | clients: Unable to find user secret
2021-03-26T08:40:11.1687925Z ceph_base_object_test.go:87:
2021-03-26T08:40:11.1688924Z Error Trace: ceph_base_object_test.go:87
2021-03-26T08:40:11.1689961Z ceph_smoke_test.go:132
2021-03-26T08:40:11.1691113Z Error: Should be true
2021-03-26T08:40:11.1692167Z Test: TestCephSmokeSuite/TestObjectStorage_SmokeTest
2021-03-26T08:40:11.1693751Z 2021-03-26 08:40:11.168502 D | ceph-object-controller: getting s3 user "rook-user"
2021-03-26T08:40:11.1695580Z 2021-03-26 08:40:11.168514 D | exec: Running command: kubectl exec -i rook-ceph-tools-78cdfd976c-2nbp2 -n smoke-ns -- timeout 15 radosgw-admin user info --uid rook-user --rgw-realm= --rgw-zonegroup= --rgw-zone=
2021-03-26T08:40:11.6090406Z ceph_base_object_test.go:89:
2021-03-26T08:40:11.6093015Z Error Trace: ceph_base_object_test.go:89
2021-03-26T08:40:11.6093968Z ceph_smoke_test.go:132
2021-03-26T08:40:11.6095116Z Error: Received unexpected error:
2021-03-26T08:40:11.6095798Z failed to get user info: warn: s3 user not found
2021-03-26T08:40:11.6096724Z github.com/rook/rook/pkg/operator/ceph/object.GetUser
2021-03-26T08:40:11.6097795Z /home/runner/work/rook/rook/pkg/operator/ceph/object/user.go:98
2021-03-26T08:40:11.6098764Z github.com/rook/rook/tests/framework/clients.(*ObjectUserOperation).GetUser
2021-03-26T08:40:11.6099814Z /home/runner/work/rook/rook/tests/framework/clients/object_user.go:45
2021-03-26T08:40:11.6100792Z github.com/rook/rook/tests/integration.runObjectE2ETest
2021-03-26T08:40:11.6101858Z /home/runner/work/rook/rook/tests/integration/ceph_base_object_test.go:88
2021-03-26T08:40:11.6103298Z github.com/rook/rook/tests/integration.(*SmokeSuite).TestObjectStorage_SmokeTest
2021-03-26T08:40:11.6104371Z /home/runner/work/rook/rook/tests/integration/ceph_smoke_test.go:132
2021-03-26T08:40:11.6105104Z reflect.Value.call
2021-03-26T08:40:11.6105856Z /opt/hostedtoolcache/go/1.15.10/x64/src/reflect/value.go:476
2021-03-26T08:40:11.6106662Z reflect.Value.Call
2021-03-26T08:40:11.6107412Z /opt/hostedtoolcache/go/1.15.10/x64/src/reflect/value.go:337
2021-03-26T08:40:11.6108217Z github.com/stretchr/testify/suite.Run.func1
2021-03-26T08:40:11.6109104Z /home/runner/go/pkg/mod/github.com/stretchr/testify@v1.6.1/suite/suite.go:158
2021-03-26T08:40:11.6109844Z testing.tRunner
2021-03-26T08:40:11.6110578Z /opt/hostedtoolcache/go/1.15.10/x64/src/testing/testing.go:1123
2021-03-26T08:40:11.6111271Z runtime.goexit
2021-03-26T08:40:11.6112037Z /opt/hostedtoolcache/go/1.15.10/x64/src/runtime/asm_amd64.s:1374
2021-03-26T08:40:11.6112995Z Test: TestCephSmokeSuite/TestObjectStorage_SmokeTest
```
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When verifying OBC creation, validating that secret and configmap exist
is good, but the definitive validation is to check that the OBC's phase
is "Bound".
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Always attempt to delete the object store in case some test
fails and aborts the remainder of the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The object bucket tests will now always check if the conditions being waited
upon for creating users and buckets were completed, instead of allowing
the test to continue even when they didn't succeed.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The object bucket check is sometimes not completed yet
that causes the integration test to fail intermittently.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Instead of setting the Phase of the CR once the reconcile is done, let's
actually set in from the healthcheck so it is more accurate.
Closes: https://github.com/rook/rook/issues/5249
Signed-off-by: Sébastien Han <seb@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Now, when an S3 user gets created, Rook will add the S3 endpoint to the
Secret along with the credentials.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.
A good status will look like:
status:
endpointStatus:
lastChanged: "2020-06-25T13:47:45Z"
lastChecked: "2020-06-25T13:48:46Z"
phase: Connected
A failed status:
status:
endpointStatus:
details: |-
error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
health: ERROR
This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.
Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
This fix is added to ensure that everything works as expected
when user tries to create ObjectStoreUser before creating
ObjectStore itself. ObjectStoreUser will wait for ObjectStore
to be up and running.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
Purge the object store properly, otherwise the cephobjectstore CRD won't
be deleted since a finalizer is in place for the CR.
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.
A number of other test improvements are also made for code
readability, test stability, and removing unused options.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
retry ObjectUser creation as long as valid input is provided
or it times out. Retry every 15sec for 5min.
integration test creates ObjectUser before ObjectStore to
ensure that ObjectUser waits for ObjectStore to be created
and available
Closes: https://github.com/rook/rook/issues/3937
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
With the storage providers growing, we should make it
clear in the source which tests belong to which storage
provider. This change renames the ceph integration tests.
Signed-off-by: travisn <tnielsen@redhat.com>