Commit Graph
426 Commits
Author SHA1 Message Date
Olivier Bouffet 5edeff4df8 object: fix search user in objectstore
avoid failed reconcile when multiple multisite objectstore are configured

Signed-off-by: Olivier Bouffet <olivier.bouffet@infomaniak.com>
2021-11-29 20:38:22 +01:00
Olivier 561cede1ef object: fix rgw ceph config
use Zone and ZoneGroup instead of storename for rgw_zone and rgw_zonegroup

Signed-off-by: Olivier Bouffet <olivier.bouffet@infomaniak.com>
(cherry picked from commit c92270cd66)
2021-11-25 18:38:42 +01:00
Sébastien Han e9f9a40138 Merge pull request #9212 from travisn/admin-test-cluster
core: Ensure cluster name is available on cluster info
2021-11-22 11:50:55 +01:00
Blaine Gardner 03ba7dec64 pool: file: object: clean up stop health checkers
Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.

Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-19 10:29:12 -07:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Jiffin Tony Thottan aba50d3ca9 object: add support in RGW to communicate vault with TLS
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-17 10:19:28 +05:30
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Yuichiro Ueno 0b575703c7 core: add context parameter to k8sutil deployment
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 14:58:13 +09:00
Travis Nielsen 9ecd0cbc75 rgw: raise errors when rgw daemon fails to be created
The generation of the rgw deployment spec was swallowing errors
if any issues are raised such as the tls cert not being found
as expected in some configurations. We need to fail the reconcile
so the error will be logged and the admin can identify the issue.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-10 09:13:03 -07:00
Yuval Lifshitz 71ed45b69b rgw: implement bucket notifications for object storage
following the design from here:
https://github.com/rook/rook/blob/master/design/ceph/object/ceph-bucket-notification-crd.md

Closes: https://github.com/rook/rook/issues/5313
Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-11-04 11:20:40 +02:00
Yuzuki Mimura 536b59ef0f rgw: change the way to livenessProbe and introduce readinessProbe
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.

Closes: #8407

Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-10-29 15:07:29 +00:00
Sébastien Han 7fbc298498 Merge pull request #9049 from leseb/fix-8683
rgw: add support for updating user caps
2021-10-28 15:07:18 +02:00
Sébastien Han ecd01bb452 Merge pull request #9020 from leseb/fix-8993
rgw: read tls secret hint for insecure tls
2021-10-28 15:07:07 +02:00
Sébastien Han f2cb792e9f rgw: add support for updating user caps
User's capabilities can now be updated from the admin ops API.

Closes: https://github.com/rook/rook/issues/8683
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-28 14:17:13 +02:00
Sébastien Han 86c9a8f3de rgw: read tls secret hint for insecure tls
If the admin wants to use insecure TLS to validate connections to rgw
internally, the TLS secret can have another entry "insecureSkipVerify"
and set it to "true".

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-28 11:27:00 +02:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Sébastien Han c54a5556f3 rgw: stop using context.TODO() and use parent ctx
The clusterInfo has the parent Context so let's use it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-27 11:52:52 +02:00
Blaine Gardner 2f850b6ae6 Merge pull request #8613 from subhamkrai/remove-nautilus
ceph: remove ceph nautilus, ceph octopus to default
2021-10-25 09:09:18 -06:00
Sébastien Han fc9c4b6fc9 Merge pull request #9010 from thotz/rgw-external-endpoint
ceph: update endpoint with IP for external RGW server
2021-10-25 15:49:08 +02:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
Jiffin Tony Thottan d4562f6b83 ceph: update endpoint with IP for external RGW server
For external RGW server use the IP mentioned in Gateway for admin Ops
operattions.

Fixes: #8916
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-10-21 19:13:10 +05:30
Blaine Gardner 063c714c97 Merge pull request #9007 from BlaineEXE/remove-dynamic-clientset
ceph: get rid of dynamic clientset
2021-10-20 09:40:28 -06:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
PixelJonas 8733bf6272 ceph: use normal quotation marks for comment
replaces all occurences of lduo/rduo quotation marks to make
sure that using the snippets in a k8s manifest will work

This fixes an issue with ArgoCD not being able
to apply `common.yaml` because of an encoding issue

Signed-off-by: PixelJonas <jonas@janz.digital>
2021-10-20 10:09:22 +02:00
Blaine Gardner 01e2feaef5 ceph: get rid of dynamic clientset
Stop using the dynamic clientset in favor of the controller-runtime
clientset.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-19 17:05:38 -06:00
Blaine Gardner 34a8b42e97 test: try to un-flake multi-cluster-mirror test
Try to un-flake the multi-cluster-mirror test that keeps failing on this
PR.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-13 11:06:51 -06:00
Blaine Gardner 65c4972038 rgw: make CephObjectRealm controller idempotent
The CephObjectRealm controller would fail all subsequent reconciles if
the first reconcile created the Kubernetes Secret containing the access
keys for the realm but where the radosgw-admin command failed to create
the realm. This was the only idempotency issue found after reviewing the
CephObjectRealm controller.

Resolves #8954

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-13 11:06:51 -06:00
Blaine Gardner a1dd256d4b Merge pull request #8911 from BlaineEXE/rgw-commands-use-staging-flag
rgw: replace period update --commit with function
2021-10-12 08:22:41 -06:00
Blaine Gardner 956430826c rgw: add integration test for committing period
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.

Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.

To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
   of its operation that didn't cause failures in the past

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 15:24:59 -06:00
Blaine Gardner eadcd757b3 rgw: replace period update --commit with function
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.

Resolves #8879

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 14:45:39 -06:00
parth-gr fc7905a7bd ceph: fixing ClientID of log-collector for RGW instance
The Client_ID generated by operator was
different from the log rotate file created
The Clinet_ID= rgwceph.client.rook.ceph.rgw.my.store.a
and log file name= ceph-client.rgw.my.store.a.log
So changed the CLient_ID to ceph-client.rgw.my.store.a for
correct working and this follow the patterns how other modules
Client_ID is generated

Closes: https://github.com/rook/rook/issues/8692
Signed-off-by: parth-gr <paarora@redhat.com>
2021-10-07 19:30:21 +05:30
Blaine Gardner 7586cea049 core: create TRACE_INSECURE log level
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.

If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.

Resolves #8778

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-01 12:09:19 -06:00
Blaine Gardner e10fb751e5 rgw: add period does not exist debug message
Log as a debug message when the RGW period will be updated because it
does not exist.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-29 13:14:52 -06:00
Blaine Gardner 7b9293624a rgw: update period if period does not exist
Rook should update the RGW object store's period if the period doesn't
yet exist. This protects us from the case where the
'radosgw-admin period update --commit` command fails and the
CephObjectStore controller reconciles again.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-28 14:17:04 -06:00
Sébastien Han cda5dad291 rgw: use insecure TLS for bucket health check
We have seen cases where the signed certificate used for the RGW does not
contain the internal DNS endpoint, resulting in the health check to fail
since the certificate is not valid for this domain.
People consuming the gateways by external clients and for specific
domains do not necessarily have the internal DNS configured in the
certificate.
So let's be a bit more flexible and simply ensure a connectivity check
and bypass the certificate validation.

Also, this is fixing the tls code in newS3Agent and adds unit tests.

Closes: #8663
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-28 14:48:07 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Travis Nielsen 09a5b08fde Merge pull request #8766 from thotz/regionfixobcprovisioning
ceph: pass region to newS3agent()
2021-09-23 11:24:55 -06:00
Blaine Gardner c8b26e458c rgw: fix misleading log line in rgw health checker
There was a log line that informed that the object store status would
not be updated because the status was deleting erroneously. Move the
line to the correct position.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-22 07:33:57 +00:00
Sébastien Han 8786b40d64 rgw: do not create the rgw ops user on the secondary cluster
If the cluster where the rgw is started is secondary and not primary,
trying to create the admin ops user will fail with:

```
Please run the command on master zone.
Performing this operation on non-master zone
leads to inconsistent metadata between zones
```

So we need to force the creation regardless, it is fine the creation will
return UserAlreadyExist and then we just read the current user.

Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 16:58:03 +02:00
Sébastien Han 470fbfd341 Merge pull request #8743 from leseb/next-pacific
ceph: use next ceph v16.2.6 pacific version
2021-09-21 16:41:45 +02:00
Sébastien Han c1a88f34d4 mds: change init sequence
The MDS core team suggested with deploy the MDS daemon first and then do
the filesystem creation and configuration. Reversing the sequence lets
us avoid spurious FS_DOWN warnings when creating the filesystem.

Closes: #8745
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 15:43:34 +02:00
Jiffin Tony Thottan 280c29f330 ceph: pass region to newS3agent()
If the region is specified in the storage class of OBC, use that in the
newS3agent() than using constant "us-east-1".

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-21 13:08:18 +05:30
Blaine Gardner 5383ba2df2 ceph: retry object health check if creation fails
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:24:18 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Jiffin Tony Thottan 50ecff8f13 ceph: addressing nits from #8211
Addressing remaining nits from the PR #8211

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-09 12:57:32 +05:30
Jiffin Tony Thottan ca43800119 ceph: add options for cephobjectstore user
Adding options for quota, bucket limit, caps for the
`cephobjectstoreuser`.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-07 22:43:09 +05:30
Sébastien Han 8f42bee563 ci: fix object store test
We just need to wait longer when the status is not ready. We needed
another sleep otherwise the status was never nil and the loop went too
fast. See:

```
2021-09-02 16:29:11.372249 I | integrationTest:
2021-09-02 16:29:11.374427 I | integrationTest:
2021-09-02 16:29:11.377764 I | integrationTest:
2021-09-02 16:29:11.379950 I | integrationTest:
2021-09-02 16:29:11.382084 I | integrationTest:
2021-09-02 16:29:11.385383 I | integrationTest:
2021-09-02 16:29:11.388499 I | integrationTest:
2021-09-02 16:29:11.391301 I | integrationTest:
2021-09-02 16:29:11.393545 I | integrationTest:
2021-09-02 16:29:11.396249 I | integrationTest:
```

Signed-off-by: Sébastien Han <seb@redhat.com>

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-07 17:15:20 +02:00
Sébastien Han d4dd0577f9 Merge pull request #8529 from sp98/id-mapping
ceph: add ClusterID and PoolID mappings between local and peer cluster
2021-08-31 17:32:14 +02:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
Sébastien Han d675969567 ceph: fix vault kv secret engine auto-detection
Passing struct by value essentially gives you a copy, so when modified
within a function, the scope is then reduced to that function. Using
pointers solves that you mutate the struct as many times as you want from
anywhere.
As a result, the auto-detection of the Vault KV backend was not working
correctly.
Also, added a ton of unit tests for Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-31 15:29:04 +02:00