use Zone and ZoneGroup instead of storename for rgw_zone and rgw_zonegroup
Signed-off-by: Olivier Bouffet <olivier.bouffet@infomaniak.com>
(cherry picked from commit c92270cd66)
Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.
Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
The generation of the rgw deployment spec was swallowing errors
if any issues are raised such as the tls cert not being found
as expected in some configurations. We need to fail the reconcile
so the error will be logged and the admin can identify the issue.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
If the admin wants to use insecure TLS to validate connections to rgw
internally, the TLS secret can have another entry "insecureSkipVerify"
and set it to "true".
Signed-off-by: Sébastien Han <seb@redhat.com>
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
For external RGW server use the IP mentioned in Gateway for admin Ops
operattions.
Fixes: #8916
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
replaces all occurences of lduo/rduo quotation marks to make
sure that using the snippets in a k8s manifest will work
This fixes an issue with ArgoCD not being able
to apply `common.yaml` because of an encoding issue
Signed-off-by: PixelJonas <jonas@janz.digital>
The CephObjectRealm controller would fail all subsequent reconciles if
the first reconcile created the Kubernetes Secret containing the access
keys for the realm but where the radosgw-admin command failed to create
the realm. This was the only idempotency issue found after reviewing the
CephObjectRealm controller.
Resolves#8954
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.
Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.
To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
of its operation that didn't cause failures in the past
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.
Resolves#8879
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The Client_ID generated by operator was
different from the log rotate file created
The Clinet_ID= rgwceph.client.rook.ceph.rgw.my.store.a
and log file name= ceph-client.rgw.my.store.a.log
So changed the CLient_ID to ceph-client.rgw.my.store.a for
correct working and this follow the patterns how other modules
Client_ID is generated
Closes: https://github.com/rook/rook/issues/8692
Signed-off-by: parth-gr <paarora@redhat.com>
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.
If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.
Resolves#8778
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Rook should update the RGW object store's period if the period doesn't
yet exist. This protects us from the case where the
'radosgw-admin period update --commit` command fails and the
CephObjectStore controller reconciles again.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
We have seen cases where the signed certificate used for the RGW does not
contain the internal DNS endpoint, resulting in the health check to fail
since the certificate is not valid for this domain.
People consuming the gateways by external clients and for specific
domains do not necessarily have the internal DNS configured in the
certificate.
So let's be a bit more flexible and simply ensure a connectivity check
and bypass the certificate validation.
Also, this is fixing the tls code in newS3Agent and adds unit tests.
Closes: #8663
Signed-off-by: Sébastien Han <seb@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
There was a log line that informed that the object store status would
not be updated because the status was deleting erroneously. Move the
line to the correct position.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
If the cluster where the rgw is started is secondary and not primary,
trying to create the admin ops user will fail with:
```
Please run the command on master zone.
Performing this operation on non-master zone
leads to inconsistent metadata between zones
```
So we need to force the creation regardless, it is fine the creation will
return UserAlreadyExist and then we just read the current user.
Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
The MDS core team suggested with deploy the MDS daemon first and then do
the filesystem creation and configuration. Reversing the sequence lets
us avoid spurious FS_DOWN warnings when creating the filesystem.
Closes: #8745
Signed-off-by: Sébastien Han <seb@redhat.com>
If the region is specified in the storage class of OBC, use that in the
newS3agent() than using constant "us-east-1".
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
We just need to wait longer when the status is not ready. We needed
another sleep otherwise the status was never nil and the loop went too
fast. See:
```
2021-09-02 16:29:11.372249 I | integrationTest:
2021-09-02 16:29:11.374427 I | integrationTest:
2021-09-02 16:29:11.377764 I | integrationTest:
2021-09-02 16:29:11.379950 I | integrationTest:
2021-09-02 16:29:11.382084 I | integrationTest:
2021-09-02 16:29:11.385383 I | integrationTest:
2021-09-02 16:29:11.388499 I | integrationTest:
2021-09-02 16:29:11.391301 I | integrationTest:
2021-09-02 16:29:11.393545 I | integrationTest:
2021-09-02 16:29:11.396249 I | integrationTest:
```
Signed-off-by: Sébastien Han <seb@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters
This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Passing struct by value essentially gives you a copy, so when modified
within a function, the scope is then reduced to that function. Using
pointers solves that you mutate the struct as many times as you want from
anywhere.
As a result, the auto-detection of the Vault KV backend was not working
correctly.
Also, added a ton of unit tests for Vault.
Signed-off-by: Sébastien Han <seb@redhat.com>