RGW can only serve a single certificate. This limitation means that the
prior behavior of using the default service for admin ops when TLS is
enabled may mean it requires additional complex certificate management
to make sure the object store uses a certificate valid for Rook internal
admin ops and user connections.
This is needlessly complex for users. Instead, change Rook's behavior
and documentation to clarify that it will use the same endpoint intended
for S3 client applications. This means that users have a more
straightforward path to enabling both Rook and consuming applications.
More info: https://github.com/rook/rook/issues/14530
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Add CephObjectStore spec.hosting.advertiseEndpoint configuration. This
provides a clear documented default for which endpoint Rook "advertises"
to dependent resources like CephObjectStores, OBCs, and COSI
Buckets/Accesses and allows users to override the default behavior if
desired.
The current default is to round-robin an endpoint from
spec.hosting.dnsNames, which has proven to be troublesome for some
users' object store configurations. This change provides much-needed
disambiguation for users.
This may be a breaking change for some existing spec.hosting.dnsNames
users. This is unexpected but is documented.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The radosgw-admin command uses the network spec from ceph cluster spec
in object context but it is not filled properly in the object package.
But with PR 10898, network spec is available in clusterinfo which can
be used directly. Also removed cluserspec from object context.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.
This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.
The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The `periodWillChange()` implementation to detect map keys to ignore.
Previously, we relied on the `GoString()` method to output the same
format always to determine the current path to the
map[string]interface{} node, but that may change depending on the underlying
implementation of go-cmp.
Instead, parse the go-cmp `Path` step by step and construct a JSON path
to the node in a stable format.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Check the radosgw-admin realm user list per object store instead of relying
on the ceph dashboard get-rgw-api- command.
Resolves#9099
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.
Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.
To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
of its operation that didn't cause failures in the past
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.
Resolves#8879
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.
If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.
Resolves#8778
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
If the cluster where the rgw is started is secondary and not primary,
trying to create the admin ops user will fail with:
```
Please run the command on master zone.
Performing this operation on non-master zone
leads to inconsistent metadata between zones
```
So we need to force the creation regardless, it is fine the creation will
return UserAlreadyExist and then we just read the current user.
Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The go-ceph library has removed the `Debug` field from the API type in
https://github.com/ceph/go-ceph/pull/543. Since the HTTP Client can be
mutated we now have our own client to dump requests and responses when
the operator log level is DEBUG.
Signed-off-by: Sébastien Han <seb@redhat.com>
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
We now pass the networking spec to the clusterInfo so that the executor
can make the right decision on how to execute a command.
This is a small change that allows us to remove the CephBlockPool CR
since it's checking for rbd images. On a Multus deployment, the operator
does not have the network annotations, thus has no access to the OSD
network then rbd commands are hanging forever.
Signed-off-by: Sébastien Han <seb@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
The debug for adminOps client can be enabled by setting `Debug` flag.
In this PR, it is enabled for OBC and cephobjectstore not healthchecker.
Also add similar changes while calling `NewS3Agent` in the bucket
provisioner code path.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.
Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
radosgw-admin may output logs before and after JSON when this command succeeds. We should
get rid of these logs before unmarshaling output. Although the logs before JSON were
treated in PR7354, the logs after JSON weren't.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Work around issue https://github.com/rook/rook/issues/7573
and make sure integration tests check for regressions.
Eventually we should use the RADOS Gateway admin REST API, but for now
we need to work around an issue where the built version of
'radosgw-admin' has incompatibilities with the RADOS Gateway version
running in the Ceph cluster.
Of note, Rook built on the Ceph Pacific image will not support
some 'radosgw-admin' commands to Ceph Nautilus (v14) or Octopus (v15)
clusters.
Further complicating matters, the flag used for the workaround changes
between Ceph v16.2.0 and v16.2.1 (both Pacific).
This bug needs to be treated a little differently than most of the ways
Rook handles different commands for different Ceph versions because this
is based on the Ceph version that is installed in the container with the
Rook operator primarily and not the version of Ceph running in the
cluster.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Sometimes `radosgw-admin` succeeds after showing logs to stderr. We should
skip non-json strings if parsing output as json.
Here is an example.
```
2021-02-26 04:10:44.190418 I | op-bucket-prov: creating Ceph user "ceph-user-aSzNqgE7"
E0226 04:11:37.901310 8 controller.go:199] error syncing 'logging/loki-bucket': error provisioning bucket: Provision: can't create ceph user: error creating ceph user "ceph-user-aSzNqgE7": failed to unmarshal json. 2021-02-26T04:11:21.425+0000 7f6714be4980 1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:21.426+0000 7f6714be4980 0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.uid:ceph-user-aSzNqgE7
2021-02-26T04:11:32.168+0000 7f6714be4980 1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:32.168+0000 7f6714be4980 0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.keys:23Z8GUEXR0TJDO86PSJR
{
"user_id": "ceph-user-aSzNqgE7",
"display_name": "ceph-user-aSzNqgE7",
...
"mfa_ids": []
}: invalid character '-' after top-level value: failed to unmarshal json.
```
In this case, some logs like "robust_notify:..." was shown in stderr.
Unmarsharing was failed due to tried to parse these logs as json.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
When creating object store, `radosgw-admin realm get ..` command is stuck forever when required number of OSDs are not available. Because of this the uninstall of object store is also stuck. User has to manually remove the finalizer to delete the object store.This PR uses `ExecuteCommandWithTimeout` for running `radosgw-admin` command. Timeout during installation will be reconciled. Cleanup will be treated as best effort. Any errors during uninstalling of single site object store will only be logged.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.
- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up
- Added more debug logging all around the object-store
code related to multisite
- files generated by rerun of `make codegen`
- change back the edit on the Copyright in object.go
Signed-off-by: Ali Maredia <amaredia@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For the CephObjectStore, the Status field of the CR has a new property
called "info", which contains a couple of information about the CR.
The first new addition is the endpoint of the object store.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.
A good status will look like:
status:
endpointStatus:
lastChanged: "2020-06-25T13:47:45Z"
lastChecked: "2020-06-25T13:48:46Z"
phase: Connected
A failed status:
status:
endpointStatus:
details: |-
error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
health: ERROR
This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.
Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone
Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.
Signed-off-by: Ali Maredia <amaredia@redhat.com>
When a user or a bucket does not exist anymore, let's stop the reconcile
loop instead of waiting forever.
Also, unlink the user before deleting it otherwise the deletion fails.
Signed-off-by: Sébastien Han <seb@redhat.com>
Currently, we need to configure the Ceph external Admin keyi
in the Rook deployment to be able to connect to an external Ceph cluster.
If we wanted to run in a multi-tenant fashion
were several K8s/Rook clusters wanted to connect to the same external ceph cluster,
each K8s deployment would have the access to the External Ceph Admin key
and could potentially access or delete the Data
from pools that belong to other k8s/Rook Clusters.
Now the admin key is optional but the helper script create-external-cluster-resources.sh
will help create the necessary keys/users to connect to that cluster.
Closes: https://github.com/rook/rook/issues/4917 and https://github.com/rook/rook/pull/5227
Signed-off-by: Sébastien Han <seb@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
Configure the Ceph rgw daemon completely from the operator a la the
recent changes to the Ceph mon, mgr, and mds operators.
Create the rgw deployment or daemonset first, and then create the
keyring secret for the object store with its owner reference as the
corresponding deployment or daemonset. When the replication controller
is deleted, the secret is also deleted.
The RGW's mime.types file is now stored in a configmap with a different
file created for each object store. This is primarily just a means to
get the mime.types file into the rgw pod, but the added benefit is that
the administrator can modify the configmap, which could reduce
susceptibility to file type execution vulnerabilities (worst case).
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>