Commit Graph
35 Commits
Author SHA1 Message Date
Blaine Gardner 956430826c rgw: add integration test for committing period
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.

Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.

To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
   of its operation that didn't cause failures in the past

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 15:24:59 -06:00
Blaine Gardner eadcd757b3 rgw: replace period update --commit with function
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.

Resolves #8879

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 14:45:39 -06:00
Blaine Gardner 7586cea049 core: create TRACE_INSECURE log level
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.

If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.

Resolves #8778

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-01 12:09:19 -06:00
Sébastien Han 8786b40d64 rgw: do not create the rgw ops user on the secondary cluster
If the cluster where the rgw is started is secondary and not primary,
trying to create the admin ops user will fail with:

```
Please run the command on master zone.
Performing this operation on non-master zone
leads to inconsistent metadata between zones
```

So we need to force the creation regardless, it is fine the creation will
return UserAlreadyExist and then we just read the current user.

Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 16:58:03 +02:00
Blaine Gardner 5383ba2df2 ceph: retry object health check if creation fails
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:24:18 -06:00
Sébastien Han ad45924807 ceph: change the debug implementation of the admin ops API
The go-ceph library has removed the `Debug` field from the API type in
https://github.com/ceph/go-ceph/pull/543. Since the HTTP Client can be
mutated we now have our own client to dump requests and responses when
the operator log level is DEBUG.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-29 13:59:50 +02:00
Satoru Takeuchi 30e4fbb01f ceph: make the timeout of ceph commands cofigurable
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-27 12:42:16 +00:00
Sébastien Han 1900160939 ceph: proxy rbd commands when multus is enabled
We now pass the networking spec to the clusterInfo so that the executor
can make the right decision on how to execute a command.
This is a small change that allows us to remove the CephBlockPool CR
since it's checking for rbd images. On a Multus deployment, the operator
does not have the network annotations, thus has no access to the OSD
network then rbd commands are hanging forever.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-20 12:28:00 +02:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Jiffin Tony Thottan 560f44b07e ceph: enable debug for adminops client if the rook loglevel <= debug
The debug for adminOps client can be enabled by setting `Debug` flag.
In this PR, it is enabled for OBC and cephobjectstore not healthchecker.
Also add similar changes while calling `NewS3Agent` in the bucket
provisioner code path.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-07-01 11:36:51 +05:30
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Satoru Takeuchi ca9aa1a11c ceph: parse json correctly
radosgw-admin may output logs before and after JSON when this command succeeds. We should
get rid of these logs before unmarshaling output. Although the logs before JSON were
treated in  PR7354, the logs after JSON weren't.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-05-27 04:46:13 +00:00
Satoru Takeuchi 709e8d5ba8 ceph: unity timeout values of commands
There is no strong reason to have different timeout values.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-04-26 12:57:40 +00:00
Blaine Gardner 8aef8b756d ceph: change radosgw-admin workaround log to debug
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-16 12:51:54 -06:00
Blaine Gardner d78b2a41d6 ceph: work around radosgw-admin fifo file io error
Work around issue https://github.com/rook/rook/issues/7573
and make sure integration tests check for regressions.

Eventually we should use the RADOS Gateway admin REST API, but for now
we need to work around an issue where the built version of
'radosgw-admin' has incompatibilities with the RADOS Gateway version
running in the Ceph cluster.

Of note, Rook built on the Ceph Pacific image will not support
some 'radosgw-admin' commands to Ceph Nautilus (v14) or Octopus (v15)
clusters.

Further complicating matters, the flag used for the workaround changes
between Ceph v16.2.0 and v16.2.1 (both Pacific).

This bug needs to be treated a little differently than most of the ways
Rook handles different commands for different Ceph versions because this
is based on the Ceph version that is installed in the container with the
Rook operator primarily and not the version of Ceph running in the
cluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-13 12:08:02 -06:00
Satoru Takeuchi a89d35e5b6 ceph: fix improper json parsing in radosgw-admin
Sometimes `radosgw-admin` succeeds after showing logs to stderr. We should
skip non-json strings if parsing output as json.

Here is an example.

```
2021-02-26 04:10:44.190418 I | op-bucket-prov: creating Ceph user "ceph-user-aSzNqgE7"
E0226 04:11:37.901310       8 controller.go:199] error syncing 'logging/loki-bucket': error provisioning bucket: Provision: can't create ceph user: error creating ceph user "ceph-user-aSzNqgE7": failed to unmarshal json. 2021-02-26T04:11:21.425+0000 7f6714be4980  1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:21.426+0000 7f6714be4980  0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.uid:ceph-user-aSzNqgE7
2021-02-26T04:11:32.168+0000 7f6714be4980  1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:32.168+0000 7f6714be4980  0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.keys:23Z8GUEXR0TJDO86PSJR
{
    "user_id": "ceph-user-aSzNqgE7",
    "display_name": "ceph-user-aSzNqgE7",
...
    "mfa_ids": []
}: invalid character '-' after top-level value: failed to unmarshal json.
```

In this case, some logs like "robust_notify:..." was shown in stderr.
Unmarsharing was failed due to tried to parse these logs as json.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-10 16:06:56 +00:00
Santosh Pillai 113376f1ab ceph: timeout radosgw-admin cli commands
When creating object store, `radosgw-admin realm get ..` command is stuck forever when required number of OSDs are not available. Because of this the uninstall of object store is also stuck. User has to manually remove the finalizer to delete the object store.This PR uses `ExecuteCommandWithTimeout` for running `radosgw-admin` command. Timeout during installation will be reconciled. Cleanup will be treated as best effort. Any errors during uninstalling of single site object store will only be logged.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-02-01 10:57:04 +05:30
Ali Maredia e6ed4ff8ea ceph: minor fixes + add realm/zg/zone to object context
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.

- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up

- Added more debug logging all around the object-store
code related to multisite

- files generated by rerun of `make codegen`

- change back the edit on the Copyright in object.go

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-07-21 16:47:59 -04:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Sébastien Han 5c5008e27c ceph: add s3 endpoint to CephObjectStore CRD
For the CephObjectStore, the Status field of the CR has a new property
called "info", which contains a couple of information about the CR.
The first new addition is the endpoint of the object store.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-10 12:06:25 +02:00
Sébastien Han e4eaa91ede ceph: add rgw endpoint healthcheck
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.

A good status will look like:

status:
  endpointStatus:
    lastChanged: "2020-06-25T13:47:45Z"
    lastChecked: "2020-06-25T13:48:46Z"
  phase: Connected

A failed status:

status:
  endpointStatus:
    details: |-
      error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
      caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
    health: ERROR

This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.

Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-02 16:34:10 +02:00
Sébastien Han 84d1e28c99 ceph: add external support for objectstoreuser
Now, the object store user is capable of creating s3 users on an
external Ceph cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 16:34:01 +02:00
Ali Maredia fc579f4520 ceph: initial commit for ceph rgw multisite resources
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone

Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-06-09 16:28:53 -04:00
Sébastien Han 58577a3537 ceph: obc relax error handling and user deletion
When a user or a bucket does not exist anymore, let's stop the reconcile
loop instead of waiting forever.
Also, unlink the user before deleting it otherwise the deletion fails.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-05-12 16:28:47 +02:00
Sébastien Han b86bb9c528 ceph: add support for OBC on external cluster
We can now consume an external s3 endpoint which is not managed by Rook.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-05-04 18:14:26 +02:00
Madhu Rajanna 81688398f2 cleanup: use err.Wrap when the formatting is not required
Replaced err.Wrapf with err.Wrap when the formatting
is not required.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-29 17:42:59 +05:30
Sébastien Han 7c994441c3 ceph: make admin key optional for external cluster
Currently, we need to configure the Ceph external Admin keyi
in the Rook deployment to be able to connect to an external Ceph cluster.
If we wanted to run in a multi-tenant fashion
were several K8s/Rook clusters wanted to connect to the same external ceph cluster,
each K8s deployment would have the access to the External Ceph Admin key
and could potentially access or delete the Data
from pools that belong to other k8s/Rook Clusters.

Now the admin key is optional but the helper script create-external-cluster-resources.sh
will help create the necessary keys/users to connect to that cluster.

Closes: https://github.com/rook/rook/issues/4917 and https://github.com/rook/rook/pull/5227
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-20 09:59:04 +02:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 18b0e7d295 ceph: scrub ceph commands to write actions to the log
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Sébastien Han 9f2867e12a ceph: separate controller for CephObjectStoreUser CRD
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-06 11:53:40 +01:00
Sébastien Han 5ce2ed220e ceph: use "github.com/pkg/errors"
We now use the error package.
Kubernetes errors have been renamed kerrors since they are lower than
'errors'.

Closes: https://github.com/rook/rook/issues/4054
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-09 16:58:32 +01:00
c852e86fe9 ceph-rook object bucket provisioner
Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: jeffvance <jeff.h.vance@gmail.com>
Signed-off-by: Jon Cope <jcope@redhat.com>

Co-authored-by: Jon Cope <copejon@users.noreply.github.com>
Co-authored-by: Jeff Vance <jeff.h.vance@gmail.com>
2019-08-22 19:54:43 -06:00
Blaine Gardner 086fa8231c rgw: configure entirely in operator
Configure the Ceph rgw daemon completely from the operator a la the
recent changes to the Ceph mon, mgr, and mds operators.

Create the rgw deployment or daemonset first, and then create the
keyring secret for the object store with its owner reference as the
corresponding deployment or daemonset. When the replication controller
is deleted, the secret is also deleted.

The RGW's mime.types file is now stored in a configmap with a different
file created for each object store. This is primarily just a means to
get the mime.types file into the rgw pod, but the added benefit is that
the administrator can modify the configmap, which could reduce
susceptibility to file type execution vulnerabilities (worst case).

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-03-07 08:02:42 -07:00