Commit Graph
63 Commits
Author SHA1 Message Date
Sébastien Han f2cb792e9f rgw: add support for updating user caps
User's capabilities can now be updated from the admin ops API.

Closes: https://github.com/rook/rook/issues/8683
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-28 14:17:13 +02:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Jiffin Tony Thottan 50ecff8f13 ceph: addressing nits from #8211
Addressing remaining nits from the PR #8211

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-09 12:57:32 +05:30
Jiffin Tony Thottan ca43800119 ceph: add options for cephobjectstore user
Adding options for quota, bucket limit, caps for the
`cephobjectstoreuser`.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-07 22:43:09 +05:30
Sébastien Han 2d55e69416 ceph: move scheme initialization to the same place
Let's initialize the schemes in a single place instead of doing it
when each controller initializes.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-06 11:33:10 +02:00
Sébastien Han 733c0f72f6 ceph: mock object store user create
Since the mock client was added to the rgw/admin package from the
go-ceph project in https://github.com/ceph/go-ceph/pull/532 we can now
mock the HTTP client and validate the entire reconcile.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 10:01:53 +02:00
Sébastien Han ad54ed8aac ceph: add missing spec to the object context
The CephClusterSpec was missing from the object context, so the check
for the network provider in RunAdminCommandNoMultisite() was not
discovering the network property correctly.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-12 15:46:09 +02:00
Sébastien Han f074c12c4d ceph: get the s3 user first instead of create
The Rados Gateway Admin OPS API has changed its behavior from Nautilus
to Pacific. Calling user create on an existing user wil generate
additional keys to the user on Nautilus. Where in Pacific it will report
an error with UserAlreadyExists.
So to handle both scenarios, let's first get the user, and if the user
does not exist (NoSuchUser) we then create it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-09 18:32:48 +02:00
Sébastien Han 1e45eaa436 ceph: always rehydrate the access and secret keys
Prior to that the access and secret keys were left empty if the user
already existed, which led to updating the secret with empty values when
the operator restarts.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-09 10:39:52 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Jiffin Tony Thottan 867474e405 ceph: initialise httpclient for bucketchecker and objectstoreuser
For the TLS communication for AdminOps Api, httpclient is required and
filled with TLS certs, currently it is set to nil pointer in
`buckethealthchecker` and `cephobjectstoreuser`.

Thanks @Krast76 finding the issue even proposing the fix.

Fixes: #8132
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-06-21 10:59:58 +05:30
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Ali Maredia c9fee1f883 ceph: ensure object store endpoint is initialized for user
when store.Status was not initialized it caused
crashes when assigning store.Status.Info["endpoint"]
to objContext.Endpoint.

This commits makes sure store.Status is not Nil
before setting objContext.Endpoint and
objContext.SecureEndpoint in the ObjectStoreUser
controller.

Closes: https://github.com/rook/rook/issues/6916
Signed-off-by: Ali Maredia <amaredia@redhat.com>
2021-04-28 21:47:34 -04:00
Lars Lehtonen f1b341ba23 ceph: fix multiple imports
This fixes libraries that were being imported multiple times in
pkg/operator/ceph/object and its subpackages.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-12 02:57:13 -07:00
Travis Nielsen f50e49be3a ceph: object store user initialization check for nil
If the object store status is not yet initialized, the object store user should
fail its initialization and requeue the reconcile.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 08:39:26 -06:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Jiffin Tony Thottan b836aab13e ceph: fill up endpoint for cephobjectstore user
Currently endpoint is filled from `Status.Info["endpoint"]` string map for cephobjectstore user's secret and
have two issues:
1.) endpoint is nil and only secureendpoint is available
2.) Chance of small race window in which endpoint is not filled before this ceph object user is handling.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-02-12 10:46:30 +05:30
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Sébastien Han 97be23e374 ceph: apply finalizer before updating object status
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-18 18:16:59 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
Ali Maredia e6ed4ff8ea ceph: minor fixes + add realm/zg/zone to object context
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.

- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up

- Added more debug logging all around the object-store
code related to multisite

- files generated by rerun of `make codegen`

- change back the edit on the Copyright in object.go

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-07-21 16:47:59 -04:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen 4681f9e73d ceph: consolidate ceph config and client packages
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:43 -06:00
Travis Nielsen 631b13b906 ceph: refactor creds used by operator
The operator should only connect to ceph with a single set of creds.
In a converged cluster this will be the admin creds and in an external
cluster it will be lower-privileged creds. Independent clusters were
implemented with a separate set of creds. To simplify the code these
are now merged to a single set.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:42 -06:00
Jiffin Tony Thottan e37a6ee122 ceph: display name updation for existing ceph object stores users
Closes: #5713
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-07-16 16:14:34 +05:30
subhamkrai 3e33fc55a3 ceph: report secret in cephobjectstoreuser status field
adding method to generate secret name in
cephobjectstoreuser status field in order
to programmatically retrieve the secret name
generated by the cephobjectstoreuser crd.
Also added unit test for the method.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-15 14:27:23 +05:30
Sébastien Han 07c7dd457f ceph: expose endpoint in the CephObjectStoreUser
Now, when an S3 user gets created, Rook will add the S3 endpoint to the
Secret along with the credentials.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-10 15:59:47 +02:00
Sébastien Han 5f74e493ef ceph: small user delete refactor
Do not return error code, interpret it directly, put the error as part
of the output on failures.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 18:49:03 +02:00
Sébastien Han 84d1e28c99 ceph: add external support for objectstoreuser
Now, the object store user is capable of creating s3 users on an
external Ceph cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 16:34:01 +02:00
Travis Nielsen f47bb945c2 ceph: remove duplicate controller wait setting
The controller setting for requeuing an event moved to the
opcontroller package and was no longer needed in the main
controller package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Travis Nielsen 269cfe1199 ceph: update status on current version of resources
When updating the status on resources sometimes the update fails due to
the resource being an outdated version. This frequently occurs when
the same reconcile loop updates the status multiple times, or the finalizer
is added, or some other update to the resource. Upon the next reconcile
the error would generally go away since the status or finalizer didn't
need to be updated multiple times. But now the error is not expected
since the controller will retrieve the latest version of the resource
immediately before attempting to update it.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-09 11:20:45 -06:00
Travis Nielsen 0e55aa348b tests: reduce unhelpful debug logging
Removed some extremely verbose debug logging that makes
the logs difficult to read when debug is enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-30 13:32:47 -06:00
Sébastien Han 8db14885b5 ceph: controller fix misleading debug log
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 19:12:17 +01:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Sébastien Han 711ec6095b ceph: controllers just log status error
Let's return the original error instead and just log the status change
error.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:15 -06:00
Sébastien Han b3a61bf4b9 ceph: more precise watcher object user controller
Only watch and react on resources that are matching the Kind of
CephObjectStoreUser.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:15 -06:00
Sébastien Han 9f2867e12a ceph: separate controller for CephObjectStoreUser CRD
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-06 11:53:40 +01:00
Nizamudeen 53883f68cf ceph: Handling Unhandled errors
This commit is to handle all those unhandled errors which raises the gosec warning.

Fixed G104: Unhandled Errors are handled now

Signed-off-by: Nizamudeen <nia@redhat.com>
2020-02-21 22:48:24 +05:30
Travis Nielsen d0c7d28a63 build: remove operator kit dependency
The operator kit had more utility originally when the operator
was creating and managing the TPRs and CRDs directly. Since
the CRDs are now created from a manifest and no longer by the
operators, the utility of operator kit is limited to the
controller watcher. Since we are moving to the controller runtime
we simplify the code to make the transition smoother. Now
there is only a simple WatchCR method that will need to be
replaced as we maek that transition.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-07 08:30:14 -07:00
Ashish Ranjan 8274aa2426 enhance(ceph): Adds status field for ceph related CRs
This commit adds status field for ceph related CRs which will be useful for knowing the ceph component status without checking the logs.

Signed-off-by: Ashish Ranjan <ashishranjan738@gmail.com>
2019-12-17 12:01:17 +05:30
Sébastien Han ad95c7296f ceph: do not print extended format for loggers.
When using the "errors" package, using `%+v` (extended format),
each Frame of the error's StackTrace will be printed in detail.
Let's only print `%v` to print the error.
If the error has a Cause it will be printed recursively.

Basically `%+v` has been replaced with `%v` for all `error` type
interfaces, whether the logger is Info, Warning or Error.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-16 18:35:54 +01:00
Sébastien Han 5ce2ed220e ceph: use "github.com/pkg/errors"
We now use the error package.
Kubernetes errors have been renamed kerrors since they are lower than
'errors'.

Closes: https://github.com/rook/rook/issues/4054
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-09 16:58:32 +01:00
Umanga Chapagain 6dd474d6f4 ceph: adds retry to ObjectUser creation
retry ObjectUser creation as long as valid input is provided
or it times out. Retry every 15sec for 5min.

integration test creates ObjectUser before ObjectStore to
ensure that ObjectUser waits for ObjectStore to be created
and available

Closes: https://github.com/rook/rook/issues/3937
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2019-11-05 12:45:45 +05:30
travisn 879971f19a ceph: create object user immediately if obj store available
Creating an object user was always waiting at least 15 seconds before
attempting to create the user. With this change the operator will check
immediately if the object store is already initialized instead
of waiting for the initial 15 seconds delay to elapse.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-19 22:46:17 -06:00
Sébastien Han 148f8da1fb ceph: relax pre-requisite for external cluster
We now differentiate the cases where:

* we only consume the external cluster
* we consume the external cluster as well as creating stateless
resources in Kubernetes (bootstrap mds,rgw, nfs)

This is mostly controlled via the image property spec. If not defined,
not extra CRs won't be able to be created.

Now the external cluster feature supports Ceph cluster as of Luminous 12.2.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-10-08 10:52:01 +02:00