The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Since the mock client was added to the rgw/admin package from the
go-ceph project in https://github.com/ceph/go-ceph/pull/532 we can now
mock the HTTP client and validate the entire reconcile.
Signed-off-by: Sébastien Han <seb@redhat.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
The CephClusterSpec was missing from the object context, so the check
for the network provider in RunAdminCommandNoMultisite() was not
discovering the network property correctly.
Signed-off-by: Sébastien Han <seb@redhat.com>
The Rados Gateway Admin OPS API has changed its behavior from Nautilus
to Pacific. Calling user create on an existing user wil generate
additional keys to the user on Nautilus. Where in Pacific it will report
an error with UserAlreadyExists.
So to handle both scenarios, let's first get the user, and if the user
does not exist (NoSuchUser) we then create it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Prior to that the access and secret keys were left empty if the user
already existed, which led to updating the secret with empty values when
the operator restarts.
Signed-off-by: Sébastien Han <seb@redhat.com>
For the TLS communication for AdminOps Api, httpclient is required and
filled with TLS certs, currently it is set to nil pointer in
`buckethealthchecker` and `cephobjectstoreuser`.
Thanks @Krast76 finding the issue even proposing the fix.
Fixes: #8132
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.
Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
when store.Status was not initialized it caused
crashes when assigning store.Status.Info["endpoint"]
to objContext.Endpoint.
This commits makes sure store.Status is not Nil
before setting objContext.Endpoint and
objContext.SecureEndpoint in the ObjectStoreUser
controller.
Closes: https://github.com/rook/rook/issues/6916
Signed-off-by: Ali Maredia <amaredia@redhat.com>
This fixes libraries that were being imported multiple times in
pkg/operator/ceph/object and its subpackages.
Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
If the object store status is not yet initialized, the object store user should
fail its initialization and requeue the reconcile.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Currently endpoint is filled from `Status.Info["endpoint"]` string map for cephobjectstore user's secret and
have two issues:
1.) endpoint is nil and only secureendpoint is available
2.) Chance of small race window in which endpoint is not filled before this ceph object user is handling.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
When creating object store, `radosgw-admin realm get ..` command is stuck forever when required number of OSDs are not available. Because of this the uninstall of object store is also stuck. User has to manually remove the finalizer to delete the object store.This PR uses `ExecuteCommandWithTimeout` for running `radosgw-admin` command. Timeout during installation will be reconciled. Cleanup will be treated as best effort. Any errors during uninstalling of single site object store will only be logged.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.
Signed-off-by: Sébastien Han <seb@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.
- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up
- Added more debug logging all around the object-store
code related to multisite
- files generated by rerun of `make codegen`
- change back the edit on the Copyright in object.go
Signed-off-by: Ali Maredia <amaredia@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator should only connect to ceph with a single set of creds.
In a converged cluster this will be the admin creds and in an external
cluster it will be lower-privileged creds. Independent clusters were
implemented with a separate set of creds. To simplify the code these
are now merged to a single set.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
adding method to generate secret name in
cephobjectstoreuser status field in order
to programmatically retrieve the secret name
generated by the cephobjectstoreuser crd.
Also added unit test for the method.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
Now, when an S3 user gets created, Rook will add the S3 endpoint to the
Secret along with the credentials.
Signed-off-by: Sébastien Han <seb@redhat.com>
The controller setting for requeuing an event moved to the
opcontroller package and was no longer needed in the main
controller package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
The name of a CephCluster is commonly the same as the namespace,
but not always. When looking up the ceph cluster we can look for
the first one in the namespace rather than requiring the name
to be the same as the namespace.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When updating the status on resources sometimes the update fails due to
the resource being an outdated version. This frequently occurs when
the same reconcile loop updates the status multiple times, or the finalizer
is added, or some other update to the resource. Upon the next reconcile
the error would generally go away since the status or finalizer didn't
need to be updated multiple times. But now the error is not expected
since the controller will retrieve the latest version of the resource
immediately before attempting to update it.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Removed some extremely verbose debug logging that makes
the logs difficult to read when debug is enabled.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.
Signed-off-by: Sébastien Han <seb@redhat.com>
Let's not wait for the CephCluster to be done reconciling but instead
check for the Ceph cluster status, if it's closed to "ok" then we
proceed so HEALTH_OK and HEALTH_WARN are accepted.
Signed-off-by: Sébastien Han <seb@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>