Commit Graph
30 Commits
Author SHA1 Message Date
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Blaine Gardner 34a8b42e97 test: try to un-flake multi-cluster-mirror test
Try to un-flake the multi-cluster-mirror test that keeps failing on this
PR.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-13 11:06:51 -06:00
Blaine Gardner 65c4972038 rgw: make CephObjectRealm controller idempotent
The CephObjectRealm controller would fail all subsequent reconciles if
the first reconcile created the Kubernetes Secret containing the access
keys for the realm but where the radosgw-admin command failed to create
the realm. This was the only idempotency issue found after reviewing the
CephObjectRealm controller.

Resolves #8954

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-13 11:06:51 -06:00
Blaine Gardner eadcd757b3 rgw: replace period update --commit with function
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.

Resolves #8879

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 14:45:39 -06:00
Blaine Gardner 7b9293624a rgw: update period if period does not exist
Rook should update the RGW object store's period if the period doesn't
yet exist. This protects us from the case where the
'radosgw-admin period update --commit` command fails and the
CephObjectStore controller reconciles again.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-28 14:17:04 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Jiffin Tony Thottan ca43800119 ceph: add options for cephobjectstore user
Adding options for quota, bucket limit, caps for the
`cephobjectstoreuser`.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-07 22:43:09 +05:30
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
subhamkrai bc48dad19a ceph: update rgw ceph dashboard command
in the latest ceph releases, there are few
changes in accessing the rgw ceph dashboard
command. This commit updates the commands.

old commands
```
$ ceph dashboard set-rgw-api-access-key <access_key>
$ ceph dashboard set-rgw-api-secret-key <secret_key>
```
new commands
```
$ ceph dashboard set-rgw-api-access-key -i <file-containing-access-key>
$ ceph dashboard set-rgw-api-secret-key -i <file-containing-secret-key>
```

Signed-off-by: subhamkrai <srai@redhat.com>
2021-04-15 17:32:46 +05:30
Satoru Takeuchi a89d35e5b6 ceph: fix improper json parsing in radosgw-admin
Sometimes `radosgw-admin` succeeds after showing logs to stderr. We should
skip non-json strings if parsing output as json.

Here is an example.

```
2021-02-26 04:10:44.190418 I | op-bucket-prov: creating Ceph user "ceph-user-aSzNqgE7"
E0226 04:11:37.901310       8 controller.go:199] error syncing 'logging/loki-bucket': error provisioning bucket: Provision: can't create ceph user: error creating ceph user "ceph-user-aSzNqgE7": failed to unmarshal json. 2021-02-26T04:11:21.425+0000 7f6714be4980  1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:21.426+0000 7f6714be4980  0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.uid:ceph-user-aSzNqgE7
2021-02-26T04:11:32.168+0000 7f6714be4980  1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:32.168+0000 7f6714be4980  0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.keys:23Z8GUEXR0TJDO86PSJR
{
    "user_id": "ceph-user-aSzNqgE7",
    "display_name": "ceph-user-aSzNqgE7",
...
    "mfa_ids": []
}: invalid character '-' after top-level value: failed to unmarshal json.
```

In this case, some logs like "robust_notify:..." was shown in stderr.
Unmarsharing was failed due to tried to parse these logs as json.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-10 16:06:56 +00:00
Santosh Pillai 113376f1ab ceph: timeout radosgw-admin cli commands
When creating object store, `radosgw-admin realm get ..` command is stuck forever when required number of OSDs are not available. Because of this the uninstall of object store is also stuck. User has to manually remove the finalizer to delete the object store.This PR uses `ExecuteCommandWithTimeout` for running `radosgw-admin` command. Timeout during installation will be reconciled. Cleanup will be treated as best effort. Any errors during uninstalling of single site object store will only be logged.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-02-01 10:57:04 +05:30
Travis Nielsen 464e332dc1 ceph: init rgw dashboard access key in goroutine
The command to set or disable the rgw dashboard started
hanging in some scenarios in v15.2.8. For now we start
the rgw dashboard config in a goroutine until this issue
is tracked down. After the issue is fixed in ceph, the
goroutines will no longer be necessary.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-07 07:08:20 -07:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
Jiffin Tony Thottan ec5f13318c ceph: enable/disable dashboard for rgw
Provide permission for dashboard to collect the object store metrics.

Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2020-08-03 10:26:54 +05:30
Ali Maredia e6ed4ff8ea ceph: minor fixes + add realm/zg/zone to object context
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.

- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up

- Added more debug logging all around the object-store
code related to multisite

- files generated by rerun of `make codegen`

- change back the edit on the Copyright in object.go

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-07-21 16:47:59 -04:00
Ali Maredia 5d9c2f09a9 ceph: enable the realm pull on ceph-object-realm
This commit:
- adds the pull section on the CephObjectRealm spec
to enables realms to be pulled instead of created.

- creates the system user and generates the access
key and secret key for the user and zones in a realm.

- adds "omitempty" to fields in multisite CRs that
are not required

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-07-21 13:30:19 -04:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Ali Maredia fc579f4520 ceph: initial commit for ceph rgw multisite resources
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone

Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-06-09 16:28:53 -04:00
Santosh PillaiandElise Gafford 5ef88f8f09 ceph: finalizer for OBC cleanup
This PR is for rebasing #4683 with master. With new controller runtime changes most the changes with #4683 got redundant.
Only valid change is waiting for all the object buckets to be cleaned up before removing the finalizer.

Co-authored-by: Elise Gafford <egafford@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-05-18 23:54:59 +05:30
Travis Nielsen 5d2db6a9f1 ceph: allow creation of object store with pre-existing pools
The pools may already be created by the mgr module. If the pool specs
are not specified in the object CR, skip pool creation for the
object store. The pools are required to exist and the reconcile
will fail until they do exist.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-23 23:54:20 -06:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Sébastien Han 5ce2ed220e ceph: use "github.com/pkg/errors"
We now use the error package.
Kubernetes errors have been renamed kerrors since they are lower than
'errors'.

Closes: https://github.com/rook/rook/issues/4054
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-09 16:58:32 +01:00
Owen Tuz 15e3ecad54 ceph: add 'rgw.buckets.non-ec' to list of RGW metadataPools
This is used for S3 multipart uploads and should be configured in the same way as other metadata pools.

Signed-off-by: Owen Tuz <owen@segfault.re>
2019-10-17 18:44:59 +01:00
Juan Miguel Olmo Martínez 2de0787fc8 New setting to avoid delete pools on <fs>/<os> resources deletion
A new CRD property `PreservePoolsOnDelete` has been added to Filesystem(fs) and
Object Store(os) resources in order to increase protection against data loss.
If it is set to `true`, associated pools won't be deleted when the main
resource(fs/os) is deleted. Creating again the deleted fs/os with the same name
 will reuse the preserved pools.

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2019-09-30 08:13:39 +02:00
Madhu Rajanna 4efba0247d Check Pools is in use before deleting it
currently we are not checking the pool is empty or
not before deleting, This PR adds a check to check
if any images/snapshots using the pool, if yes it will
not delete the pool

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-26 11:41:47 +05:30
c852e86fe9 ceph-rook object bucket provisioner
Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: jeffvance <jeff.h.vance@gmail.com>
Signed-off-by: Jon Cope <jcope@redhat.com>

Co-authored-by: Jon Cope <copejon@users.noreply.github.com>
Co-authored-by: Jeff Vance <jeff.h.vance@gmail.com>
2019-08-22 19:54:43 -06:00
Blaine Gardner 086fa8231c rgw: configure entirely in operator
Configure the Ceph rgw daemon completely from the operator a la the
recent changes to the Ceph mon, mgr, and mds operators.

Create the rgw deployment or daemonset first, and then create the
keyring secret for the object store with its owner reference as the
corresponding deployment or daemonset. When the replication controller
is deleted, the secret is also deleted.

The RGW's mime.types file is now stored in a configmap with a different
file created for each object store. This is primarily just a means to
get the mime.types file into the rgw pod, but the added benefit is that
the administrator can modify the configmap, which could reduce
susceptibility to file type execution vulnerabilities (worst case).

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-03-07 08:02:42 -07:00