51 Commits
Author SHA1 Message Date
Joshua Hoblitt 9eec9c730e core: rm unused consts
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-05-14 09:34:52 -07:00
Travis Nielsen 023608e6fd core: enhance logging with namespaced names
For all of the controllers besides the cluster controller,
the logging now includes the namespaced name of the resource
that is being reconciled. This will help with log troubleshooting
to help analyze logs consistently for the resource being
reconciled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-12-04 12:00:42 -07:00
cuiweixie 20b3b9deef operator: refactor to use reflect.TypeFor
Signed-off-by: cuiweixie <cuiweixie@gmail.com>
2025-08-28 23:02:19 +08:00
Oded Viner cf13deee6f core: log panics in controller reconcile functions
add RecoverAndLogException() helper to log panics with stack trace.
added defer call in all rook controller Reconcile() methods for
better error visibility in operator logs

Signed-off-by: Oded Viner <oviner@redhat.com>
2025-07-29 13:20:38 +03:00
Joshua Hoblitt 3cb343f62a core: run gofumpt on all files
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2025-03-26 10:41:48 -07:00
Travis Nielsen 3bd5881fe5 core: suppress mgr module health errors during reconcile
Some ceph health errors should not block the reconcile
of the cluster. Mgr modules do not have cause to block
the reconcile, as the cluster can usually work even
if a module is failing.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-02-27 08:10:52 -07:00
Travis NielsenandDmitry Mishin d109dc9029 core: implement operator settings as env vars
The operator settings loaded from the configmap have proven
inefficient for load time and frequently checking the configmap.
To avoid this ineffenciency, the configmap is only loaded once
each time it is created or updated. The values in the configmap
are applied as environment variables, which then are very efficient
to query throughout the various controllers, without needing
to be concerned about loading the configmap again.

Co-authored-by: Dmitry Mishin <dmitry.mishin@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-02-24 12:12:45 -07:00
Blaine Gardner 0e33536539 object: disallow unsafe OBC fields by default
Implement an allow list mechanism that disables potentially unsafe OBC
fields by default. OBC fields beyond `maxObjects` and `maxSize` don't
neatly fit into the OBC framework as it was originally envisioned and
implemented.

Some of the newly added configs could allow users to cause confusion for
themselves. Others might allow users to hijack others buckets. Some
might allow bricking the entire S3 store.

Out of an abundance of safety, allow-list the known-safe options by
default, and require administrators to enable potentially troublesome
options via the new operator-level config
`ROOK_OBC_ALLOW_ADDITIONAL_CONFIG_FIELDS`.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2025-02-12 14:18:01 -07:00
Michael AdamandTravis Nielsen 70d4f5d4b7 core: fix the revisionHistoryLimit implementation and test
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
2024-11-19 10:55:26 +01:00
Michael Adam ab8fd90aa6 core: add ROOK_REVISION_HISTORY_LIMIT operator setting
This adds an operator config setting ROOK_REVISION_HISTORY_LIMIT
defaulting to kubernetes'value for RevisionHistoryLimit.

If configured, the provided value will be used as RevisionHistoryLimit

for all Deployments rook creates.

Fixes: #12722

Signed-off-by: Michael Adam <obnox@samba.org>
2024-10-02 19:40:37 +02:00
Michael Adam e378588359 network: add a new operator config setting ROOK_ENFORCE_HOSTNETWORK
This new setting is of Boolean type and defaults to "false".

    When set to "true", it changes the behavior of the
     rook operator to
    nable host network on all pods created by the cephcluster controller

     new method to check the setting:  opcontroller.EnForceHostNetwork()

Signed-off-by: Michael Adam <obnox@samba.org>
2024-09-05 17:34:13 +02:00
Tarun Gupta Akirala df0ce26923 core: typo in logs to print fullname of CephCluster
printing namespace along with name would make debugging easier

Signed-off-by: Tarun Gupta Akirala <takirala@users.noreply.github.com>
2023-06-07 14:42:00 -07:00
Shinya Hayashi 05875a3f4f osd: support loop devices for test clusters
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.

This feature is intended to be used for testing purposes only.

Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
2022-11-09 06:45:41 +00:00
Sébastien Han 05506e7a68 core: reload go routine after CR is edited
Previously, the struct maintaining the list of cluster was still
initialized with a cluster item. Then the monitoring check will see that
the cluster is part of the struct already and thus won't run the
monitoring go routine again.
Now each time we cancel the context, we also remove the cluster item
from the map so that when the controller runs again, the monitoring
struct is re-populated and the go routine runs and statuses are updated.

Closes: https://github.com/rook/rook/issues/9911
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-03-31 08:49:42 +02:00
Travis Nielsen 7ee9cc9d56 pool: allow configuration of built-in pools with non-k8s names
The built-in pools device_health_metrics and .nfs created by ceph
need to be configured for replicas, failure domain, etc.
To support this, we allow the pool to be created as a CR.
Since K8s does not support underscores in the resource names
the operator must translate this special pool name into
the name expected by ceph.

This also sets the basis for allowing filesystem data
pools to specify the desired pool name instead of requiring
a generated name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-15 15:32:02 -07:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Satoru Takeuchi 30e4fbb01f ceph: make the timeout of ceph commands cofigurable
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-27 12:42:16 +00:00
Travis Nielsen c39c1c7ddf ceph: retry reconcile immediately after cancellation
If the reconcile is cancelled due to a CR update, we want to retry the next
reconcile immediately rather than wait for the exponential backoff timeout
if the reconcile was already failing. The wait can easily be minutes
if the reconcile was in this state, which makes it appear the operator
is ignoring the request to start a new reconcile.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-02 08:23:25 -06:00
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
subhamkrai 05d4c2776c ceph: update cephCluster CR with ceph versions output
update cephCluster CR with ceph versions output.
this output will contain ceph version of the
ceph daemons which will help with upgrade status.

ceph versions command output
```
ceph versions
{
    "mon": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 3
    },
    "mgr": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "osd": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "mds": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 2
    },
    "rbd-mirror": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "rgw": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "overall": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 9
    }
}
```

CR status
```
status:
  ceph:
   ---
    versions:
      mds:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 2
      mgr:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      mon:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 3
      osd:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      overall:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 9
      rbd-mirror:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      rgw:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
```

Signed-off-by: subhamkrai <srai@redhat.com>
2021-03-29 21:04:35 +05:30
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
Travis Nielsen 64e28af741 ceph: allow flex driver and discovery to be enabled with configmap
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 08:39:26 -06:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Sébastien Han 8d033efb5a ceph: silence harmless errors
If the ceph cli outputs an error with "error calling conf_read_file"
this means that the operator has not written its ceph configuration
file. Thus ceph cli commands will fail, so we can just ignore that since
the operator will soon write this file in its initialization sequence.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-03 17:28:44 +01:00
Travis Nielsen afeb38417d ceph: suppress reconcile error after operator restart
After the operator restarts, the controllers will not all be able
to reconcile until the ceph config has been generated by the reconcile
of the CephCluster controller. The message printed to the operator log
is frequently seen as an error condition even though it is a normal
condition where we requeue the reconcile until the config is available.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-07 17:24:13 -07:00
Travis Nielsen 156774c459 ceph: remove obsolete topology comments and fix comment typo
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-04 09:10:52 -07:00
Sébastien Han ad24990473 ceph: ability to abort orchestration
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):

* request for cluster deletion
* request for cluster upgrade

If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.

This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.

Operator log example:

```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
                Image:            "ceph/ceph:v15.2.5",
-               AllowUnsupported: true,
+               AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
-               Image:            "ceph/ceph:v15.2.5",
+               Image:            "ceph/ceph:v15.2.6",
                AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```

Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-25 15:32:04 +01:00
Lalit Maganti 88f16e4980 ceph: ignore MDS_ALL_DOWN during reconciliation
This allows the catch-22 situation where the filesystem cannot be
reconciled because there is no MDS but there is no MDS because the
operator has not reconciled the filesystem and brought up the MDS pods.

Closes #5967, #5846

Signed-off-by: Lalit Maganti <lalitm@google.com>
2020-10-29 01:11:36 +00:00
Sébastien Han af1e8c320a ceph: do not log an error if no clusters
If the list of cluster is empty there is no need to report an error in
the logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-05 10:05:51 +02:00
Sébastien Han aca541803b Merge pull request #5664 from leseb/ceph-object-user-external
ceph: add external support for object store user
2020-06-18 20:10:05 +02:00
Sébastien Han 84d1e28c99 ceph: add external support for objectstoreuser
Now, the object store user is capable of creating s3 users on an
external Ceph cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 16:34:01 +02:00
Sébastien Han 1e4a4b477a Merge pull request #5111 from leseb/lang-neutral-ceph
ceph: use octopus base image
2020-06-18 10:46:54 +02:00
Sébastien Han 68e62836c5 ceph: use newer octopus time format
radosgw-admin as of Octopus uses a different time format, it uses
"2006-01-02T15:04:05.999999999Z". It's close from RFC3339 but not quite
the same. This change is needed in order for the operator image (with a
Ceph Octopus based image) to perform radosgw-admin call correctly.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-17 23:00:17 +02:00
Sébastien Han c7f255a0e8 ceph: add the ability to set any pool property
We can now explicitly set any property on a given pool by using the new
Property field in the CephBlockPool Spec.

Also, this fixes the case where both `CephBlockPool` and `CephCluster`
are created at the same time. When Rook creates the pool, the cluster is
still being bootstrapped and the global option
`osd_pool_default_pg_autoscale_mode` has not bee set yet. So the pool
gets created but its `pg_autoscale_mode` property is set to `warn`
instead of `on`.

Closes: https://github.com/rook/rook/issues/5608V
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-17 10:08:00 +02:00
Travis Nielsen d1da12c6ac ceph: wait indefinitely for cleanup before removing cluster finalizer
During cluster deletion, we currently only retry for a couple minutes
to wait for the pvcs to be deleted. After the timeout, we proceed
with the cluster deletion. To properly protect the pvcs for proper
cleanup, the finalizer should not be removed until the pvcs
are all confirmed to be deleted. In order to not block other cluster
events, we re-queue the deletion event to run again every 10s
until the pvcs are deleted or the finalizer is manually removed.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Travis Nielsen f47bb945c2 ceph: remove duplicate controller wait setting
The controller setting for requeuing an event moved to the
opcontroller package and was no longer needed in the main
controller package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Travis Nielsen 8d72e788c1 ceph: lookup of the cluster cannot rely on namespace name
The name of a CephCluster is commonly the same as the namespace,
but not always. When looking up the ceph cluster we can look for
the first one in the namespace rather than requiring the name
to be the same as the namespace.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-22 13:42:40 -06:00
Travis Nielsen eeb687ad14 ceph: log skipped reconcile based on ceph health
When the Ceph health is HEALTH_ERR the controllers will skip the
reconcile. With info level logging there needs to be an indication
of this decision to skip the reconcile, otherwise users will
wonder why their resources aren't being created.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-06 16:07:01 +02:00
Sébastien Han 8db14885b5 ceph: controller fix misleading debug log
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 19:12:17 +01:00
Sébastien Han c84d66de1b ceph: controller, reconcile faster
Let's not wait for the CephCluster to be done reconciling but instead
check for the Ceph cluster status, if it's closed to "ok" then we
proceed so HEALTH_OK and HEALTH_WARN are accepted.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 18:45:19 +01:00
Sébastien Han b5e50481ff ceph: controller: retry after 10sec even if no cluster
Even if there is no CephCluster we still want to retry every 10sec as
one will likely show up soon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Sébastien Han 7b544f04d4 ceph: reconcile more often when cluster is not ready
Let's not apply the exponential backoff when waiting for the cluster to
be ready, let's only apply it when there is no cluster.

Closes: https://github.com/rook/rook/issues/5059
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Travis Nielsen 18b0e7d295 ceph: scrub ceph commands to write actions to the log
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Umanga Chapagain 06ba93e629 Ceph: refactor GetOperatorSetting for reuse
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-09 21:41:56 +05:30
Sébastien Han 9f2867e12a ceph: separate controller for CephObjectStoreUser CRD
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-06 11:53:40 +01:00
Sébastien Han a137b31e1a ceph: refactor controller helper
Clean and refactor code helper for controller-runtime.
Implement those into the block pool controller.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-05 12:37:44 -07:00