Commit Graph
64 Commits
Author SHA1 Message Date
Sébastien Han ebbf332d8d ceph: update rgw and mds deployment for logCollector
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.

Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-22 18:19:07 +01:00
Sébastien Han 758226294b ceph: convert CephClient CRD to controller-runtime
This was the last remaining CRD to not use the controller-runtime
library.
Small additions were added with the transition:

* the Kubernetes Secret that contains the CephX key has now an owner
  reference to the CephClient object
* the secret name is present in the Status field of the CephClient:

```
status:
  info:
    secretName: rook-ceph-client-glance
  phase: Ready
```

The controller will reconcile on CR updates and also if the Kubernetes
Secret is deleted.

Closes: https://github.com/rook/rook/issues/4938
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-15 16:17:38 +01:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Travis Nielsen afeb38417d ceph: suppress reconcile error after operator restart
After the operator restarts, the controllers will not all be able
to reconcile until the ceph config has been generated by the reconcile
of the CephCluster controller. The message printed to the operator log
is frequently seen as an error condition even though it is a normal
condition where we requeue the reconcile until the config is available.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-07 17:24:13 -07:00
Travis Nielsen 156774c459 ceph: remove obsolete topology comments and fix comment typo
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-04 09:10:52 -07:00
Sébastien Han c6a87203ca ceph: add log collector
We can now collect logs directly into a side-car container.
A new CRD spec has been added:

spec:
  logCollector:
    enabled: true
    periodicity: 24h

Every 24h we will rotate log files for each Ceph daemon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-12-01 16:30:17 +01:00
Sébastien Han ad24990473 ceph: ability to abort orchestration
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):

* request for cluster deletion
* request for cluster upgrade

If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.

This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.

Operator log example:

```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
                Image:            "ceph/ceph:v15.2.5",
-               AllowUnsupported: true,
+               AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
-               Image:            "ceph/ceph:v15.2.5",
+               Image:            "ceph/ceph:v15.2.6",
                AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```

Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-25 15:32:04 +01:00
Sébastien Han 70c6752e0e ceph: add dedicated predicate for CephCluster object
Since we want to pass a context to it, let's extract the logic into its
own predicate.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-24 18:21:07 +01:00
Sébastien Han 9b9f45b792 ceph: export isDoNotReconcile function
Since the predicate for the CephCluster object will soon move into its
own predidacte we need to export isDoNotReconcile so that it can be
consummed by the "cluster" package.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-24 18:21:02 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Sébastien Han ea1d71cbfb ceph: add vault kms support for osd encryption
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/

The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:

security:
  kms:
    tokenSecretName: <name of the secret containing a Vault token, used
    to authenticate>
    connectionDetails: < a map of strings containing connection
    information>

Refer to the ceph-cluster-crd documentation to lear more.

Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-30 16:16:33 +01:00
Lalit Maganti 88f16e4980 ceph: ignore MDS_ALL_DOWN during reconciliation
This allows the catch-22 situation where the filesystem cannot be
reconciled because there is no MDS but there is no MDS because the
operator has not reconciled the filesystem and brought up the MDS pods.

Closes #5967, #5846

Signed-off-by: Lalit Maganti <lalitm@google.com>
2020-10-29 01:11:36 +00:00
Santosh Pillai 70c5f70701 ceph: support IPv6 single-stack for ceph
Add --ms-bind-ipv6 true args to mom, osd, mgr, object, mds, rbd daemons.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-10-02 01:03:03 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
Sébastien Han 204ffd4065 Merge pull request #6250 from leseb/mirroring-config
ceph: add rbd-mirror configuration
2020-09-18 09:16:08 +02:00
Sébastien Han 451622a955 ceph: add rbd-mirror configuration
Rook is now capable of configuring mirroring between sites. The
implementation works at different levels:

* CephBlockPool: which introduces a new `mirroring` configuration as well
as `statusCheck`. When turned on, Rook will enable mirroring on the
pool. It will also create a bootstrap peer token and store it in a
Kubernetes Secret. The name of that Secret can be found in the Status
field of the CephBlockPool CRD. This token can be fetched and used by
other clusters to configure the site as a peer. Mirroring can be
configured either at the pool or the image level.

* CephRBDMirror: which introduces a new `peers` configuration allowing
Rook to connect to peers by passing a Secret name. The administrator will
create a Kubernetes Secret with 2 keys: 'token' for the bootstrap peer
token and 'pool' for the name of pool. Once detected the rbd-mirror
controller will go ahead and import the peer configuration.

Pool mirroring status example:

```
status:
  info:
    rbdMirrorBootstrapPeerSecretName: pool-peer-token-test
  mirroringInfo:
    lastChanged: "2020-09-17T14:47:27Z"
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        mode: image
        peers:
        - client_name: client.rbd-mirror-peer
          direction: rx-tx
          mirror_uuid: ""
          site_name: rhcs
          uuid: c50522a4-28a4-4bd3-ba68-e11780308882
        site_name: 91eae0dd-06b1-4d2c-91f3-1311c9df382b-rook-ceph
  mirroringStatus:
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        daemon_health: OK
        health: OK
        image_health: OK
        states:
          replaying: 1
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-17 19:05:04 +02:00
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
Sébastien Han bd90f9afe9 ceph: add more error handling and debug logging
The predicate was lacking from very useful errors as well as logging
messages.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Sébastien Han e7043edbd8 ceph: reconcile if cm is config override
Earlier, we were returning only when the object was not the config
override configmap, we want the opposite.
Also, this was blocking all subsequent conditions and add the correct
check to avoid cm to reconcile.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Travis Nielsen ce92249725 ceph: continue with memory limits below min settings
The operator will now allow the resource limits to be applied below the
recommended minimums. In small clusters, even the recommended minimums
may not be necessary. A warning is still printed to the operator log,
but we allow the configuration to continue.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-19 11:07:16 -06:00
Sébastien Han af1e8c320a ceph: do not log an error if no clusters
If the list of cluster is empty there is no need to report an error in
the logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-05 10:05:51 +02:00
Travis Nielsen 06c21954e1 ceph: labels are immutable for match selectors
The labels cannot be updated on a deployment match selector.
In v1.4 a new label was added for ceph_daemon_type that is intended
to be on the pod labels, but cannot be applied to the match selectors.
Therefore, we suppress any new labels that are added to the daemons.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-04 16:47:00 -06:00
Blaine Gardner dcf8ae2526 ceph: rename pod labels function for more clarity
Rename PodLabels function to CephDaemonAppLabels for more clarity about
what the function's purpose is.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:30 -06:00
Blaine Gardner 9b6114c192 ceph: add 'ceph_daemon_type' label to pod labels
This can help users identify the daemon type similarly to how
'ceph_daemon_id' helps identify the daemon ID. This also makes it so
scripts users create don't have to parse "app=rook-ceph-<daemonType>"
into "<daemonType>" if they desire that bit of info.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:22 -06:00
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen 4681f9e73d ceph: consolidate ceph config and client packages
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:43 -06:00
Sébastien Han 56954d5cc8 ceph: fix CRs not reconciling on updates
Most of the CRs expect CephCluster and CephBlockPool were not
reconciling on updates. The predicate must return true on CR diff.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-15 11:15:33 +02:00
Travis Nielsen 8f9df75f7b ceph: capture logging from the osd daemon
The OSD daemon has been missing critical flags for logging to stderr
where k8s can capture the logs. Without the --log-to-stderr=true,
all the OSD logging was essentially lost until now.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-06 14:38:37 -06:00
Sébastien Han aca541803b Merge pull request #5664 from leseb/ceph-object-user-external
ceph: add external support for object store user
2020-06-18 20:10:05 +02:00
Sébastien Han 84d1e28c99 ceph: add external support for objectstoreuser
Now, the object store user is capable of creating s3 users on an
external Ceph cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 16:34:01 +02:00
Sébastien Han 1e4a4b477a Merge pull request #5111 from leseb/lang-neutral-ceph
ceph: use octopus base image
2020-06-18 10:46:54 +02:00
Sébastien Han 68e62836c5 ceph: use newer octopus time format
radosgw-admin as of Octopus uses a different time format, it uses
"2006-01-02T15:04:05.999999999Z". It's close from RFC3339 but not quite
the same. This change is needed in order for the operator image (with a
Ceph Octopus based image) to perform radosgw-admin call correctly.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-17 23:00:17 +02:00
Sébastien Han c7f255a0e8 ceph: add the ability to set any pool property
We can now explicitly set any property on a given pool by using the new
Property field in the CephBlockPool Spec.

Also, this fixes the case where both `CephBlockPool` and `CephCluster`
are created at the same time. When Rook creates the pool, the cluster is
still being bootstrapped and the global option
`osd_pool_default_pg_autoscale_mode` has not bee set yet. So the pool
gets created but its `pg_autoscale_mode` property is set to `warn`
instead of `on`.

Closes: https://github.com/rook/rook/issues/5608V
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-17 10:08:00 +02:00
Ali Maredia fc579f4520 ceph: initial commit for ceph rgw multisite resources
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone

Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-06-09 16:28:53 -04:00
Sébastien Han c8de7009de ceph: increase liveness probe start delay on OSD
When the cluster is loaded and we restart an OSD, it will need some time
to respond to socket calls, basically more to be ready.
Increasing the initialDelaySeconds of the liveness probe fixes that
issue. For OSD, it waits for 45 sec, where other daemons 10 sec.

Closes: https://github.com/rook/rook/issues/5492
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-05 15:53:49 +02:00
Travis Nielsen d1da12c6ac ceph: wait indefinitely for cleanup before removing cluster finalizer
During cluster deletion, we currently only retry for a couple minutes
to wait for the pvcs to be deleted. After the timeout, we proceed
with the cluster deletion. To properly protect the pvcs for proper
cleanup, the finalizer should not be removed until the pvcs
are all confirmed to be deleted. In order to not block other cluster
events, we re-queue the deletion event to run again every 10s
until the pvcs are deleted or the finalizer is manually removed.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Travis Nielsen f47bb945c2 ceph: remove duplicate controller wait setting
The controller setting for requeuing an event moved to the
opcontroller package and was no longer needed in the main
controller package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Madhu Rajanna 81688398f2 cleanup: use err.Wrap when the formatting is not required
Replaced err.Wrapf with err.Wrap when the formatting
is not required.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-29 17:42:59 +05:30
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Travis Nielsen 8d72e788c1 ceph: lookup of the cluster cannot rely on namespace name
The name of a CephCluster is commonly the same as the namespace,
but not always. When looking up the ceph cluster we can look for
the first one in the namespace rather than requiring the name
to be the same as the namespace.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-22 13:42:40 -06:00
Sébastien Han c1b8a7aa9e ceph: extract rbd-mirror to its own crd
Previously, the rbd-mirror daemon was integrated into the `CephCluster`
CRD. This wasn't really practical since we would have to wait for the
whole orchestration to be done to actually set it up. The same goes for
any CR update. Let's say you want to change the number of daemons, Rook
would go through mons, mgrs and osds until it get to rbd-mirror.
This triggers an undesired full orchestration.

With its own CRD this component just gains a lot more flexibility.

Closes: https://github.com/rook/rook/issues/5084
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-15 09:12:56 +02:00
Travis Nielsen ff1ab64d8d ceph: skip reconcile events where the spec not updated
The controllers should ignore events for status updates. The reconcile
only needs to happen when the spec is updated or the resource is marked
for deletion. Otherwise, the controllers may stay in an update loop as
the reconcile events are triggered again every time there is a status
update during the reconcile.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-09 11:20:45 -06:00
Sébastien Han 71799682ca ceph: print the reason for reconciling
General debugging helper without having to turn on DEBUG logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-06 16:07:01 +02:00
Travis Nielsen eeb687ad14 ceph: log skipped reconcile based on ceph health
When the Ceph health is HEALTH_ERR the controllers will skip the
reconcile. With info level logging there needs to be an indication
of this decision to skip the reconcile, otherwise users will
wonder why their resources aren't being created.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-06 16:07:01 +02:00
Sébastien Han c5075d55d1 ceph: trigger update on child controllers
When the CephCluster CR gets updated with a new image version, we need
to notify the child controllers of that upgrade.
For this, we set a 'ceph_version' label on the CR itself which will
effectively trigger a reconcile based on an update event.
This needs to be revisited and see how we can make use of
`EnqueueRequestsFromMapFunc` which might be a better approach using
controller-runtime built-in.

Closes: https://github.com/rook/rook/issues/5153
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-03 13:56:49 -06:00
Travis Nielsen 96dd92cc7f ceph: change verbose info log to debug in controller
Keeping the operator log clean...

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-03 09:15:25 -06:00
Sébastien Han ed37a6a699 Merge pull request #5128 from leseb/liveness-other-daemons
ceph: add liveness probe to mon, mds and osd daemons
2020-04-01 10:53:12 +02:00
Sébastien Han 040193bb5a ceph: remove DaemonType type
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-01 09:08:18 +02:00