Commit Graph
88 Commits
Author SHA1 Message Date
Sébastien Han db91fd8bc3 ceph: do not configure external metric endpoint is not present
If no endpoint are configured let's simply return.

Closes: https://github.com/rook/rook/issues/7963
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-24 14:41:26 +02:00
Sébastien Han 5e1e9f4c92 ceph: actively update the service endpoint for external mgr
If the cluster is external we want to periodically rehydrate the mgr
endpoint. This handles the scenarion where the active manager changes,
 so we need to update the endpoint with the new IP address.
The create-external-cluster-resources.py script now requires an extra
permission to query the manager service so Rook can discover the active
one and its IP.

Testing:

```
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 37,
    "available": true,
    "active_name": "b",
    "num_standby": 1
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.12:9283   3m10s

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] k scale --replicas=0 deployment rook-ceph-mgr-b
deployment.apps/rook-ceph-mgr-b scaled

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 40,
    "available": true,
    "active_name": "a",
    "num_standby": 0
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.13:9283   3m55s
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-11 16:08:40 +02:00
parth-gr bb23622b72 core: enhancement in debug logs
When DEBUG logging is enabled in the operator, there are a number of overwhelming messages in the log that are very overwhelming and don't seem useful, which makes debug mode difficult to use
Updated and Cutted down the messages that are not useful in debug mode

Closes: https://github.com/rook/rook/issues/7499
Signed-off-by: parth-gr <paarora@redhat.com>
2021-04-28 22:28:10 +05:30
Blaine Gardner 9d657f464c ceph: redact secret info from reconcile diffs
To avoid logging sensitive information, do not output diffs from secrets
when reconciling resources.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-15 08:14:34 -06:00
Lars Lehtonen 41c567beaf ceph: fix multiple imports
This fixes additional double-imports.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-13 02:02:09 -07:00
Travis Nielsen 721acd1a8a Revert "ceph: added rook-ceph-default service account"
This reverts commit 737fb099fe.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-08 11:46:12 -06:00
Sébastien Han 51e7310e09 ceph: add dualstack support on pacific
With Pacific comes the support for dualstack where ceph daemons can
listen on both ipv4 and ipv6 stacks.
A new field in the network spec has been added: `dualStack`

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-31 18:04:29 +02:00
Travis Nielsen 3c37def52f Merge pull request #7406 from parth-gr/service-account
ceph: added rook-ceph-default service account
2021-03-30 08:06:40 -06:00
Sébastien Han 0c799c3b8f Merge pull request #7476 from subhamkrai/ceph-versions
ceph: update cephCluster CR with ceph versions output
2021-03-30 09:10:01 +02:00
Blaine Gardner 1df4336a68 Merge pull request #7386 from BlaineEXE/update-osds-in-parallel
ceph: Update osds in parallel
2021-03-29 16:32:07 -06:00
Blaine Gardner 795124b7a8 ceph: update osds in parallel
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md

The max number of OSDs updated in parallel is currently fixed at 20.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-29 10:55:28 -06:00
Blaine Gardner 660d2fe442 Merge pull request #7493 from BlaineEXE/codespell-tidy
test: document or clean up codespell flags
2021-03-29 10:07:53 -06:00
subhamkrai 05d4c2776c ceph: update cephCluster CR with ceph versions output
update cephCluster CR with ceph versions output.
this output will contain ceph version of the
ceph daemons which will help with upgrade status.

ceph versions command output
```
ceph versions
{
    "mon": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 3
    },
    "mgr": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "osd": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "mds": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 2
    },
    "rbd-mirror": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "rgw": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "overall": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 9
    }
}
```

CR status
```
status:
  ceph:
   ---
    versions:
      mds:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 2
      mgr:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      mon:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 3
      osd:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      overall:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 9
      rbd-mirror:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      rgw:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
```

Signed-off-by: subhamkrai <srai@redhat.com>
2021-03-29 21:04:35 +05:30
parth-grandTareq Sharafy 737fb099fe ceph: added rook-ceph-default service account
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.

Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
2021-03-29 20:20:40 +05:30
Blaine Gardner 71c41180be test: document or clean up codespell flags
Document codespell flags that are necessary. Remove codespell flags with
minor code changes if possible.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-26 15:00:49 -06:00
Satoru Takeuchi 9eb3160d8b ceph: validate all owner references
Remaining work of #7259

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-26 06:43:54 +00:00
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
Travis Nielsen 64e28af741 ceph: allow flex driver and discovery to be enabled with configmap
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 08:39:26 -06:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Travis Nielsen d93ff428ec ceph: reconcile the mgr services with the active mgr
The active mgr should match the labels on the services that
are available for prometheus and the dashboard. The selector
labels must be updated whenever there is a new active mgr.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:29:16 -07:00
Travis Nielsen 15d0537cec ceph: only raise conditions that represent current state
The conditions on the cephcluster CR were being set to false
when not in progress, which is misleading to the purpose of
conditions. The conditions will now always represent current
state of the cluster. Transient conditions that show progress of
a reconcile will be removed from the status.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-02 14:15:47 -07:00
Sébastien Han 8d033efb5a ceph: silence harmless errors
If the ceph cli outputs an error with "error calling conf_read_file"
this means that the operator has not written its ceph configuration
file. Thus ceph cli commands will fail, so we can just ignore that since
the operator will soon write this file in its initialization sequence.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-03 17:28:44 +01:00
Travis Nielsen 5371c943ca ceph: simplify the log-collector container name
The log-collector container name was exceeding the 63-char limit
if the parent CR name was too long. The container name just needs
to be unique to the pod spec, so we simplify it to remove the parent
name and avoid hitting the limit in the pod spec.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-02-02 12:42:53 -07:00
Sébastien Han c0123cf182 ceph: add cephfs mirroring support
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.

The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050

The same goes for integration tests, they will get added later once we
start testing on Pacific.

Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-28 19:21:18 +01:00
Sébastien Han ebbf332d8d ceph: update rgw and mds deployment for logCollector
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.

Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-22 18:19:07 +01:00
Sébastien Han 758226294b ceph: convert CephClient CRD to controller-runtime
This was the last remaining CRD to not use the controller-runtime
library.
Small additions were added with the transition:

* the Kubernetes Secret that contains the CephX key has now an owner
  reference to the CephClient object
* the secret name is present in the Status field of the CephClient:

```
status:
  info:
    secretName: rook-ceph-client-glance
  phase: Ready
```

The controller will reconcile on CR updates and also if the Kubernetes
Secret is deleted.

Closes: https://github.com/rook/rook/issues/4938
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-15 16:17:38 +01:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Travis Nielsen afeb38417d ceph: suppress reconcile error after operator restart
After the operator restarts, the controllers will not all be able
to reconcile until the ceph config has been generated by the reconcile
of the CephCluster controller. The message printed to the operator log
is frequently seen as an error condition even though it is a normal
condition where we requeue the reconcile until the config is available.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-07 17:24:13 -07:00
Travis Nielsen 156774c459 ceph: remove obsolete topology comments and fix comment typo
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-04 09:10:52 -07:00
Sébastien Han c6a87203ca ceph: add log collector
We can now collect logs directly into a side-car container.
A new CRD spec has been added:

spec:
  logCollector:
    enabled: true
    periodicity: 24h

Every 24h we will rotate log files for each Ceph daemon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-12-01 16:30:17 +01:00
Sébastien Han ad24990473 ceph: ability to abort orchestration
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):

* request for cluster deletion
* request for cluster upgrade

If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.

This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.

Operator log example:

```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
                Image:            "ceph/ceph:v15.2.5",
-               AllowUnsupported: true,
+               AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
-               Image:            "ceph/ceph:v15.2.5",
+               Image:            "ceph/ceph:v15.2.6",
                AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```

Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-25 15:32:04 +01:00
Sébastien Han 70c6752e0e ceph: add dedicated predicate for CephCluster object
Since we want to pass a context to it, let's extract the logic into its
own predicate.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-24 18:21:07 +01:00
Sébastien Han 9b9f45b792 ceph: export isDoNotReconcile function
Since the predicate for the CephCluster object will soon move into its
own predidacte we need to export isDoNotReconcile so that it can be
consummed by the "cluster" package.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-24 18:21:02 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Sébastien Han ea1d71cbfb ceph: add vault kms support for osd encryption
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/

The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:

security:
  kms:
    tokenSecretName: <name of the secret containing a Vault token, used
    to authenticate>
    connectionDetails: < a map of strings containing connection
    information>

Refer to the ceph-cluster-crd documentation to lear more.

Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-30 16:16:33 +01:00
Lalit Maganti 88f16e4980 ceph: ignore MDS_ALL_DOWN during reconciliation
This allows the catch-22 situation where the filesystem cannot be
reconciled because there is no MDS but there is no MDS because the
operator has not reconciled the filesystem and brought up the MDS pods.

Closes #5967, #5846

Signed-off-by: Lalit Maganti <lalitm@google.com>
2020-10-29 01:11:36 +00:00
Santosh Pillai 70c5f70701 ceph: support IPv6 single-stack for ceph
Add --ms-bind-ipv6 true args to mom, osd, mgr, object, mds, rbd daemons.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-10-02 01:03:03 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
Sébastien Han 204ffd4065 Merge pull request #6250 from leseb/mirroring-config
ceph: add rbd-mirror configuration
2020-09-18 09:16:08 +02:00
Sébastien Han 451622a955 ceph: add rbd-mirror configuration
Rook is now capable of configuring mirroring between sites. The
implementation works at different levels:

* CephBlockPool: which introduces a new `mirroring` configuration as well
as `statusCheck`. When turned on, Rook will enable mirroring on the
pool. It will also create a bootstrap peer token and store it in a
Kubernetes Secret. The name of that Secret can be found in the Status
field of the CephBlockPool CRD. This token can be fetched and used by
other clusters to configure the site as a peer. Mirroring can be
configured either at the pool or the image level.

* CephRBDMirror: which introduces a new `peers` configuration allowing
Rook to connect to peers by passing a Secret name. The administrator will
create a Kubernetes Secret with 2 keys: 'token' for the bootstrap peer
token and 'pool' for the name of pool. Once detected the rbd-mirror
controller will go ahead and import the peer configuration.

Pool mirroring status example:

```
status:
  info:
    rbdMirrorBootstrapPeerSecretName: pool-peer-token-test
  mirroringInfo:
    lastChanged: "2020-09-17T14:47:27Z"
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        mode: image
        peers:
        - client_name: client.rbd-mirror-peer
          direction: rx-tx
          mirror_uuid: ""
          site_name: rhcs
          uuid: c50522a4-28a4-4bd3-ba68-e11780308882
        site_name: 91eae0dd-06b1-4d2c-91f3-1311c9df382b-rook-ceph
  mirroringStatus:
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        daemon_health: OK
        health: OK
        image_health: OK
        states:
          replaying: 1
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-17 19:05:04 +02:00
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
Sébastien Han bd90f9afe9 ceph: add more error handling and debug logging
The predicate was lacking from very useful errors as well as logging
messages.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Sébastien Han e7043edbd8 ceph: reconcile if cm is config override
Earlier, we were returning only when the object was not the config
override configmap, we want the opposite.
Also, this was blocking all subsequent conditions and add the correct
check to avoid cm to reconcile.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Travis Nielsen ce92249725 ceph: continue with memory limits below min settings
The operator will now allow the resource limits to be applied below the
recommended minimums. In small clusters, even the recommended minimums
may not be necessary. A warning is still printed to the operator log,
but we allow the configuration to continue.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-19 11:07:16 -06:00
Sébastien Han af1e8c320a ceph: do not log an error if no clusters
If the list of cluster is empty there is no need to report an error in
the logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-05 10:05:51 +02:00
Travis Nielsen 06c21954e1 ceph: labels are immutable for match selectors
The labels cannot be updated on a deployment match selector.
In v1.4 a new label was added for ceph_daemon_type that is intended
to be on the pod labels, but cannot be applied to the match selectors.
Therefore, we suppress any new labels that are added to the daemons.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-04 16:47:00 -06:00
Blaine Gardner dcf8ae2526 ceph: rename pod labels function for more clarity
Rename PodLabels function to CephDaemonAppLabels for more clarity about
what the function's purpose is.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:30 -06:00
Blaine Gardner 9b6114c192 ceph: add 'ceph_daemon_type' label to pod labels
This can help users identify the daemon type similarly to how
'ceph_daemon_id' helps identify the daemon ID. This also makes it so
scripts users create don't have to parse "app=rook-ceph-<daemonType>"
into "<daemonType>" if they desire that bit of info.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:22 -06:00
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30