Commit Graph
205 Commits
Author SHA1 Message Date
Madhu Rajanna 123025f22c csi: update csi-addons to v0.9.0
As we have new csi-addons v0.9.0
updating the same here as well.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-08-16 07:57:16 +02:00
subhamkrai 4b0b3a55d9 csi: add code for new CSI operator CR cephcluster
adding code changes,rbac changes required for create the new
Ceph-CSI operator CR named cephCluster in api group 'csi.ceph.io'.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-08-08 11:21:45 +05:30
subhamkrai 9f1b64be54 mds: block of active mds ip only
currently, it was blocking both the mds ip
instead its should block ip of of mds which is in
the same node which is down.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-07-17 13:30:35 +05:30
subhamkrai d429ed8be4 build: update controller runtime to v0.18.4
this commit update cntrl runtime to v0.18.4 and other related deps/

Signed-off-by: subhamkrai <srai@redhat.com>
2024-07-05 09:05:12 +05:30
subhamkrai 5bff860380 core: remove namespace/ownerRef from networkFence
since networkFence is a cluster-based resource so that we don't need
the namespace and ownerReferences as it cause garbage-collector errors.
Also, now we create the networkFence with clusteUID label so when doing
cleanup we match the cephCluster uid and networkFence label clusterUID.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-02-13 21:47:04 +05:30
subhamkrai fb39580067 operator: move most of discover pod setting to cm
It's better to move most of discover daemon setting
from env to configmap rook-ceph-operator-config.
Although, we are moving to configmap, we keep reading settings
from env but the priority will be configmap settings.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-16 22:17:47 +05:30
subhamkrai 39b5c057ce core: faster recovery from rbd rwo node loss
in the existing node watcher, we'll check for node update
event and see if there are `out-of-service` taints are applied
and `ROOK_WATCH_FOR_NODE_FAILURE` is enabled in rook-ceph-operator-configmap,
if then we'll create the networkFence cr and delete the cr if nodes come back.
And, added the unit test too.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-07 21:16:25 +05:30
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
Travis Nielsen 0ede2946a8 core: remove finalizer from cluster cr last
Removing the finalizers from cluster resources can intermittently fail
in the CI if other updates are made at similar times. The controller
runtime will retry after the failure, but if something external
removes the finalizer on the cephcluster CR, the deletion will not
continue with the removal of finalizers from the other configmap
and secret critical resources.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-10-26 14:17:03 -06:00
parth-gr 26584fc6e5 core: update loadclusterInfo with multus check
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor

Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-22 14:46:51 +05:30
Rakshith R f21c12234d osd: add kmip encryption support
The cluster-wide encryption feature now
supports KMIP (Key Management Interoperability Protocol)
kms.
KMIP is an extensible communication protocol that defines
message formats for the manipulation of cryptographic keys
on a key management server.

For more information, refer:
https://en.wikipedia.org/wiki/Key_Management_Interoperability_Protocol

Signed-off-by: Rakshith R <rar@redhat.com>
2022-09-09 11:39:33 +05:30
Jiffin Tony Thottan 5c8ca01bd0 object: adding support for sse s3 for RGW
The RGW support server side encryption with help of s3 protocol, till
now the `sse:kms` was support in which keys will be provided by the user
and but it will be saved in external management service like vault. Now
the support for `sse:s3` is added so the entire encryption key
management is performed by RGW itsels.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-07-28 11:50:48 +05:30
Sébastien Han a6036074db core: more checkpoint for context cancelled
Some context canceled were missed and not reported correctly. We had
multiple reconcilers running in //. Since they all run in goroutine then
will continue to run until they timeout (if timeout is available,
    otherwise would run forever until the op is restarted).

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-07-12 12:07:59 +02:00
Sébastien Han 05506e7a68 core: reload go routine after CR is edited
Previously, the struct maintaining the list of cluster was still
initialized with a cluster item. Then the monitoring check will see that
the cluster is part of the struct already and thus won't run the
monitoring go routine again.
Now each time we cancel the context, we also remove the cluster item
from the map so that when the controller runs again, the monitoring
struct is re-populated and the go routine runs and statuses are updated.

Closes: https://github.com/rook/rook/issues/9911
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-03-31 08:49:42 +02:00
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
parth-gr 2dfd64a97c core: add observedGeneration to CR status
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller

Closes: https://github.com/rook/rook/issues/9673

Signed-off-by: parth-gr <paarora@redhat.com>
2022-03-16 19:58:35 +05:30
Blaine Gardner c92c6fbc36 core: rework usage of ReportReconcileResult
ReportReconcileResult should never be given an object that is nil. It is
impossible to force compile-time checking for this because the object's
type is an interface. We can't force this, but we can rework the
function definition to handle more cases where the object may not be
returned fully-complete, and we can rework callers to return structs
(not pointers-to-structs) so it is less likely a caller will pass nil
to the function.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-03-10 16:53:25 -07:00
Sébastien Han 739abf6662 csi: change rook-ceph-csi-config to expose clusterID for subvolume
The rook-ceph-csi-config config map now includes CephFS subvolumegroups
when a subvolumegroup CR is created. It is useful for ceph-csi to
understand which subvolumegroup to use for a given cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-21 10:12:56 +01:00
Sébastien Han a7acf490b6 osd: add ibm key protect kms
The cluster-wide encryption feature now has a new supported key
management system: IBM Key protect. More information about the backend
can be found in Rook documentation under the

Today's implementation stores OSD encryption keys has "Standard" keys.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-13 16:02:10 +01:00
Sébastien Han fffb862956 Merge pull request #9457 from leseb/fix-9452
core: disallow multiple clusters in the same namespace
2021-12-20 17:18:10 +01:00
Sébastien Han da76a772f8 core: disallow multiple clusters in the same namespace
Rook does not support running multiple clusters in the same namespace,
so the operator should not reconcile if a new cluster is added.
The scenario where a cluster is added while the operator is down is also
handled. CR updates are also handled.
When the operator detects more than one cluster it will refuse to
reconcile the CephCluster, and child CRDs will block too until the
operator is ready.
The user must remove one of the clusters before can continue to perform
any reconcile.

Closes: https://github.com/rook/rook/issues/9452
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-17 10:56:15 +01:00
Blaine Gardner da61ac1ae8 operator: always report events
The original PR which added event reporting unnecessarily "optimized" to
prevent spamming the API controller[1].

It is sometimes important to get events as they happen and not hide new
events behind preexisting older events. For example, in integration
tests, we may often want to wait for a controller to finish processing
an update, and the best way to do that is to wait for the
"ReconcileSucceeded" event. In order for this to be useful, the events
must be reported each time.

If we begin having problems with events being reported too often, then
we should fix the underlying issue of reconciles happening too often
instead of relying on a time-based "optimization" that hides recent
event reports that may be useful.

[1]: https://github.com/rook/rook/pull/7222

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-15 15:24:17 -07:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
parth-gr 7c99858a77 ceph: add finalizers to rook-ceph-mon secrets and configmap
Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery

Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
2021-10-07 19:44:17 +00:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 2d55e69416 ceph: move scheme initialization to the same place
Let's initialize the schemes in a single place instead of doing it
when each controller initializes.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-06 11:33:10 +02:00
Sébastien Han 630c2f6a8b ceph: add an rbd-mirror bootstrap token on cluster creation
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-02 18:34:27 +02:00
Sébastien Han 65aea09e7d ceph: small bootstrap sequence refactor
Moving and factoring code here and there to hopefully make the
initialization sequence easier to go by.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-30 15:03:25 +02:00
Travis Nielsen c39c1c7ddf ceph: retry reconcile immediately after cancellation
If the reconcile is cancelled due to a CR update, we want to retry the next
reconcile immediately rather than wait for the exponential backoff timeout
if the reconcile was already failing. The wait can easily be minutes
if the reconcile was in this state, which makes it appear the operator
is ignoring the request to start a new reconcile.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-02 08:23:25 -06:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Blaine Gardner a5c29a3847 ceph: handle CephCluster stopCleanupCh better
Investigate a possible memory leak around CephCluster cleanup's
stopCleanupCh channel. Add more close() statements and comment that Go's
garbage collector is able to close the channel when there are no more
referents to the channel.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-10 10:50:30 -06:00
Blaine Gardner 21e290e003 ceph: implement dependencies for CephCluster
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-10 10:06:10 -06:00
Nitin Goyal 25d2df4018 ceph: generate cephCluster reconcile result as events
Generate cephCluster reconcile result as an k8s event if it is not
already present on the cluster

Signed-off-by: Nitin Goyal <nigoyal@redhat.com>
2021-05-18 18:27:52 +05:30
Sébastien Han 5e1e9f4c92 ceph: actively update the service endpoint for external mgr
If the cluster is external we want to periodically rehydrate the mgr
endpoint. This handles the scenarion where the active manager changes,
 so we need to update the endpoint with the new IP address.
The create-external-cluster-resources.py script now requires an extra
permission to query the manager service so Rook can discover the active
one and its IP.

Testing:

```
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 37,
    "available": true,
    "active_name": "b",
    "num_standby": 1
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.12:9283   3m10s

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] k scale --replicas=0 deployment rook-ceph-mgr-b
deployment.apps/rook-ceph-mgr-b scaled

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 40,
    "available": true,
    "active_name": "a",
    "num_standby": 0
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.13:9283   3m55s
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-11 16:08:40 +02:00
Lars Lehtonen 600a1498fe ceph: fix multiple imports
This fixes double imports within and beneath pkg/operator/ceph/cluster.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-14 01:51:35 -07:00
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
Travis Nielsen 64e28af741 ceph: allow flex driver and discovery to be enabled with configmap
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 08:39:26 -06:00
Travis Nielsen 4d969a406f Merge pull request #7414 from sp98/skip-cluster-cleanup
ceph: skip cleanup if cluster is not configured correctly.
2021-03-16 08:24:24 -06:00
Santosh Pillai 7bd05d0d7c ceph: skip cleanup if cluster is not configured correctly
deletion of cluster is stuck when cluster is not configured. The PR skips the cleanup and removes the finalizers if the cluster was not configured, that is, if clusterInfo could not be loaded.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-03-16 18:59:43 +05:30
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Travis Nielsen 15d0537cec ceph: only raise conditions that represent current state
The conditions on the cephcluster CR were being set to false
when not in progress, which is misleading to the purpose of
conditions. The conditions will now always represent current
state of the cluster. Transient conditions that show progress of
a reconcile will be removed from the status.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-02 14:15:47 -07:00
Travis Nielsen 4b43e7afa5 ceph: allow removal of arbitrary osds on pvcs and simplify pvc names
OSDs on PVCs previously were always replaced with PVCs of a given name.
For on-prem scenarios where the OSDs might need to be removed instead
of replaced, the operator now allows holes to exist in the PVC index
names as long as there are a sufficient number of PVCs to meet the
criteria for the deviceSet.count.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-26 14:38:12 -07:00
Sébastien Han ebbf332d8d ceph: update rgw and mds deployment for logCollector
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.

Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-22 18:19:07 +01:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Sébastien Han ad24990473 ceph: ability to abort orchestration
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):

* request for cluster deletion
* request for cluster upgrade

If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.

This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.

Operator log example:

```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
                Image:            "ceph/ceph:v15.2.5",
-               AllowUnsupported: true,
+               AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
-               Image:            "ceph/ceph:v15.2.5",
+               Image:            "ceph/ceph:v15.2.6",
                AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```

Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-25 15:32:04 +01:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Sébastien Han ea1d71cbfb ceph: add vault kms support for osd encryption
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/

The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:

security:
  kms:
    tokenSecretName: <name of the secret containing a Vault token, used
    to authenticate>
    connectionDetails: < a map of strings containing connection
    information>

Refer to the ceph-cluster-crd documentation to lear more.

Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-30 16:16:33 +01:00
Raghavendra TalurandTravis Nielsen 0c21abd535 ceph: prevent closing of channel more than once
Once we close the channel for the monitoring daemons, we update the
monitoring status so as to not call them again.

This is to prevent "close of closed channel" panic.

Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Nitin Goyal <nigoyal@redhat.com>
Signed-off-by: Raghavendra Talur <raghavendra.talur@gmail.com>
2020-10-05 22:14:16 +05:30