This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery
Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.
Signed-off-by: Sébastien Han <seb@redhat.com>
If the reconcile is cancelled due to a CR update, we want to retry the next
reconcile immediately rather than wait for the exponential backoff timeout
if the reconcile was already failing. The wait can easily be minutes
if the reconcile was in this state, which makes it appear the operator
is ignoring the request to start a new reconcile.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Investigate a possible memory leak around CephCluster cleanup's
stopCleanupCh channel. Add more close() statements and comment that Go's
garbage collector is able to close the channel when there are no more
referents to the channel.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
If the cluster is external we want to periodically rehydrate the mgr
endpoint. This handles the scenarion where the active manager changes,
so we need to update the endpoint with the new IP address.
The create-external-cluster-resources.py script now requires an extra
permission to query the manager service so Rook can discover the active
one and its IP.
Testing:
```
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
"epoch": 37,
"available": true,
"active_name": "b",
"num_standby": 1
}
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME ENDPOINTS AGE
rook-ceph-mgr-external 172.17.0.12:9283 3m10s
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] k scale --replicas=0 deployment rook-ceph-mgr-b
deployment.apps/rook-ceph-mgr-b scaled
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
"epoch": 40,
"available": true,
"active_name": "a",
"num_standby": 0
}
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME ENDPOINTS AGE
rook-ceph-mgr-external 172.17.0.13:9283 3m55s
```
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
deletion of cluster is stuck when cluster is not configured. The PR skips the cleanup and removes the finalizers if the cluster was not configured, that is, if clusterInfo could not be loaded.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The conditions on the cephcluster CR were being set to false
when not in progress, which is misleading to the purpose of
conditions. The conditions will now always represent current
state of the cluster. Transient conditions that show progress of
a reconcile will be removed from the status.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
OSDs on PVCs previously were always replaced with PVCs of a given name.
For on-prem scenarios where the OSDs might need to be removed instead
of replaced, the operator now allows holes to exist in the PVC index
names as long as there are a sufficient number of PVCs to meet the
criteria for the deviceSet.count.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.
Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):
* request for cluster deletion
* request for cluster upgrade
If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.
This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.
Operator log example:
```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff= v1.ClusterSpec{
CephVersion: v1.CephVersionSpec{
Image: "ceph/ceph:v15.2.5",
- AllowUnsupported: true,
+ AllowUnsupported: false,
},
DriveGroups: nil,
Storage: {UseAllNodes: true, Selection: {UseAllDevices: &true}},
... // 20 identical fields
}
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff= v1.ClusterSpec{
CephVersion: v1.CephVersionSpec{
- Image: "ceph/ceph:v15.2.5",
+ Image: "ceph/ceph:v15.2.6",
AllowUnsupported: false,
},
DriveGroups: nil,
Storage: {UseAllNodes: true, Selection: {UseAllDevices: &true}},
... // 20 identical fields
}
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```
Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/
The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:
security:
kms:
tokenSecretName: <name of the secret containing a Vault token, used
to authenticate>
connectionDetails: < a map of strings containing connection
information>
Refer to the ceph-cluster-crd documentation to lear more.
Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
Once we close the channel for the monitoring daemons, we update the
monitoring status so as to not call them again.
This is to prevent "close of closed channel" panic.
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Nitin Goyal <nigoyal@redhat.com>
Signed-off-by: Raghavendra Talur <raghavendra.talur@gmail.com>
If AllowUninstallWithVolumes is set in the cleanup policy then we don't
check and wait for the deletion of existing volumes from the
cephCluster.
Co-authored-by: Nitin Goyal <nigoyal@redhat.com>
Signed-off-by: Raghavendra Talur <raghavendra.talur@gmail.com>
Now, we can not only cleanup monitor data, logs and crashes but the
disks too. As part of the cleanupPolicy CR spec, we have a new setting
called sanitizeDisks which holds more details:
* confirmation: the confirmation message to sanitize disks,
use "yes-really-sanitize-disks" to confirm
This will **only** wipe the metadata, so it's a fast cleanup allowing
you to re-install later but won't remove all the data from the drive.
Signed-off-by: Sébastien Han <seb@redhat.com>
During cluster delection, the finalizer will not be removed until
all the PVCs are removed. Since this could wait for a long time,
as long as the deletion event is re-queued, we need to cancel
the goroutine that is watching for cluster deletion so it can
avoid running multiple goroutines as the event runs again
and again.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When a cluster is requested for deletion there is no longer
a reason to monitor the cluster health such as whether mons
are in quorum, osds are up and in, or monitor the ceph status.
The cluster will wait potentially for a long time if the
PVCs are not deleted by the admin, therefore we stop monitoring
the cluster even before the finalizer is removed.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The Ceph version information is not needed to be passed
to the osd health checker since mimic support was dropped.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The monitoring goroutines that check for mon and osd health,
the ceph cluster status, and the bucket provisioner should
only be started once for each cluster. Previous to this change
and since the controller runtime conversion the goroutines were
being started again every time a reconcile was run. This means
you could potentially have many goroutines monitoring the same
thing.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
During cluster deletion, we currently only retry for a couple minutes
to wait for the pvcs to be deleted. After the timeout, we proceed
with the cluster deletion. To properly protect the pvcs for proper
cleanup, the finalizer should not be removed until the pvcs
are all confirmed to be deleted. In order to not block other cluster
events, we re-queue the deletion event to run again every 10s
until the pvcs are deleted or the finalizer is manually removed.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If the OSD pod is portable and is on a failed node, the node
will be stuck indefinitely in the terminating state. The pod
can be force deleted in this scenario in order to allow the
underlying volume to be detached and then mounted on another
node.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
Currently, we need to configure the Ceph external Admin keyi
in the Rook deployment to be able to connect to an external Ceph cluster.
If we wanted to run in a multi-tenant fashion
were several K8s/Rook clusters wanted to connect to the same external ceph cluster,
each K8s deployment would have the access to the External Ceph Admin key
and could potentially access or delete the Data
from pools that belong to other k8s/Rook Clusters.
Now the admin key is optional but the helper script create-external-cluster-resources.sh
will help create the necessary keys/users to connect to that cluster.
Closes: https://github.com/rook/rook/issues/4917 and https://github.com/rook/rook/pull/5227
Signed-off-by: Sébastien Han <seb@redhat.com>
You can now use Rook along with Multus. Multus must be up and running
and the right ressources must exist such as NetworkAttachmentDefinition
CR.
The Cluster CR spec has new fields to work with multus:
network:
provider: multus (or 'host' for hostNetworking)
selectors:
public: NetworkAttachmentDefinition name
cluster: NetworkAttachmentDefinition name
If only a single NetworkAttachmentDefinition is provided Rook will use
both anyway for the Ceph traffic.
Please refer to the doc to learn more.
Closes: https://github.com/rook/rook/issues/4716
Signed-off-by: Sébastien Han <seb@redhat.com>
In case of multiple clusters, we don't want to delete all the mon directories under the dataDirHostPath during cluster cleanup.
This PR deletes the mon directory only if the montior secret key matches.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
In order to perform full automation of cephcluster deletion safely
and protect against catastrophic data loss, the end user must be
able to signify that they intend to irrecoverably delete the data
in their cluster. The cleanupPolicy field of the cluster spec is
intended to communicate this. This patch does not implement
automated deletion, but only creates the spec field and the
safety feature of halting orchestration other than deletion on a
cluster with a set cleanup policy value.
Partially-fixes: #3222
Signed-off-by: Elise Gafford <egafford@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4941
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>