Commit Graph
63 Commits
Author SHA1 Message Date
Travis Nielsen 82c82bbe13 core: remove unnecessary param to watch namespaces
Cleanup obsolete code where the context no longer needs
to be passed as a parameter to determine the namespaces
to watch.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-01-08 16:02:27 -07:00
Sébastien Han 121c2987e3 ceph: stop using tini
We don't need to use tini.
We don't have anything in the rook operator that would
either create zombie processes (no threads) or use
exec (to fork). The Go binary has a really good
signal handling mechanism.

Closes: https://github.com/rook/rook/issues/8794
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 10:54:16 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
Sébastien Han 656dd0f334 ceph: move the admission webhook to the operator
Our admission webhooks will now run as part of the Operator container
and not an additional deployment. This has the advantage of consuming
fewer resources in the cluster and not having to manage affinities and
tolerations. This only drawback is that the Secret containing the
certificates is not mounted anymore and the content needs to be written
inside the Operator. This is not practical since we also need to watch
for the Secret content to change. Meaning that the certificates have
been renewed and the webhook server needs to use them.
A new approach is on its way to hopefully simplify this last issue and
implement a watcher for the Secret.
In the meantime, users need to use the cert-manager or renew
certificates manually. Additionally, they must update the
ValidatingWebhookConfiguration object with the new CA bundle.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-24 19:07:04 +02:00
Sébastien Han 4e2879a154 ceph: signal signals handling with context
As of Golang 1.16, we can use `signal.NotifyContext()` from the signal
package. Essentially it combines signal handling for termination as well
as canceling the context to stop any ongoing operations.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-02 14:19:50 +02:00
Satoru Takeuchi 30e4fbb01f ceph: make the timeout of ceph commands cofigurable
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-27 12:42:16 +00:00
Rakshith R 66c9ccab86 ceph: read and validate CSI params in Go routine
This commit moves SetParams and ValidateCSIParam func
into go routine/lock, since it modifies/reads global CSIParam
variable.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-06-21 15:44:54 +05:30
Sébastien Han 7e7be7ea56 ceph: force upper case for Operator log level
When the operator configmap is edited, it is likely that some user will
simply put "debug" instead of "DEBUG", so let's always transform the
value to upper case to ease user experience.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-01 10:51:44 +02:00
subhamkrai 0d17dede9c ceph: change operator log level dynamically
earlier, to change the operator log level we need
to change the deployment which requires a restart
the operator pod.

Now, we are changing the log level dynamically
by reading the configmap. For, backward compatibility
we still load from env var.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-05-27 12:40:50 +05:30
Travis Nielsen 64e28af741 ceph: allow flex driver and discovery to be enabled with configmap
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 08:39:26 -06:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Satoru Takeuchi 1be47ea0b8 ceph: delete discovery daemon if it is disabled
discovery-daemon still exists even if it's disabled.

Closes: https://github.com/rook/rook/issues/6936

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-01-15 19:44:40 +00:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Travis Nielsen e56c68c9e9 ceph: remove unused operator namespace parameter
The operator namespace is not used in the StartOperatorSettingsWatch()
method. Instead, the method looks up the namespace from the
POD_NAMESPACE env var.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-18 11:45:14 -07:00
Arun Kumar Mohan ded16f779d ceph: changes for 'sigs.k8s.io/sig-storage-lib-external-provisioner/v6'
Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:03 +05:30
rohan47 cb9f947c3b ceph: support ceph cluster and CSI on multus in diffrent namespace
Support ceph cluster and CSI on multus deployed in diffrent namespace.
Previously csi was looking for multus config from the cluster deployed
in the namespace in which rook-ceph-operator/csi was deployed.
Now it will look for multus configuration from ceph clusters from all
the namespaces.

Signed-off-by: rohan47 <rohgupta@redhat.com>
2020-11-16 20:47:09 +05:30
rohan47 8c4ede1edc ceph: added support for multus for csi
CSI pods now utilize multus networking and connect to public
network specified in the CephCluster CR.

Closes: https://github.com/rook/rook/issues/5356
Signed-off-by: rohan47 <rohgupta@redhat.com>
2020-08-20 18:25:09 +05:30
Madhu Rajanna 1989e0d8c5 ceph: remove csi support for kubernetes 1.13
as the kubernetes 1.13 is EOL and there is no major
functionalities available in 1.13 (resize,snapshot,clone
metrics etc) we are removing the support for kubernetes
for the same.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-08-05 21:19:59 +05:30
subhamkrai 465a0f0aec ceph: handling gosec error code g601
this commit handles all the gosec g601
error code (i.e Implicit memory aliasing
of items from a range statement).

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-28 11:34:49 +05:30
subhamkrai 7f9f1690d2 ceph: remove csi drivers when disable
remove csi drivers and delete k8s
services that drivers create,when
the default setting of drivers changed
to false.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-06 13:33:52 +05:30
Vineet Badrinath ce1003aef8 ceph: adds scripts and components to support admission controllers
adds deploy.sh script to deploy validatingwebhookconfiguration and create secrets.
adds new command ceph admission-controller to start webhook servers.
adds validation for various rook custom resources

Signed-off-by: Vineet Badrinath <vbadrina@redhat.com>
2020-06-24 14:59:00 +05:30
Travis Nielsen 6c249e751b ceph: start csi driver in parallel of cluster
The csi driver does not need to be started before the cluster
is created. When the csi driver fails to load, neither should
it prevent the cluster from being created. Therefore, we
initialize the basic constructs in the csi driver that are required
for cluster creation, then start the csi driver in a goroutine
so the cluster creation can continue. If the csi driver fails
or takes a long time to load, it will no longer affect cluster
creation.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-29 16:42:22 -06:00
Satoru Takeuchi f6a3da2760 ceph: gave up to run the operator if the controller-runtime failed to start
The operator continues to run even though the controller-runtime
failed to start. Since the controller-runtime is essential, the operator
is non-functional after that. It makes troubleshooting more difficult.

Closes: https://github.com/rook/rook/issues/5434

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-05-18 17:38:35 +00:00
Travis Nielsen f47bb945c2 ceph: remove duplicate controller wait setting
The controller setting for requeuing an event moved to the
opcontroller package and was no longer needed in the main
controller package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Madhu Rajanna 81688398f2 cleanup: use err.Wrap when the formatting is not required
Replaced err.Wrapf with err.Wrap when the formatting
is not required.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-29 17:42:59 +05:30
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Umanga Chapagain 0e932c15eb Ceph: add CSI configurations to ConfigMap
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-27 15:17:30 +05:30
Sébastien Han ca0a30f38d ceph: convert Filesystem controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 23:34:58 +01:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Umanga Chapagain c54c0d16cf Ceph: adds watch to rook-ceph-operator-config ConfigMap
watches the rook-ceph-operator-config ConfigMap and updates
CSI driver

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-10 12:46:54 +05:30
Sébastien Han 9f2867e12a ceph: separate controller for CephObjectStoreUser CRD
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-06 11:53:40 +01:00
Sébastien Han a3068dee0b ceph: separate controller for CephBlockPool CRD
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-02 17:31:03 +01:00
Stefan Haas b3cc4aeb44 ceph: ceph-csi version detection #3824
Checks the version of the configured ceph-csi image while starting the operator. The operator will fail if the image is not supported.
Added an additional parameter to operator to disable the check e.g. to test not yet supported csi images.

Signed-off-by: Stefan Haas <shaas@suse.com>
2020-02-28 14:14:40 +01:00
Elise Gafford de2c409d79 ceph: delete CSI drivers with deletion of last cluster
Prior attempts to delete CSI drivers using owner references to the
Ceph ConfigMap resulted in garbage collection of the driver
resources as described in rook#4590. This patch manually deletes these
resources by name as a stopgap to provide an appropriate user
experience while the team investigates this issue further.

Resolves: rook#4824
Signed-off-by: Elise Gafford <egafford@redhat.com>
2020-02-27 19:06:23 -05:00
Andrew DeMaria 60ec9c35fa ceph: limit watches to configured namespace
When ROOK_CURRENT_NAMESPACE_ONLY is true, watches should be limited to
the configured namespace

Signed-off-by: Andrew DeMaria <lostonamountain@gmail.com>
2020-02-14 18:34:57 -07:00
Travis Nielsen d0c7d28a63 build: remove operator kit dependency
The operator kit had more utility originally when the operator
was creating and managing the TPRs and CRDs directly. Since
the CRDs are now created from a manifest and no longer by the
operators, the utility of operator kit is limited to the
controller watcher. Since we are moving to the controller runtime
we simplify the code to make the transition smoother. Now
there is only a simple WatchCR method that will need to be
replaced as we maek that transition.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-07 08:30:14 -07:00
Sébastien Han 5ce2ed220e ceph: use "github.com/pkg/errors"
We now use the error package.
Kubernetes errors have been renamed kerrors since they are lower than
'errors'.

Closes: https://github.com/rook/rook/issues/4054
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-09 16:58:32 +01:00
Elise Gafford 32e5d09336 ceph: remove CSI resources on deletion of last cluster
We create CSI resources when we create the first Ceph cluster.
This change deletes all CSI resources when we delete the last Ceph
cluster.

Partially resolves: #4234
Signed-off-by: Elise Gafford <egafford@redhat.com>
2019-11-22 10:33:39 -05:00
Elise Gafford 3adcb67755 ceph: add removeCallbacks hook to Operator and ClusterController
In order to remove Ceph-CSI resources after deletion of the last
Ceph cluster, the operator must be able to perform post-processing
actions on cluster deletion. This change creates a hook for the
addition of post-deletion operator callback functions.

Partially resolves: #4234
Signed-off-by: Elise Gafford <egafford@redhat.com>
2019-11-21 11:11:25 -05:00
Juan Miguel Olmo Martínez 7c942604f6 ceph: Get <ceph-volume inventory> data in dev. configmaps
**Description of your changes:**
This modification adds the information extracted from 'ceph-volume inventory':
command to the device configmaps generated by the discovery daemon when
"rook discover" starts with the new boolean "--use-ceph-volume" parameter.

Resolves #
https://github.com/rook/rook/issues/2606

Now the <cephVolumeData> field contains all the information returned
from <ceph-volume inventory> command.

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2019-11-06 10:22:56 +01:00
Rohan CJ 4072d14557 Add tests for the clusterdisruption controller and update the docs and
pending releas notes for managed PDBs.

Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-08-30 14:04:23 +05:30
Rohan CJ 079c24883b Add a controller-runtime scaffolding
Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-08-30 08:59:53 +05:30
Sébastien Han 50e8fced8c Merge pull request #3698 from travisn/skip-start-discovery
Start the ceph device discovery only based on env var
2019-08-26 22:27:22 +02:00
travisn f3e9a5adac ceph: starting the device discovery only based on env var
The decision to start the device discovery daemonset is made
by an env var on the operator pod since in some scenarios
the discovery needs to be started immediately with the operator
instead of being delayed to start with the cluster.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-08-26 11:20:38 -06:00
Madhu Rajanna 78fb35a93d Use Deployment with leader election instead of StatefulSet
Deployment behaves better when a node gets disconnected
from the rest of the cluster - new provisioner leader
is elected in ~15 seconds, while it may take up to
5 minutes for StatefulSet to start a new replica.

if kube version is 1.13.x deploy provisioner as statefulset.
if kube version is higher than 1.14+ deploy provisioner
as deployment.

Refer: kubernetes-csi/external-provisioner@52d1fbc
Refer: ceph/ceph-csi#497

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-08-26 19:29:37 +05:30
travisn 7ab689ee8e ceph: skip local setup for an external ceph cluster
When configuring an external cluster the orchestration of discover and all the ceph
daemons will be skipped. The operator will still watch for the creation of
crds for this namespace. When a filesystem, object, object user, or
ganesha crd are created, the operator will simply print an error
to the log.

Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-08-22 17:19:32 +02:00
John Mulligan 521680d100 ceph csi: disable csi if server version is too low
Previously, if csi was not supported by the k8s version the csi set-up
code was skipped but csi was "left on". This change ensures that if
csi is not supported the csi enablement flags are set to false so
that other code that needs csi support will not run.

Signed-off-by: John Mulligan <jmulligan@redhat.com>
2019-08-13 06:42:22 -04:00
John Mulligan 2b1ae54b84 ceph csi: clean up how templates are expressed
Avoid using global values that are not used anywhere in the code.
Split the object used to express the ceph csi templates into two
so only values that are set globally need to be managed globally,
all other values can be stored in the temporary local struct.

Signed-off-by: John Mulligan <jmulligan@redhat.com>
2019-08-12 16:41:26 -04:00