Cleanup obsolete code where the context no longer needs
to be passed as a parameter to determine the namespaces
to watch.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We don't need to use tini.
We don't have anything in the rook operator that would
either create zombie processes (no threads) or use
exec (to fork). The Go binary has a really good
signal handling mechanism.
Closes: https://github.com/rook/rook/issues/8794
Signed-off-by: Sébastien Han <seb@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters
This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Our admission webhooks will now run as part of the Operator container
and not an additional deployment. This has the advantage of consuming
fewer resources in the cluster and not having to manage affinities and
tolerations. This only drawback is that the Secret containing the
certificates is not mounted anymore and the content needs to be written
inside the Operator. This is not practical since we also need to watch
for the Secret content to change. Meaning that the certificates have
been renewed and the webhook server needs to use them.
A new approach is on its way to hopefully simplify this last issue and
implement a watcher for the Secret.
In the meantime, users need to use the cert-manager or renew
certificates manually. Additionally, they must update the
ValidatingWebhookConfiguration object with the new CA bundle.
Signed-off-by: Sébastien Han <seb@redhat.com>
As of Golang 1.16, we can use `signal.NotifyContext()` from the signal
package. Essentially it combines signal handling for termination as well
as canceling the context to stop any ongoing operations.
Signed-off-by: Sébastien Han <seb@redhat.com>
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This commit moves SetParams and ValidateCSIParam func
into go routine/lock, since it modifies/reads global CSIParam
variable.
Signed-off-by: Rakshith R <rar@redhat.com>
When the operator configmap is edited, it is likely that some user will
simply put "debug" instead of "DEBUG", so let's always transform the
value to upper case to ease user experience.
Signed-off-by: Sébastien Han <seb@redhat.com>
earlier, to change the operator log level we need
to change the deployment which requires a restart
the operator pod.
Now, we are changing the log level dynamically
by reading the configmap. For, backward compatibility
we still load from env var.
Signed-off-by: subhamkrai <srai@redhat.com>
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The operator namespace is not used in the StartOperatorSettingsWatch()
method. Instead, the method looks up the namespace from the
POD_NAMESPACE env var.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Support ceph cluster and CSI on multus deployed in diffrent namespace.
Previously csi was looking for multus config from the cluster deployed
in the namespace in which rook-ceph-operator/csi was deployed.
Now it will look for multus configuration from ceph clusters from all
the namespaces.
Signed-off-by: rohan47 <rohgupta@redhat.com>
as the kubernetes 1.13 is EOL and there is no major
functionalities available in 1.13 (resize,snapshot,clone
metrics etc) we are removing the support for kubernetes
for the same.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
this commit handles all the gosec g601
error code (i.e Implicit memory aliasing
of items from a range statement).
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
remove csi drivers and delete k8s
services that drivers create,when
the default setting of drivers changed
to false.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
adds deploy.sh script to deploy validatingwebhookconfiguration and create secrets.
adds new command ceph admission-controller to start webhook servers.
adds validation for various rook custom resources
Signed-off-by: Vineet Badrinath <vbadrina@redhat.com>
The csi driver does not need to be started before the cluster
is created. When the csi driver fails to load, neither should
it prevent the cluster from being created. Therefore, we
initialize the basic constructs in the csi driver that are required
for cluster creation, then start the csi driver in a goroutine
so the cluster creation can continue. If the csi driver fails
or takes a long time to load, it will no longer affect cluster
creation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator continues to run even though the controller-runtime
failed to start. Since the controller-runtime is essential, the operator
is non-functional after that. It makes troubleshooting more difficult.
Closes: https://github.com/rook/rook/issues/5434
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The controller setting for requeuing an event moved to the
opcontroller package and was no longer needed in the main
controller package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
Checks the version of the configured ceph-csi image while starting the operator. The operator will fail if the image is not supported.
Added an additional parameter to operator to disable the check e.g. to test not yet supported csi images.
Signed-off-by: Stefan Haas <shaas@suse.com>
Prior attempts to delete CSI drivers using owner references to the
Ceph ConfigMap resulted in garbage collection of the driver
resources as described in rook#4590. This patch manually deletes these
resources by name as a stopgap to provide an appropriate user
experience while the team investigates this issue further.
Resolves: rook#4824
Signed-off-by: Elise Gafford <egafford@redhat.com>
When ROOK_CURRENT_NAMESPACE_ONLY is true, watches should be limited to
the configured namespace
Signed-off-by: Andrew DeMaria <lostonamountain@gmail.com>
The operator kit had more utility originally when the operator
was creating and managing the TPRs and CRDs directly. Since
the CRDs are now created from a manifest and no longer by the
operators, the utility of operator kit is limited to the
controller watcher. Since we are moving to the controller runtime
we simplify the code to make the transition smoother. Now
there is only a simple WatchCR method that will need to be
replaced as we maek that transition.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We create CSI resources when we create the first Ceph cluster.
This change deletes all CSI resources when we delete the last Ceph
cluster.
Partially resolves: #4234
Signed-off-by: Elise Gafford <egafford@redhat.com>
In order to remove Ceph-CSI resources after deletion of the last
Ceph cluster, the operator must be able to perform post-processing
actions on cluster deletion. This change creates a hook for the
addition of post-deletion operator callback functions.
Partially resolves: #4234
Signed-off-by: Elise Gafford <egafford@redhat.com>
**Description of your changes:**
This modification adds the information extracted from 'ceph-volume inventory':
command to the device configmaps generated by the discovery daemon when
"rook discover" starts with the new boolean "--use-ceph-volume" parameter.
Resolves #
https://github.com/rook/rook/issues/2606
Now the <cephVolumeData> field contains all the information returned
from <ceph-volume inventory> command.
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
The decision to start the device discovery daemonset is made
by an env var on the operator pod since in some scenarios
the discovery needs to be started immediately with the operator
instead of being delayed to start with the cluster.
Signed-off-by: travisn <tnielsen@redhat.com>
Deployment behaves better when a node gets disconnected
from the rest of the cluster - new provisioner leader
is elected in ~15 seconds, while it may take up to
5 minutes for StatefulSet to start a new replica.
if kube version is 1.13.x deploy provisioner as statefulset.
if kube version is higher than 1.14+ deploy provisioner
as deployment.
Refer: kubernetes-csi/external-provisioner@52d1fbc
Refer: ceph/ceph-csi#497
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
When configuring an external cluster the orchestration of discover and all the ceph
daemons will be skipped. The operator will still watch for the creation of
crds for this namespace. When a filesystem, object, object user, or
ganesha crd are created, the operator will simply print an error
to the log.
Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
Previously, if csi was not supported by the k8s version the csi set-up
code was skipped but csi was "left on". This change ensures that if
csi is not supported the csi enablement flags are set to false so
that other code that needs csi support will not run.
Signed-off-by: John Mulligan <jmulligan@redhat.com>
Avoid using global values that are not used anywhere in the code.
Split the object used to express the ceph csi templates into two
so only values that are set globally need to be managed globally,
all other values can be stored in the temporary local struct.
Signed-off-by: John Mulligan <jmulligan@redhat.com>