Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery
Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
During cluster deletion, we currently only retry for a couple minutes
to wait for the pvcs to be deleted. After the timeout, we proceed
with the cluster deletion. To properly protect the pvcs for proper
cleanup, the finalizer should not be removed until the pvcs
are all confirmed to be deleted. In order to not block other cluster
events, we re-queue the deletion event to run again every 10s
until the pvcs are deleted or the finalizer is manually removed.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
Currently, we need to configure the Ceph external Admin keyi
in the Rook deployment to be able to connect to an external Ceph cluster.
If we wanted to run in a multi-tenant fashion
were several K8s/Rook clusters wanted to connect to the same external ceph cluster,
each K8s deployment would have the access to the External Ceph Admin key
and could potentially access or delete the Data
from pools that belong to other k8s/Rook Clusters.
Now the admin key is optional but the helper script create-external-cluster-resources.sh
will help create the necessary keys/users to connect to that cluster.
Closes: https://github.com/rook/rook/issues/4917 and https://github.com/rook/rook/pull/5227
Signed-off-by: Sébastien Han <seb@redhat.com>
Prior attempts to delete CSI drivers using owner references to the
Ceph ConfigMap resulted in garbage collection of the driver
resources as described in rook#4590. This patch manually deletes these
resources by name as a stopgap to provide an appropriate user
experience while the team investigates this issue further.
Resolves: rook#4824
Signed-off-by: Elise Gafford <egafford@redhat.com>
The rook types used across the storage providers moved from the v1alpha2
package to the v1 package. This commit points the packages at their new
location. Implementation is expected to remain unchanged.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is to handle all those unhandled errors which raises the gosec warning.
Fixed G104: Unhandled Errors are handled now
Signed-off-by: Nizamudeen <nia@redhat.com>
The ceph/ceph:v14.2.7 image is released so we can pick these fixes
up as the recommended version of ceph.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Fixed conditions getting resetted after the operator restart.
Did the changes which required to implement conditions on the rook ceph cluster
Conditions will eliminate the current status.State and incorporates a type which
provides much more description to the current status of the cluster.
Signed-off-by: Nizamudeen <nia@redhat.com>
In order to remove Ceph-CSI resources after deletion of the last
Ceph cluster, the operator must be able to perform post-processing
actions on cluster deletion. This change creates a hook for the
addition of post-deletion operator callback functions.
Partially resolves: #4234
Signed-off-by: Elise Gafford <egafford@redhat.com>
We now differentiate the cases where:
* we only consume the external cluster
* we consume the external cluster as well as creating stateless
resources in Kubernetes (bootstrap mds,rgw, nfs)
This is mostly controlled via the image property spec. If not defined,
not extra CRs won't be able to be created.
Now the external cluster feature supports Ceph cluster as of Luminous 12.2.
Signed-off-by: Sébastien Han <seb@redhat.com>
When the flex driver is disabled, the check for
volume attachments will fail and the finalizer
is never removed. To avoid this, just log the
failure to list volumes and remove the
finalizer anyway.
Resolves#3912
Signed-off-by: Kristoffer Grönlund <kgronlund@suse.com>
The last image has disabled ephemeral repositories so it's now possible
for images older than 15 days to install packages without having an
error from non-existing repositories.
Closes: https://github.com/rook/rook/issues/3662
Signed-off-by: Sébastien Han <seb@redhat.com>
When configuring an external cluster the orchestration of discover and all the ceph
daemons will be skipped. The operator will still watch for the creation of
crds for this namespace. When a filesystem, object, object user, or
ganesha crd are created, the operator will simply print an error
to the log.
Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
In the 0.9 release rook converted the v1beta1 CRD resources
to v1 resources. This conversion code is no longer necessary
as we will be using v1 resources going forward.
The code paths will not be triggered anymore, thus
removing the dead code.
Signed-off-by: travisn <tnielsen@redhat.com>
This commit does:
* remove dead code SetCrushTunables()
* stop creating initial crushmap
There is no need to create an initial crushmap, since Ceph natively does
that for us.
We don't delete the existing configmap containing the initial crush map
of already deployed cluster, perhaps people are using for whatever
reason. It does not harm to keep it around.
Closes: https://github.com/rook/rook/issues/3138
Signed-off-by: Sébastien Han <seb@redhat.com>
When the operator first starts, the only operation needed
is to watch for new cephcluster crds to be created and
start the discovery to find available devices. The flexvolume
agent, the csi driver, and the volume provisioning can all be delayed
starting until the first cluster is created.
Signed-off-by: travisn <tnielsen@redhat.com>
Only the args necessary have traditionally been passed to the mon package.
This has led to many properties in the mon struct, whereas we will
now simply pass the cluster crd spec to simplify.
Signed-off-by: travisn <tnielsen@redhat.com>
The orchestrations need to be triggered on any update to the crd.
Users today restart the operator for various scenarios when the
operator could take care of it automatically. If unsupported options
are updated, they should just be ignored.
There are two goroutines that are working with the mons.
The first is an orchestration that is triggered by the operator
at startup, or when the cluster crd is updated. The second
is the health check that triggers periodically by default every
45 seconds. These goroutines must not try to make updates at the
same time. A mutex is added so one will block if the other is
still active. The mutex is only active for the duration of working
with mons, and not an entire orchestration of mgr, osd, etc
Signed-off-by: travisn <tnielsen@redhat.com>
Create a cephconfig module in Ceph's daemon pkg source, and refactor the
config and keyring generation that exists in the mon package into the
new cephconfig package. The config/keyring generation code is used by
most all daemons and not just mon, so a new package is a more
appropriate place for this.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Fix some basic spellcheck errors. Also remove trailing spaces and make
sure files have a newline (my editor does automatically).
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>