Commit Graph
45 Commits
Author SHA1 Message Date
Blaine Gardner 01e2feaef5 ceph: get rid of dynamic clientset
Stop using the dynamic clientset in favor of the controller-runtime
clientset.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-19 17:05:38 -06:00
parth-gr 7c99858a77 ceph: add finalizers to rook-ceph-mon secrets and configmap
Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery

Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
2021-10-07 19:44:17 +00:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Blaine Gardner 21e290e003 ceph: implement dependencies for CephCluster
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-10 10:06:10 -06:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen d1da12c6ac ceph: wait indefinitely for cleanup before removing cluster finalizer
During cluster deletion, we currently only retry for a couple minutes
to wait for the pvcs to be deleted. After the timeout, we proceed
with the cluster deletion. To properly protect the pvcs for proper
cleanup, the finalizer should not be removed until the pvcs
are all confirmed to be deleted. In order to not block other cluster
events, we re-queue the deletion event to run again every 10s
until the pvcs are deleted or the finalizer is manually removed.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Sébastien Han 7c994441c3 ceph: make admin key optional for external cluster
Currently, we need to configure the Ceph external Admin keyi
in the Rook deployment to be able to connect to an external Ceph cluster.
If we wanted to run in a multi-tenant fashion
were several K8s/Rook clusters wanted to connect to the same external ceph cluster,
each K8s deployment would have the access to the External Ceph Admin key
and could potentially access or delete the Data
from pools that belong to other k8s/Rook Clusters.

Now the admin key is optional but the helper script create-external-cluster-resources.sh
will help create the necessary keys/users to connect to that cluster.

Closes: https://github.com/rook/rook/issues/4917 and https://github.com/rook/rook/pull/5227
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-20 09:59:04 +02:00
Travis Nielsen 348690ea64 ceph: update defaults to the ceph v14.2.9 release
With the release of ceph v14.2.9 we set the default examples to this version.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-17 12:54:06 -06:00
Sébastien Han 14cb46ae39 ceph: bump to 14.2.8
Ceph 14.2.8 just got released so let's use it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-11 19:32:48 +01:00
Umanga Chapagain c54c0d16cf Ceph: adds watch to rook-ceph-operator-config ConfigMap
watches the rook-ceph-operator-config ConfigMap and updates
CSI driver

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-10 12:46:54 +05:30
Elise Gafford de2c409d79 ceph: delete CSI drivers with deletion of last cluster
Prior attempts to delete CSI drivers using owner references to the
Ceph ConfigMap resulted in garbage collection of the driver
resources as described in rook#4590. This patch manually deletes these
resources by name as a stopgap to provide an appropriate user
experience while the team investigates this issue further.

Resolves: rook#4824
Signed-off-by: Elise Gafford <egafford@redhat.com>
2020-02-27 19:06:23 -05:00
Travis Nielsen bcb99d86ce crds: pick up the rook types in the v1 package
The rook types used across the storage providers moved from the v1alpha2
package to the v1 package. This commit points the packages at their new
location. Implementation is expected to remain unchanged.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-26 11:18:17 -07:00
Nizamudeen 53883f68cf ceph: Handling Unhandled errors
This commit is to handle all those unhandled errors which raises the gosec warning.

Fixed G104: Unhandled Errors are handled now

Signed-off-by: Nizamudeen <nia@redhat.com>
2020-02-21 22:48:24 +05:30
Travis Nielsen e35939b26c ceph: bump ceph version to v14.2.7
The ceph/ceph:v14.2.7 image is released so we can pick these fixes
up as the recommended version of ceph.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-20 11:06:40 -07:00
Nizamudeen 4b55a74ac4 ceph: Implementing Conditions on rook-ceph
Fixed conditions getting resetted after the operator restart.
Did the changes which required to implement conditions on the rook ceph cluster
Conditions will eliminate the current status.State and incorporates a type which
provides much more description to the current status of the cluster.

Signed-off-by: Nizamudeen <nia@redhat.com>
2020-01-30 10:21:11 +05:30
Sébastien Han dec0c7d4b0 ceph: bump to 14.2.6
14.2.6 is released, let's use it!

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-15 11:02:44 +01:00
Travis Nielsen 3294384da7 Merge pull request #4415 from leseb/use-minor-version
rook-ceph: stop using ceph version with timestamp and bump to 14.2.5
2019-12-10 08:14:51 -07:00
Sébastien Han 68fd4f4df9 rook: stop using ceph version with timestamp
With https://github.com/ceph/ceph-container/pull/1529, we now build
vX.X.X version so we don't need to have the timestamp. Also we will
automatically get the last changes from the OS/packages too more
frequently.

Closes: https://github.com/rook/rook/issues/4414
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-10 11:39:35 +01:00
Sébastien Han 5ce2ed220e ceph: use "github.com/pkg/errors"
We now use the error package.
Kubernetes errors have been renamed kerrors since they are lower than
'errors'.

Closes: https://github.com/rook/rook/issues/4054
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-09 16:58:32 +01:00
Elise Gafford 3adcb67755 ceph: add removeCallbacks hook to Operator and ClusterController
In order to remove Ceph-CSI resources after deletion of the last
Ceph cluster, the operator must be able to perform post-processing
actions on cluster deletion. This change creates a hook for the
addition of post-deletion operator callback functions.

Partially resolves: #4234
Signed-off-by: Elise Gafford <egafford@redhat.com>
2019-11-21 11:11:25 -05:00
Sébastien Han 148f8da1fb ceph: relax pre-requisite for external cluster
We now differentiate the cases where:

* we only consume the external cluster
* we consume the external cluster as well as creating stateless
resources in Kubernetes (bootstrap mds,rgw, nfs)

This is mostly controlled via the image property spec. If not defined,
not extra CRs won't be able to be created.

Now the external cluster feature supports Ceph cluster as of Luminous 12.2.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-10-08 10:52:01 +02:00
Kristoffer Grönlund fb599dd02b Ceph: Remove finalizer even if flex is disabled
When the flex driver is disabled, the check for
volume attachments will fail and the finalizer
is never removed. To avoid this, just log the
failure to list volumes and remove the
finalizer anyway.

Resolves #3912

Signed-off-by: Kristoffer Grönlund <kgronlund@suse.com>
2019-09-24 08:22:30 +02:00
Sébastien Han bc6dcb9995 ceph: bump to v14.2.4
v14.2.4 just got released so let's use it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-17 15:08:15 +02:00
Sébastien Han 3390a85404 ceph: bump to 14.2.3
Let's use the latest minor release of Ceph 14.2.3.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-04 23:14:28 +02:00
Rohan CJ 4072d14557 Add tests for the clusterdisruption controller and update the docs and
pending releas notes for managed PDBs.

Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-08-30 14:04:23 +05:30
Rohan CJ 60792ebe5c Add support for PDBs on Mons, RGW, and MDS
Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-08-30 09:01:15 +05:30
Sébastien Han a42916eff0 ceph: use latest ceph images
The last image has disabled ephemeral repositories so it's now possible
for images older than 15 days to install packages without having an
  error from non-existing repositories.

Closes: https://github.com/rook/rook/issues/3662
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-08-26 23:50:55 +02:00
Sébastien Han 2d24065ddb ceph: add check for external cluster spec
Before attempting to do anything let's just make sure the spec is
correctly fulfilled.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-08-22 17:19:33 +02:00
travisn 7ab689ee8e ceph: skip local setup for an external ceph cluster
When configuring an external cluster the orchestration of discover and all the ceph
daemons will be skipped. The operator will still watch for the creation of
crds for this namespace. When a filesystem, object, object user, or
ganesha crd are created, the operator will simply print an error
to the log.

Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-08-22 17:19:32 +02:00
travisn 9e073857c5 ceph: remove legacy conversion from v1beta1 to v1
In the 0.9 release rook converted the v1beta1 CRD resources
to v1 resources. This conversion code is no longer necessary
as we will be using v1 resources going forward.
The code paths will not be triggered anymore, thus
removing the dead code.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-07-15 16:58:37 -06:00
Sébastien Han ccf2110007 ceph: stop creating initial crushmap
This commit does:

* remove dead code SetCrushTunables()
* stop creating initial crushmap

There is no need to create an initial crushmap, since Ceph natively does
that for us.
We don't delete the existing configmap containing the initial crush map
of already deployed cluster, perhaps people are using for whatever
reason. It does not harm to keep it around.

Closes: https://github.com/rook/rook/issues/3138
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-07-08 12:07:46 +02:00
travisn c6c4a9b42a ceph: delay starting the system daemons until a cluster is created
When the operator first starts, the only operation needed
is to watch for new cephcluster crds to be created and
start the discovery to find available devices. The flexvolume
agent, the csi driver, and the volume provisioning can all be delayed
starting until the first cluster is created.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-06-17 15:44:33 -06:00
travisn 3548418774 print the diff for changes to the cluster crd
Signed-off-by: travisn <tnielsen@redhat.com>
2019-03-14 13:10:52 -06:00
travisn c1628141aa minimize args to mon creation in the operator
Only the args necessary have traditionally been passed to the mon package.
This has led to many properties in the mon struct, whereas we will
now simply pass the cluster crd spec to simplify.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-03-14 13:10:52 -06:00
travisn 2fd1f7345d Trigger an orchestration on any crd update and only allow a single orchestration of mons at a time
The orchestrations need to be triggered on any update to the crd.
Users today restart the operator for various scenarios when the
operator could take care of it automatically. If unsupported options
are updated, they should just be ignored.

There are two goroutines that are working with the mons.
The first is an orchestration that is triggered by the operator
at startup, or when the cluster crd is updated. The second
is the health check that triggers periodically by default every
45 seconds. These goroutines must not try to make updates at the
same time. A mutex is added so one will block if the other is
still active. The mutex is only active for the duration of working
with mons, and not an entire orchestration of mgr, osd, etc

Signed-off-by: travisn <tnielsen@redhat.com>
2019-03-14 13:10:12 -06:00
travisn a29b876337 rename ceph v1 crds with the ceph prefix
Signed-off-by: travisn <tnielsen@redhat.com>
2018-12-05 16:23:55 -07:00
travisn 62f7e5d6ea ceph: update docs, code, and tests to use v1 crd types
Signed-off-by: travisn <tnielsen@redhat.com>
2018-12-05 14:33:11 -07:00
travisn b02cd099f1 mons: allow scaling up or down the size of quorum
Signed-off-by: travisn <tnielsen@redhat.com>
2018-10-17 11:00:05 -06:00
Blaine Gardner 3731f06176 Ceph: Refactor mon config/keyring to cephconfig
Create a cephconfig module in Ceph's daemon pkg source, and refactor the
config and keyring generation that exists in the mon package into the
new cephconfig package. The config/keyring generation code is used by
most all daemons and not just mon, so a new package is a more
appropriate place for this.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-09-05 10:17:11 -06:00
Blaine Gardner fe5fb394c7 Fix spellcheck and trailing space/newline issues
Fix some basic spellcheck errors. Also remove trailing spaces and make
sure files have a newline (my editor does automatically).

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-07-26 11:36:09 -06:00
Jared Watts 189e3eb611 ceph: migration logic for ceph.rook.io/v1alpha1 to ceph.rook.io/v1beta1, update all type references
Signed-off-by: Jared Watts <jbw976@gmail.com>
2018-07-11 09:23:46 -07:00
Jared Watts f3df68e573 operator, daemon, and cmd updates for supporting multiple storage types
Signed-off-by: Jared Watts <jbw976@gmail.com>
2018-05-18 14:32:20 -07:00