Commit Graph
37 Commits
Author SHA1 Message Date
Yuichiro Ueno 0b575703c7 core: add context parameter to k8sutil deployment
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 14:58:13 +09:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
Blaine Gardner 795124b7a8 ceph: update osds in parallel
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md

The max number of OSDs updated in parallel is currently fixed at 20.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-29 10:55:28 -06:00
Travis Nielsen 388ff3e78b ceph: start sidecar to monitor active mgr
The mgr daemon may be failed over by ceph if the active mgr is not
responding and the standby mgr is available. If the active mgr changes
the services for the dashboard and metrics will be updated with a
label selector for the new active mgr. The services cannot direct
traffic to the standby mgr or else they will be incorrectly redirected
to the active mgr.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00
Santosh Pillai 4237aa7ecc ceph: configure health check timeout for OSDs when upgrading
added waitTimeoutForHealthyOSD in the cluster crd that defines the time (in minutes) the operator would wait before an OSD can be stopped for
upgrade. This PR also removes the ok-to-continue logic for osds as its already handled by ok-to-stop. The default value is 10 minutes.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-01-27 11:13:05 +05:30
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Travis Nielsen 8664cbb8dd ceph: if spec diff checking fails during upgrade, assume it changed
If the pod spec changed, we expect an upgrade to proceed for that daemon.
If the check for a changed pod spec fails, we were skipping the update
of that daemon. Instead of skipping the update, we now assume the pod
spec changed if we fail to detect the change so that we can ensure the upgrade
even if we check for upgrades too often.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-09-17 13:22:53 -06:00
Travis Nielsen ff6a3f80cd ceph: improve osd update logging
The operator should only print helpful info messages
when an OSD is going to be updated. If the OSD hasn't
changed it is just a debug message.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-28 10:05:51 -06:00
Madhu Rajanna a16d3bd422 k8sutil: Make client as the first argument
To have parity with other functions and
also client should be the first argument to
the function.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-05-14 11:02:14 +05:30
Travis Nielsen 65377cbe63 ceph: osd on pvc node assignment is node selector
An OSD on a PVC when portable=false is assigned to a node
with a node selector. The same node assignment is expected
for the lifetime of the cluster. On subsequent reconciles,
the operator was looking up the node assignment from the
nodeName on the pod spec, which is not set if the pod is down.
The operator needs to retrieve the assignment from the
deployment spec node selector.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-07 10:51:46 -06:00
Sébastien Han c62d89c7bc k8sutil: mitigate UpdateDeploymentAndWait condition
If a deployment stays in pending we should give early by looking at
ProgressDeadlineExceeded, this will reduce the time to wait from 20 min
to 10 min because ProgressDeadlineExceeded default is 600 seconds.

Prior to this patch we would wait 20min since we take
currentDeployment.Spec.ProgressDeadlineSeconds which is typically 600
then retry every 2 seconds, which makes it 20min total.

Closes: https://github.com/rook/rook/issues/5090
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-26 14:15:45 +01:00
Sébastien Han caa8382e25 ceph: do not update deployment if no changes
We know use k8s-objectmatcher to verify whether a deployment is going to
change or not.
This avoids calling Update() from k8s utils which takes time to run.

Closes: https://github.com/rook/rook/issues/4519 and https://github.com/rook/rook/issues/4642
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-24 17:02:40 +01:00
Sébastien Han 05aa639317 ceph: add option to continue on unclean PGs
A new CRD option `continueUpgradeAfterChecksEvenIfNotHealthy` is added.
When upgrading, Rook goes OSD by OSD and then waits for PGs to be clean
before proceeding to the next OSD. Currently, Rook waits for 5 hours but
there might be circumstances where PGs need more time to settle.
Thus setting `continueUpgradeAfterChecksEvenIfNotHealthy` to true will
pursue the upgrade process, even if PGs are not 100% active+clean.

Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1786029
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-07 12:18:54 -07:00
Madhu Rajanna 8d5d9b9117 Fix liveness service creation
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-08-30 13:16:55 +05:30
Madhu Rajanna 11293393ca Fix logging in csi
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-08-30 13:16:55 +05:30
rohan47 21deea9d64 OSD: CRUSH map can be based on node or PVC in case of PVC backed OSD
Signed-off-by: rohan47 <rohgupta@redhat.com>
2019-08-29 04:01:41 +05:30
Madhu Rajanna 78fb35a93d Use Deployment with leader election instead of StatefulSet
Deployment behaves better when a node gets disconnected
from the rest of the cluster - new provisioner leader
is elected in ~15 seconds, while it may take up to
5 minutes for StatefulSet to start a new replica.

if kube version is 1.13.x deploy provisioner as statefulset.
if kube version is higher than 1.14+ deploy provisioner
as deployment.

Refer: kubernetes-csi/external-provisioner@52d1fbc
Refer: ceph/ceph-csi#497

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-08-26 19:29:37 +05:30
rohan47andAshish Ranjan d2f52aebe5 Adds support for storageClassDeviceSet in rook-ceph operator
- Added code to support StorageClassDeviceSet spec provided in the cluster-on-pvc.yaml
- The code reads the StorageClassDeviceSet spec and creates pvc based on the ‘count’ field for each device set.
- OSD prepare job is started for each PVC which activates the ceph-volume on each PVC
- Finally OSD is started on each of the PVC device.

Co-authored-by: rohan47 <rohgupta@redhat.com>
Co-authored-by: Ashish Ranjan <aranjan@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2019-08-12 09:24:13 -06:00
Travis Nielsen 7fb43d72eb Merge pull request #3546 from noahdesu/mon-pvc
ceph: base monitor storage of pvc
2019-08-02 22:25:57 -06:00
Noah Watkins 4173c5544e ceph: base monitor storage of pvc
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-08-02 10:30:23 -07:00
AllenZMC 4616f17a26 fix word ressource to resource
Signed-off-by: czm <zhongming.chang@daocloud.io>
2019-07-31 20:53:33 +08:00
Sébastien Han 9824bd2ff9 ceph: improve upgrade procedure
When a cluster is updated with a different image version, this triggers
a serialized restart of all the pods. Prior to this commit, no safety
check were performed and rook was hoping for the best outcome.

Now before doing restarting a daemon we check it can be restarted. Once
it's restarted we also check we can pursue with the rest of the
platform. For instance, with monitors we check that they are in quorum,
for OSD we check that PGs are clean and for MDS we make sure they are
 all active.

Fixes: https://github.com/rook/rook/issues/2889
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-07-17 00:37:43 +02:00
Blaine Gardner 966aef8681 ceph mds: update keyring before deployment update
Because the deployment update-and-wait function waits for the deployment
to be ready, and that cannot happen due to the keyring not
having been updated yet. The keyring has its
owner set to the deployment so that no special keyring handling must
be done, but it must be created after the deployment UID is known.
Change the order of operations so that the existing deployment's UID can
be retrieved if it exists and the keyring created/updated before the
update-and-wait function is called.

This order-of-operations issue currently only applies to the mds. Other
daemons whose keyrings are owned by the replication controller do not
use the update-and-wait function.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-15 12:26:46 -06:00
Blaine Gardner 685a8df08d ceph: add 'rook-version' label to controllers
Add a `rook-version` label to all controller resources used by Ceph:
deployments, daemonsets, and jobs. Labels are added to controller
resources only and not to the pod templates within because changing pod
templates causes updates due to the label change. The controller
resources themselves are not updated with the label addition and
therefore don't cause an unnecessary upgrade.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-12 13:56:31 -06:00
Alexander Trost 8bc26d9096 k8sclient: Update all operators to use apps/v1
All usages of k8s go client are now also using the versioned `AppsV1() `
call for the client.

Updated MySQL and Wordpress, and Kube Registy examples to use apps/v1
Deployments.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2019-04-12 09:25:42 +02:00
Blaine Gardner a80082988d ceph: remove osds only when certain
Make the Ceph operator more cautious about when it decides to remove
nodes from the Rook-Ceph cluster which are acting as osd hosts.

When `useAllNodes` is set to `true` we assume that the user wants to
have the most hands-off experience. Node removals are allowed when a
node is delted from Kubernetes and when a node has its taints/affinities
modified by the user (but not by automatic k8s modification as much as
possible).

When `useAllnodes` is set to `false` the only time a node is removed is
if it is removed from the Ceph cluster definition.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-11 09:41:54 -06:00
Blaine Gardner 35e2fbe4e0 optest: wait longer for deployment image
The Ceph upgrade test is the only one which uses the
`WaitForDeploymentImage` method, and it has to be configured to wait
longer after upgrade at this point. Since the wait time is still
hard-coded, this method is moved to the operator's test dir to make
it clear that the method is suitable only for tests currently.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-08 10:30:32 -06:00
Blaine Gardner feff587992 ceph: mon service forwards default bind port 6789
The Ceph mon flag `--public-bind-addr` does not keep the same port as
the previous daemon run when the port is left off the flag's value. A
port is undesired since that could interfere with Nautilus' msgr2
protocol, so instead configure the mon's services to forward the default
bind port on pods (6789) to the mon endpoint's port. The endpoint port
will be 6789 for new mons, but it may be 6790 for legacy clusters.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-03 15:53:56 -06:00
travisn 6ec255e269 Avoid silent failure during upgrade integration test
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-02 17:48:47 -06:00
Blaine Gardner 4eb4fb7aa6 ceph mds: configure completely from operator
Configure the Ceph mds daemon completely from the operator a la the
recent changes to the Ceph mon and mgr operators.

Create the mds deployments first and then
create the keyring secrets for them with their owner reference as the
corresponding deployment. This will mean that the secrets do not need to
be micromanaged. When the deployment is deleted, the secret is also
deleted. This has not been necessary for the mons or the manager since
the mons share a keyring with a lifespan of the cluster, as does the
mgr, which currently has single-mgr support only.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-02-28 12:31:43 -07:00
Huamin Chen 1b21e42a90 convert extensions to apps
Signed-off-by: Huamin Chen <hchen@redhat.com>
2019-02-14 16:06:05 -05:00
travisn feb33de24b tests: upgrade test is from v0.9 instead of v0.8
Signed-off-by: travisn <tnielsen@redhat.com>
2019-01-14 23:06:19 -07:00
travisn 94a7a0d4bc run daemons with the ceph image instead of rook
Signed-off-by: travisn <tnielsen@redhat.com>
2018-11-01 11:49:35 -06:00
Blaine Gardner 96f9a9c4ad ceph mds: allow mds deployment scale-down
fix #2193

Fix regression introduced by PR allowing MDS config creation in init
container. This allows the new deployment-based mds clusters to scale
down where currently they can only remain constant or scale up.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-10-12 14:53:08 -06:00
Blaine Gardner ce803fbf9d Ceph mds/file: Set up config in init container
Progress toward issue #2003.
Includes design from design doc PR #1578

Use init containers to create configuration for Ceph mgrs. There is only
one init container in this design. The init container calls the Rook
binary to create Ceph config files which are then shared with the mds
daemon main container.

Once this init is run, the main mds daemon is run. Leaving room to use
the Ceph-versioned image in the future, call `ceph-mds --foreground ...`
to run the Ceph mds.

The refactor to using an init container also necessitated refactoring
the mdses replicaset implementation to a deployment-per-pod
implementation due to a chicken-egg problem. With a single container (in
the before times) the Rook binary was able to call the ceph-mds daemon
with an id generated from the pod name. Since the pod name is not known
before runtime, and the id is one of the few params that must be
specified to Ceph daemons on run, it is necessary to know the id
beforehand; thus the move to a deployment architecture following the
likes of the mon and mgr daemons.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-10-05 19:01:12 -06:00
travisn 51e60b7b4d tests: upgrade integration test from 0.8.1 to master
Signed-off-by: travisn <tnielsen@redhat.com>
2018-08-23 16:43:10 -06:00
travisn 1d928ecd36 osd: init container to initialize config at startup
Signed-off-by: travisn <tnielsen@redhat.com>
2018-07-31 14:39:53 -06:00