This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters
This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md
The max number of OSDs updated in parallel is currently fixed at 20.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The mgr daemon may be failed over by ceph if the active mgr is not
responding and the standby mgr is available. If the active mgr changes
the services for the dashboard and metrics will be updated with a
label selector for the new active mgr. The services cannot direct
traffic to the standby mgr or else they will be incorrectly redirected
to the active mgr.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
added waitTimeoutForHealthyOSD in the cluster crd that defines the time (in minutes) the operator would wait before an OSD can be stopped for
upgrade. This PR also removes the ok-to-continue logic for osds as its already handled by ok-to-stop. The default value is 10 minutes.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
If the pod spec changed, we expect an upgrade to proceed for that daemon.
If the check for a changed pod spec fails, we were skipping the update
of that daemon. Instead of skipping the update, we now assume the pod
spec changed if we fail to detect the change so that we can ensure the upgrade
even if we check for upgrades too often.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator should only print helpful info messages
when an OSD is going to be updated. If the OSD hasn't
changed it is just a debug message.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
An OSD on a PVC when portable=false is assigned to a node
with a node selector. The same node assignment is expected
for the lifetime of the cluster. On subsequent reconciles,
the operator was looking up the node assignment from the
nodeName on the pod spec, which is not set if the pod is down.
The operator needs to retrieve the assignment from the
deployment spec node selector.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If a deployment stays in pending we should give early by looking at
ProgressDeadlineExceeded, this will reduce the time to wait from 20 min
to 10 min because ProgressDeadlineExceeded default is 600 seconds.
Prior to this patch we would wait 20min since we take
currentDeployment.Spec.ProgressDeadlineSeconds which is typically 600
then retry every 2 seconds, which makes it 20min total.
Closes: https://github.com/rook/rook/issues/5090
Signed-off-by: Sébastien Han <seb@redhat.com>
A new CRD option `continueUpgradeAfterChecksEvenIfNotHealthy` is added.
When upgrading, Rook goes OSD by OSD and then waits for PGs to be clean
before proceeding to the next OSD. Currently, Rook waits for 5 hours but
there might be circumstances where PGs need more time to settle.
Thus setting `continueUpgradeAfterChecksEvenIfNotHealthy` to true will
pursue the upgrade process, even if PGs are not 100% active+clean.
Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1786029
Signed-off-by: Sébastien Han <seb@redhat.com>
Deployment behaves better when a node gets disconnected
from the rest of the cluster - new provisioner leader
is elected in ~15 seconds, while it may take up to
5 minutes for StatefulSet to start a new replica.
if kube version is 1.13.x deploy provisioner as statefulset.
if kube version is higher than 1.14+ deploy provisioner
as deployment.
Refer: kubernetes-csi/external-provisioner@52d1fbc
Refer: ceph/ceph-csi#497
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
- Added code to support StorageClassDeviceSet spec provided in the cluster-on-pvc.yaml
- The code reads the StorageClassDeviceSet spec and creates pvc based on the ‘count’ field for each device set.
- OSD prepare job is started for each PVC which activates the ceph-volume on each PVC
- Finally OSD is started on each of the PVC device.
Co-authored-by: rohan47 <rohgupta@redhat.com>
Co-authored-by: Ashish Ranjan <aranjan@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
When a cluster is updated with a different image version, this triggers
a serialized restart of all the pods. Prior to this commit, no safety
check were performed and rook was hoping for the best outcome.
Now before doing restarting a daemon we check it can be restarted. Once
it's restarted we also check we can pursue with the rest of the
platform. For instance, with monitors we check that they are in quorum,
for OSD we check that PGs are clean and for MDS we make sure they are
all active.
Fixes: https://github.com/rook/rook/issues/2889
Signed-off-by: Sébastien Han <seb@redhat.com>
Because the deployment update-and-wait function waits for the deployment
to be ready, and that cannot happen due to the keyring not
having been updated yet. The keyring has its
owner set to the deployment so that no special keyring handling must
be done, but it must be created after the deployment UID is known.
Change the order of operations so that the existing deployment's UID can
be retrieved if it exists and the keyring created/updated before the
update-and-wait function is called.
This order-of-operations issue currently only applies to the mds. Other
daemons whose keyrings are owned by the replication controller do not
use the update-and-wait function.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Add a `rook-version` label to all controller resources used by Ceph:
deployments, daemonsets, and jobs. Labels are added to controller
resources only and not to the pod templates within because changing pod
templates causes updates due to the label change. The controller
resources themselves are not updated with the label addition and
therefore don't cause an unnecessary upgrade.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
All usages of k8s go client are now also using the versioned `AppsV1() `
call for the client.
Updated MySQL and Wordpress, and Kube Registy examples to use apps/v1
Deployments.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Make the Ceph operator more cautious about when it decides to remove
nodes from the Rook-Ceph cluster which are acting as osd hosts.
When `useAllNodes` is set to `true` we assume that the user wants to
have the most hands-off experience. Node removals are allowed when a
node is delted from Kubernetes and when a node has its taints/affinities
modified by the user (but not by automatic k8s modification as much as
possible).
When `useAllnodes` is set to `false` the only time a node is removed is
if it is removed from the Ceph cluster definition.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
The Ceph upgrade test is the only one which uses the
`WaitForDeploymentImage` method, and it has to be configured to wait
longer after upgrade at this point. Since the wait time is still
hard-coded, this method is moved to the operator's test dir to make
it clear that the method is suitable only for tests currently.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
The Ceph mon flag `--public-bind-addr` does not keep the same port as
the previous daemon run when the port is left off the flag's value. A
port is undesired since that could interfere with Nautilus' msgr2
protocol, so instead configure the mon's services to forward the default
bind port on pods (6789) to the mon endpoint's port. The endpoint port
will be 6789 for new mons, but it may be 6790 for legacy clusters.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Configure the Ceph mds daemon completely from the operator a la the
recent changes to the Ceph mon and mgr operators.
Create the mds deployments first and then
create the keyring secrets for them with their owner reference as the
corresponding deployment. This will mean that the secrets do not need to
be micromanaged. When the deployment is deleted, the secret is also
deleted. This has not been necessary for the mons or the manager since
the mons share a keyring with a lifespan of the cluster, as does the
mgr, which currently has single-mgr support only.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
fix#2193
Fix regression introduced by PR allowing MDS config creation in init
container. This allows the new deployment-based mds clusters to scale
down where currently they can only remain constant or scale up.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Progress toward issue #2003.
Includes design from design doc PR #1578
Use init containers to create configuration for Ceph mgrs. There is only
one init container in this design. The init container calls the Rook
binary to create Ceph config files which are then shared with the mds
daemon main container.
Once this init is run, the main mds daemon is run. Leaving room to use
the Ceph-versioned image in the future, call `ceph-mds --foreground ...`
to run the Ceph mds.
The refactor to using an init container also necessitated refactoring
the mdses replicaset implementation to a deployment-per-pod
implementation due to a chicken-egg problem. With a single container (in
the before times) the Rook binary was able to call the ceph-mds daemon
with an id generated from the pod name. Since the pod name is not known
before runtime, and the id is one of the few params that must be
specified to Ceph daemons on run, it is necessary to know the id
beforehand; thus the move to a deployment architecture following the
likes of the mon and mgr daemons.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>