To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
As part of the transition to support Ceph release **as of** Nautilus, we
left over that portion of code.
We don't need it anymore.
Signed-off-by: Sébastien Han <seb@redhat.com>
We now always perform ok-to-stop change when the deployment is going to
be updated. This could be for various reasons like change in the spec or
a new Ceph image.
In any case, if Ceph resources are going to be restarted this must be
done carefully.
Also trimmed the isUpgrade variable, the upcoming commit will remove the
need of that variable.
Keeping isUpgrade from the controller code since we need it to detect
whether an orchestration upgrade should be triggered or not (based on
the ceph cluster status).
Also, we keep it on the child controllers so we don't go into an
orchestration phase if the cluster image hasn't changed. The child does
not need to be updated/go into its orchestration phase because the
cluster CR changed.
Closes: https://github.com/rook/rook/issues/4379
Signed-off-by: Sébastien Han <seb@redhat.com>
A new CRD option `continueUpgradeAfterChecksEvenIfNotHealthy` is added.
When upgrading, Rook goes OSD by OSD and then waits for PGs to be clean
before proceeding to the next OSD. Currently, Rook waits for 5 hours but
there might be circumstances where PGs need more time to settle.
Thus setting `continueUpgradeAfterChecksEvenIfNotHealthy` to true will
pursue the upgrade process, even if PGs are not 100% active+clean.
Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1786029
Signed-off-by: Sébastien Han <seb@redhat.com>
On rare occasion, the ceph checks might not be 100% correct and the user
might decide to force an upgrade anyway.
Correctly, if we are in an upgrade, we perform `ok-to-stop` check for
each daemon, if one of them fails, we retry and eventually fail.
They are corner-cases where users would like to force the upgrade anyway.
Enhance the new CR property: "skipUpgradeChecks: true".
Closes: https://github.com/rook/rook/issues/3872
Signed-off-by: Sébastien Han <seb@redhat.com>
From now on, Rook will only perform check before upgrades when there is
an actual upgrade. So if the Ceph image changed and a new version is
desired Rook will go through all the daemons and update them one by one
and perform checks in between.
Closes: https://github.com/rook/rook/issues/3583
Signed-off-by: Sébastien Han <seb@redhat.com>
We do not determine the daemon type and name by string parsing anymore
but passing that info directly in the update call, it's more reliable.
Also removing a bunch of unused functions.
Signed-off-by: Sébastien Han <seb@redhat.com>
When a cluster is updated with a different image version, this triggers
a serialized restart of all the pods. Prior to this commit, no safety
check were performed and rook was hoping for the best outcome.
Now before doing restarting a daemon we check it can be restarted. Once
it's restarted we also check we can pursue with the rest of the
platform. For instance, with monitors we check that they are in quorum,
for OSD we check that PGs are clean and for MDS we make sure they are
all active.
Fixes: https://github.com/rook/rook/issues/2889
Signed-off-by: Sébastien Han <seb@redhat.com>
All usages of k8s go client are now also using the versioned `AppsV1() `
call for the client.
Updated MySQL and Wordpress, and Kube Registy examples to use apps/v1
Deployments.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Configure the Ceph mds daemon completely from the operator a la the
recent changes to the Ceph mon and mgr operators.
Create the mds deployments first and then
create the keyring secrets for them with their owner reference as the
corresponding deployment. This will mean that the secrets do not need to
be micromanaged. When the deployment is deleted, the secret is also
deleted. This has not been necessary for the mons or the manager since
the mons share a keyring with a lifespan of the cluster, as does the
mgr, which currently has single-mgr support only.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Verify that the expected deployments are updated in the Ceph mgr,
and mds unit tests.
This also allows those unit tests to pass at all since
UpdateDeploymentAndWait was blocking the unit tests from finishing due
to unexpected behavior of generated `Deployment.Update` unit test mock.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>