Commit Graph
12 Commits
Author SHA1 Message Date
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Sébastien Han f44583b2b9 ceph: remove dead code
As part of the transition to support Ceph release **as of** Nautilus, we
left over that portion of code.
We don't need it anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-16 14:01:24 +01:00
Sébastien Han 92c1696bfc ceph: cleanup/trim isUpgrade variable a bit
We now always perform ok-to-stop change when the deployment is going to
be updated. This could be for various reasons like change in the spec or
a new Ceph image.
In any case, if Ceph resources are going to be restarted this must be
done carefully.

Also trimmed the isUpgrade variable, the upcoming commit will remove the
need of that variable.

Keeping isUpgrade from the controller code since we need it to detect
whether an orchestration upgrade should be triggered or not (based on
the ceph cluster status).

Also, we keep it on the child controllers so we don't go into an
orchestration phase if the cluster image hasn't changed. The child does
not need to be updated/go into its orchestration phase because the
cluster CR changed.

Closes: https://github.com/rook/rook/issues/4379
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-24 17:02:37 +01:00
Sébastien Han 05aa639317 ceph: add option to continue on unclean PGs
A new CRD option `continueUpgradeAfterChecksEvenIfNotHealthy` is added.
When upgrading, Rook goes OSD by OSD and then waits for PGs to be clean
before proceeding to the next OSD. Currently, Rook waits for 5 hours but
there might be circumstances where PGs need more time to settle.
Thus setting `continueUpgradeAfterChecksEvenIfNotHealthy` to true will
pursue the upgrade process, even if PGs are not 100% active+clean.

Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1786029
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-07 12:18:54 -07:00
Sébastien Han 087d71106b ceph: add a new CR property to force upgrades
On rare occasion, the ceph checks might not be 100% correct and the user
might decide to force an upgrade anyway.
Correctly, if we are in an upgrade, we perform `ok-to-stop` check for
each daemon, if one of them fails, we retry and eventually fail.
They are corner-cases where users would like to force the upgrade anyway.

Enhance the new CR property: "skipUpgradeChecks: true".

Closes: https://github.com/rook/rook/issues/3872
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-19 10:55:56 +02:00
Sébastien Han bed14cfd71 ceph: improve upgrade on image change
From now on, Rook will only perform check before upgrades when there is
an actual upgrade. So if the Ceph image changed and a new version is
desired Rook will go through all the daemons and update them one by one
and perform checks in between.

Closes: https://github.com/rook/rook/issues/3583
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-08-29 10:18:19 +02:00
Sébastien Han 05d8d23f18 ceph: upgrade: pass daemonType and Name directly
We do not determine the daemon type and name by string parsing anymore
but passing that info directly in the update call, it's more reliable.
Also removing a bunch of unused functions.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-07-17 00:58:31 +02:00
Sébastien Han 9824bd2ff9 ceph: improve upgrade procedure
When a cluster is updated with a different image version, this triggers
a serialized restart of all the pods. Prior to this commit, no safety
check were performed and rook was hoping for the best outcome.

Now before doing restarting a daemon we check it can be restarted. Once
it's restarted we also check we can pursue with the rest of the
platform. For instance, with monitors we check that they are in quorum,
for OSD we check that PGs are clean and for MDS we make sure they are
 all active.

Fixes: https://github.com/rook/rook/issues/2889
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-07-17 00:37:43 +02:00
Alexander Trost 8bc26d9096 k8sclient: Update all operators to use apps/v1
All usages of k8s go client are now also using the versioned `AppsV1() `
call for the client.

Updated MySQL and Wordpress, and Kube Registy examples to use apps/v1
Deployments.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2019-04-12 09:25:42 +02:00
Blaine Gardner 4eb4fb7aa6 ceph mds: configure completely from operator
Configure the Ceph mds daemon completely from the operator a la the
recent changes to the Ceph mon and mgr operators.

Create the mds deployments first and then
create the keyring secrets for them with their owner reference as the
corresponding deployment. This will mean that the secrets do not need to
be micromanaged. When the deployment is deleted, the secret is also
deleted. This has not been necessary for the mons or the manager since
the mons share a keyring with a lifespan of the cluster, as does the
mgr, which currently has single-mgr support only.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-02-28 12:31:43 -07:00
Huamin Chen 1b21e42a90 convert extensions to apps
Signed-off-by: Huamin Chen <hchen@redhat.com>
2019-02-14 16:06:05 -05:00
Blaine Gardner 290b496069 ceph mon/mgr/mds units: verify updated deployments
Verify that the expected deployments are updated in the Ceph mgr,
and mds unit tests.

This also allows those unit tests to pass at all since
UpdateDeploymentAndWait was blocking the unit tests from finishing due
to unexpected behavior of generated `Deployment.Update` unit test mock.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-12-04 11:35:06 -07:00