Commit Graph
15 Commits
Author SHA1 Message Date
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Travis Nielsen 0c6ed25c4e core: allow downgrade of all daemons consistently
In the event a ceph image is specified that is lower than the current running
version of the daemons, the downgrade is allowed, even if not technically
supported. All of the core daemons (mon,mgr,osd) were being downgraded,
but the daemons for other controllers (rgw,mds,rbdmirror) were not being
downgraded, resulting in an inconsistent cluster. Now we log that the downgrade
is not supported and all all of the daemons to be downgraded.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-03 17:03:39 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Travis Nielsen 6f61d95dc3 ceph: clarify log message to continue with upgrade if needed
The ceph upgrade blocks if the ceph heatlh is HEALTH_ERR. If the user
wants to force the upgrade despite the health error, they will need
to set skipUpgradeChecks: true in the cephcluster CR. This message
will hopefully make that setting more discoverable.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-27 11:31:28 -06:00
Lars Lehtonen 600a1498fe ceph: fix multiple imports
This fixes double imports within and beneath pkg/operator/ceph/cluster.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-14 01:51:35 -07:00
subhamkrai 05d4c2776c ceph: update cephCluster CR with ceph versions output
update cephCluster CR with ceph versions output.
this output will contain ceph version of the
ceph daemons which will help with upgrade status.

ceph versions command output
```
ceph versions
{
    "mon": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 3
    },
    "mgr": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "osd": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "mds": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 2
    },
    "rbd-mirror": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "rgw": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "overall": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 9
    }
}
```

CR status
```
status:
  ceph:
   ---
    versions:
      mds:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 2
      mgr:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      mon:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 3
      osd:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      overall:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 9
      rbd-mirror:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      rgw:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
```

Signed-off-by: subhamkrai <srai@redhat.com>
2021-03-29 21:04:35 +05:30
subhamkraiandTravis Nielsen 319e4a41a4 ceph: placement in case of both PVC and non-PVC's
In the case of PVC,
We are giving lower priority to all placement.
We want deviceSet placement to applied and
override in case of overlapping settings and
we are merging nodeAffinity if applied in both
all placement and deviceSet.

In case of non-PVC,
we apply spec.placement

Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-18 10:32:57 +05:30
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
subhamkrai 6158b6608c ceph: do not merge nodeAffinity
we don't want to merge the node affinity if specified in both
storageClassDeviceSet and placement. Now, we are passing new
boolean argument in `ApplyToPodSpec` which will be false in
case of osd and prepare pods because we don't want to overlap
placement of osd on PVC's and non-PVC's.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-02-16 12:25:28 +05:30
Travis Nielsen ff1f74f771 ceph: correct formatting of ceph version detected
The ceph version during an upgrade was incorrectly being
formatted in the log message

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-02-10 16:36:09 -07:00
Sébastien Han 64fe9f7685 ceph: introduced unsupported releases
These are a collection of unstable ceph release that are not recommended
for production. It is recommended to rollback to the previous or
  upgrading to the next one if possible.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-05 18:35:57 +01:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Hiroshi Muraoka e2b1a2230d ceph: ignore PodAntiAffinity for detect-version placement
When a PodAntiAffinity is set to mon, detect-version may not be scheduled because detect-version has the same PodAntiAffinity with mon.
This commit fixes this issue.

Closes #5810

Signed-off-by: Hiroshi Muraoka <h.muraoka714@gmail.com>
2020-07-16 04:27:06 +00:00
Madhu Rajanna 81688398f2 cleanup: use err.Wrap when the formatting is not required
Replaced err.Wrapf with err.Wrap when the formatting
is not required.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-29 17:42:59 +05:30
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00