Commit Graph
273 Commits
Author SHA1 Message Date
Travis Nielsen d55669293a mon: set stretch tiebreaker reliably during failover
The failover of the arbiter mon in a stretch cluster was sometimes
failing due to the new tiebreaker not being set in ceph.
Rook would repeatedly try to remove the old tiebreaker mon
and keep failing because the new tiebreaker had not been set.
Now we make setting the tiebreaker idempotent in case the operator
restarts in the middle of the operation or some other corner
case causes the expected tiebreaker to be set. In that case,
the next reconcile will also ensure the tiebreaker mon is
set as expected.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-02 12:50:07 -07:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Travis Nielsen 54a2b56b0c Merge pull request #9167 from parth-gr/mons-CRname
mon: update mons Cluster ClusterInfo with CR name
2021-11-15 11:11:09 -07:00
parth-gr 242f98e94b mon: update mons Cluster ClusterInfo with CR name
the mons re-initialize its ClusterInfo which results in missing
CR name on the ClusterInfo
Set the CR name to the mons cluster ClusterInfo

Closes: https://github.com/rook/rook/issues/9159
Signed-off-by: parth-gr <paarora@redhat.com>
2021-11-15 18:52:55 +05:30
Sébastien Han ecd7fa7880 Merge pull request #9164 from y1r/add-context-k8sutil-pod
core: add context parameter to k8sutil pod
2021-11-15 11:58:36 +01:00
Yuichiro Ueno 0559977b8a core: add context parameter to k8sutil pod
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 15:39:41 +09:00
Yuichiro Ueno 0b575703c7 core: add context parameter to k8sutil deployment
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 14:58:13 +09:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Travis Nielsen 12689bd119 ceph: enable mon failover for the arbiter in stretch mode
Prior to ceph v16.2.7 the failover of the arbiter mon was
not supported. Now the new tiebreaker mon can be set during
the failover event and provide more dynamic stability to
the mon quorum if another node is available in the arbiter
zone.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-18 08:58:25 -06:00
parth-gr 7c99858a77 ceph: add finalizers to rook-ceph-mon secrets and configmap
Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery

Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
2021-10-07 19:44:17 +00:00
Sébastien Han 121c2987e3 ceph: stop using tini
We don't need to use tini.
We don't have anything in the rook operator that would
either create zombie processes (no threads) or use
exec (to fork). The Go binary has a really good
signal handling mechanism.

Closes: https://github.com/rook/rook/issues/8794
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 10:54:16 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 031be4be19 Merge pull request #8675 from subhamkrai/ok-continue
ceph: modify the log info when ok to continue fails
2021-09-09 17:22:13 +02:00
subhamkraiandZeaone 0657804491 ceph: modify the log info when ok to continue fails
correct typo in logging, it was showing `ok-to-stop`
instead of `ok-to-continue` when 'continueUpgradeAfterChecksEvenIfNotHealthy' is true

Co-Authored-by: Zeaone <zeaone@ZeaonedeMacBook-Pro.local>
Signed-off-by: subhamkrai <srai@redhat.com>
2021-09-09 19:42:06 +05:30
JrCs c606f4c488 ceph: use node externalIP if no internalIP defined
In some cases node internalIP is not defined. Then use externalIP if it
exists.

Signed-off-by: JrCs <90z7oey02@sneakemail.com>
2021-09-08 16:26:05 +02:00
Travis Nielsen 1cb97574ea ceph: allow an even number of mons
While an even number of mons can cause lower availability of
mon quorum, it also can provide higher durability for the cluster.
Mon quorum can be restored from a single mon according to the
disaster recovery guide, so there may be scenarios where
an even number of mons may be preferable.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-03 07:43:32 -06:00
Travis Nielsen f9f322683a ceph: refuse to failover the arbiter mon on stretch clusters
Stretch clusters in Ceph do not yet support failing over the arbiter
mon, so we now disable the arbiter mon from failing over. Soon
Ceph will support the arbiter mon failover and we will reenable
this failover feature.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-10 16:26:47 -06:00
Sébastien Han 07d14a93a2 ceph: remove cli unused flags
I don't know why these flags are there but it's not like we run the
operator with them and with a different value.
So removing for clarity.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-06 11:05:20 +02:00
Sébastien Han 630c2f6a8b ceph: add an rbd-mirror bootstrap token on cluster creation
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-02 18:34:27 +02:00
parth-gr 6ebd109088 ceph: updated mon health check goroutine for reconfiguring patch values
When changing the settings for the Rook mon health check
(.spec.healthCheck.mon) in the CephCluster resource,
the mon health check isn't reconfigured with the new values

Updated checkHealth goroutine so it can
updates the new patched values

Closes: https://github.com/rook/rook/issues/8363
Signed-off-by: parth-gr <paarora@redhat.com>
2021-07-26 22:51:37 +05:30
Sébastien Han d55fb8d13b Merge pull request #8351 from leseb/refact-exec-helpers
ceph: remove unnecessary exec helpers
2021-07-23 18:12:57 +02:00
Sébastien Han 0811359d28 Merge pull request #8358 from leseb/move-to-quay
ceph: move all of our docker.io reference to quay.io
2021-07-23 18:03:53 +02:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Sébastien Han 6bce1ff3e9 ceph: move all of our docker.io reference to quay.io
Recently, the builds of `ceph/ceph` image moved to quay.io, see
https://github.com/ceph/ceph-build/pull/1883 for more details.
Current images will remain but new builds will happen on quay.io only.

This means that tags such as `v14.2`, `v15.2`,`v16.2` will need to
switch to quay.io to get updates.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-22 11:17:05 +02:00
parth-gr 77813f7093 ceph: update PodDisruptionBudget from v1beta1 to v1
This commit update the PodDisruptionBudget policy to use version v1
Updated to policy/v1 as policy/v1beta1 PodDisruptionBudget is deprecated in v1.21+

Closes: https://github.com/rook/rook/issues/7917
Signed-off-by: parth-gr <paarora@redhat.com>
2021-07-22 14:09:33 +05:30
Travis Nielsen fca761c950 ceph: set the location on the mon daemon
The mon daemon in a stretch cluster now can have its location set
as a CLI param instead of setting it with a separate command.
This enables mon failover to set the location of a mon immediately
when it is joining quorum instead of having a delayed command
to set the location.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-19 14:30:43 -06:00
Sébastien Han b578f916e7 ceph: add fs mirror config
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.

So the automatic configuration of Ceph Filesystem peers is now possible.

By editing the CephFilesystem CRD, you can now turn on mirroring:

```yaml
  mirroring:
    enabled: false
    # list of Kubernetes Secrets containing the peer token
    # for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
    peers:
      secretNames:
        - secondary-cluster-peer
```

Also, the mirroring status is displayed in the CR status:

```
status:
  info:
    fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
  mirroringStatus:
    daemonsStatus:
    - daemon_id: 4186
      filesystems:
      - filesystem_id: 2
        name: myfs
    lastChecked: "2021-07-01T14:16:29Z"
  phase: Ready
  snapshotScheduleStatus:
    lastChecked: "2021-07-01T14:16:29Z"
    snapshotSchedules:
    - fs: myfs
      path: /
      rel_path: /
      retention: {}
      schedule: 24h
```

Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 17:35:19 +02:00
Travis Nielsen b1c1e296dd ceph: evict a mon if colocated with another mon
In a production cluster we never should have two mons running on
the same node. However, if two mons have ended up on the same node,
whether from a bug or some other unintentional event, the operator
will evict one of them and failover to a new mon. This only applies
when allowMultiplePerNode is false.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-06-23 09:29:04 -06:00
Travis Nielsen a7066770f7 ceph: mons in stretch cluster assign to node
When the mons are using the dataDirHostPath and not a volumeClaimTemplate,
they need to be permanently assigned to a node. The node assignment
was missing in stretch clusters when the dataDirHostPath was being
used. This resulted in mons that were moving to other nodes for example
during a node drain, which would cause the mon to lose its backing store.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-06-18 15:25:27 -06:00
Travis Nielsen 9e982a0eff ceph: remove one extra mon per health check
If the number of desired mons is reduced, for example from 5 down to 3,
the extra mons will be removed by the next health check. Only a single
mon can be removed per health check to ensure that other mons do not
become unhealthy at the same time as removing an extra mon, and so the
list of healthy mons can be refreshed. If multiple mons are removed in
a single health check, we risk removing too many mons and losing quorum.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-27 14:34:09 -06:00
Travis Nielsen b0a63711f5 build: refactor to consolidate the rook.io/v1 package
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-18 19:55:37 -06:00
Travis Nielsen bb4191a8a7 ceph: persist expected mon endpoints immediately during mon failover
After mon failover is initiated, there was a time window where if the operator
was restarted, the new mon is started and has joined quorum, but the operator
does not believe the mon should be in quorum after the operator restart.
The operator was mistakenly removing the extra mon prematurely, sometimes
causing quorum to be lost if another mon was also down at the same time.
If the mon does not come back online, steps to recover quroum would need
to be followed from the disaster guide. Now the expected list of mons
will be updated immediately during mon failover if the operator successfully
created the new mon deployment, thus removing the window where restarting
the operator can cause quorum loss.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-12 09:44:04 -06:00
Sébastien Han 5e1e9f4c92 ceph: actively update the service endpoint for external mgr
If the cluster is external we want to periodically rehydrate the mgr
endpoint. This handles the scenarion where the active manager changes,
 so we need to update the endpoint with the new IP address.
The create-external-cluster-resources.py script now requires an extra
permission to query the manager service so Rook can discover the active
one and its IP.

Testing:

```
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 37,
    "available": true,
    "active_name": "b",
    "num_standby": 1
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.12:9283   3m10s

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] k scale --replicas=0 deployment rook-ceph-mgr-b
deployment.apps/rook-ceph-mgr-b scaled

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 40,
    "available": true,
    "active_name": "a",
    "num_standby": 0
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.13:9283   3m55s
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-11 16:08:40 +02:00
Sébastien Han 19de2b28e9 ceph: allow heap dump generation when logCollector is not running
We don't need to set log_file to empty string since "log_to_file" is set
to False. This is probably an old leftover/hack we were doing in the
past when "log_to_file" and other related options were not there.

Because of this line
https://github.com/ceph/ceph/blob/master/src/perfglue/heap_profiler.cc#L99,
Ceph looks for the log_file conf option. In our case it was empty, so
the logging would default to the current directory which points to the
container runtime root when not set and we don't have permission to
write there.

We now force it to Ceph's default so that existing cluster will get the
fix after upgrading.

Just removing the option allows us to get dump in /var/log/ceph.
Phew, what a bug!

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-10 17:21:04 +02:00
Santosh Pillai f38829b57d ceph: retry once before mon failover if mon pod is unscheduled
Events like node drain can take more than 10-15 minutes. If the node is not update the default monTimeOut of 10 minutes, then rook will attempt to fail over the mon. This failover won't work as the node is still down.
This PR retries once before the mon failover if the mon pod is not scheduled

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-05-05 23:32:27 +05:30
Nitin Goyal 3575c3bf78 ceph: expand Mon Pvc's
Expand Mon Pvc's if storage request has increased for the Mons in the
cephcluster crd.

Signed-off-by: Nitin Goyal <nigoyal@redhat.com>
2021-04-23 08:47:49 +05:30
Lars Lehtonen 600a1498fe ceph: fix multiple imports
This fixes double imports within and beneath pkg/operator/ceph/cluster.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-14 01:51:35 -07:00
Sébastien Han fc1e824a8d ceph: cancel canary retry if necessary
If the cluster is deleted and the mon canaries cannot be scheduled for
any reasons, let's not wait and return.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-04-12 15:46:58 +02:00
Travis Nielsen 721acd1a8a Revert "ceph: added rook-ceph-default service account"
This reverts commit 737fb099fe.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-08 11:46:12 -06:00
Sébastien Han e8c304d9b3 ceph: allow monitor maintenance mode
If the health spec of the CephCluster CRD has a timeout set to 0 like
so:

```
  healthCheck:
    daemonHealth:
      mon:
        disabled: false
        interval: 45s
        timeout: 0
```

And the mon goes out of quorum then Rook will not fail over the mon.
This is interesting when doing maintenance on a monitor and we don't
want to create a new one.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-04-08 16:27:53 +02:00
Sébastien Han 62d45239d4 ceph: use minutes instead of seconds
So we don't have to do the math :)

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-04-07 12:31:21 +02:00
Blaine Gardner 0f6a5bac76 Merge pull request #7442 from sp98/mon-drain-on-failover
ceph: prevent voluntary mon drain while another mon is failing over
2021-03-30 09:40:14 -06:00
Travis Nielsen 3c37def52f Merge pull request #7406 from parth-gr/service-account
ceph: added rook-ceph-default service account
2021-03-30 08:06:40 -06:00
Santosh Pillai 4af14f268c ceph: prevent voluntary mon drain while another mon is failing over
With simultaneous node drains, a mon can go down while another mon is failing over. This results in two mons down and ceph become inaccessible. The commit makes maxUnavailable=0 while a mon is failing over and updates it back to maxUnavailable=1 after the failover. This update action is best effort. Any errors while updating the maxUnavailable mon pdb is only logged

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-03-30 12:21:58 +05:30
Santosh Pillai 42121e909e ceph: update mon PDBs
- Update mon PDB to use maxUnavailable=1 instead of minAvailable. The maxUnavailable will always be 1 irrespective of the number of mons in the cluster
- Move mon PDB reconcile logic from Disruption Controller to Cluster Controller

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-03-30 12:21:21 +05:30
parth-grandTareq Sharafy 737fb099fe ceph: added rook-ceph-default service account
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.

Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
2021-03-29 20:20:40 +05:30
Sébastien Han d8de013443 ceph: use metav1.Duration for time CRD field
Previously, `interval` was just a string so no special validation was
done by the OpenAPI validator. Now that it is advertised as metav1.Duration
 the server will introspect it correctly. Underneath the type is still a
 string so no update issue to be worried about.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-25 12:12:31 +01:00
subhamkraiandTravis Nielsen 319e4a41a4 ceph: placement in case of both PVC and non-PVC's
In the case of PVC,
We are giving lower priority to all placement.
We want deviceSet placement to applied and
override in case of overlapping settings and
we are merging nodeAffinity if applied in both
all placement and deviceSet.

In case of non-PVC,
we apply spec.placement

Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-18 10:32:57 +05:30
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00