The failover of the arbiter mon in a stretch cluster was sometimes
failing due to the new tiebreaker not being set in ceph.
Rook would repeatedly try to remove the old tiebreaker mon
and keep failing because the new tiebreaker had not been set.
Now we make setting the tiebreaker idempotent in case the operator
restarts in the middle of the operation or some other corner
case causes the expected tiebreaker to be set. In that case,
the next reconcile will also ensure the tiebreaker mon is
set as expected.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
the mons re-initialize its ClusterInfo which results in missing
CR name on the ClusterInfo
Set the CR name to the mons cluster ClusterInfo
Closes: https://github.com/rook/rook/issues/9159
Signed-off-by: parth-gr <paarora@redhat.com>
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
Prior to ceph v16.2.7 the failover of the arbiter mon was
not supported. Now the new tiebreaker mon can be set during
the failover event and provide more dynamic stability to
the mon quorum if another node is available in the arbiter
zone.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery
Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
We don't need to use tini.
We don't have anything in the rook operator that would
either create zombie processes (no threads) or use
exec (to fork). The Go binary has a really good
signal handling mechanism.
Closes: https://github.com/rook/rook/issues/8794
Signed-off-by: Sébastien Han <seb@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
correct typo in logging, it was showing `ok-to-stop`
instead of `ok-to-continue` when 'continueUpgradeAfterChecksEvenIfNotHealthy' is true
Co-Authored-by: Zeaone <zeaone@ZeaonedeMacBook-Pro.local>
Signed-off-by: subhamkrai <srai@redhat.com>
While an even number of mons can cause lower availability of
mon quorum, it also can provide higher durability for the cluster.
Mon quorum can be restored from a single mon according to the
disaster recovery guide, so there may be scenarios where
an even number of mons may be preferable.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Stretch clusters in Ceph do not yet support failing over the arbiter
mon, so we now disable the arbiter mon from failing over. Soon
Ceph will support the arbiter mon failover and we will reenable
this failover feature.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
I don't know why these flags are there but it's not like we run the
operator with them and with a different value.
So removing for clarity.
Signed-off-by: Sébastien Han <seb@redhat.com>
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.
Signed-off-by: Sébastien Han <seb@redhat.com>
When changing the settings for the Rook mon health check
(.spec.healthCheck.mon) in the CephCluster resource,
the mon health check isn't reconfigured with the new values
Updated checkHealth goroutine so it can
updates the new patched values
Closes: https://github.com/rook/rook/issues/8363
Signed-off-by: parth-gr <paarora@redhat.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
Recently, the builds of `ceph/ceph` image moved to quay.io, see
https://github.com/ceph/ceph-build/pull/1883 for more details.
Current images will remain but new builds will happen on quay.io only.
This means that tags such as `v14.2`, `v15.2`,`v16.2` will need to
switch to quay.io to get updates.
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit update the PodDisruptionBudget policy to use version v1
Updated to policy/v1 as policy/v1beta1 PodDisruptionBudget is deprecated in v1.21+
Closes: https://github.com/rook/rook/issues/7917
Signed-off-by: parth-gr <paarora@redhat.com>
The mon daemon in a stretch cluster now can have its location set
as a CLI param instead of setting it with a separate command.
This enables mon failover to set the location of a mon immediately
when it is joining quorum instead of having a delayed command
to set the location.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.
So the automatic configuration of Ceph Filesystem peers is now possible.
By editing the CephFilesystem CRD, you can now turn on mirroring:
```yaml
mirroring:
enabled: false
# list of Kubernetes Secrets containing the peer token
# for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
peers:
secretNames:
- secondary-cluster-peer
```
Also, the mirroring status is displayed in the CR status:
```
status:
info:
fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
mirroringStatus:
daemonsStatus:
- daemon_id: 4186
filesystems:
- filesystem_id: 2
name: myfs
lastChecked: "2021-07-01T14:16:29Z"
phase: Ready
snapshotScheduleStatus:
lastChecked: "2021-07-01T14:16:29Z"
snapshotSchedules:
- fs: myfs
path: /
rel_path: /
retention: {}
schedule: 24h
```
Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
In a production cluster we never should have two mons running on
the same node. However, if two mons have ended up on the same node,
whether from a bug or some other unintentional event, the operator
will evict one of them and failover to a new mon. This only applies
when allowMultiplePerNode is false.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the mons are using the dataDirHostPath and not a volumeClaimTemplate,
they need to be permanently assigned to a node. The node assignment
was missing in stretch clusters when the dataDirHostPath was being
used. This resulted in mons that were moving to other nodes for example
during a node drain, which would cause the mon to lose its backing store.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If the number of desired mons is reduced, for example from 5 down to 3,
the extra mons will be removed by the next health check. Only a single
mon can be removed per health check to ensure that other mons do not
become unhealthy at the same time as removing an extra mon, and so the
list of healthy mons can be refreshed. If multiple mons are removed in
a single health check, we risk removing too many mons and losing quorum.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
After mon failover is initiated, there was a time window where if the operator
was restarted, the new mon is started and has joined quorum, but the operator
does not believe the mon should be in quorum after the operator restart.
The operator was mistakenly removing the extra mon prematurely, sometimes
causing quorum to be lost if another mon was also down at the same time.
If the mon does not come back online, steps to recover quroum would need
to be followed from the disaster guide. Now the expected list of mons
will be updated immediately during mon failover if the operator successfully
created the new mon deployment, thus removing the window where restarting
the operator can cause quorum loss.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If the cluster is external we want to periodically rehydrate the mgr
endpoint. This handles the scenarion where the active manager changes,
so we need to update the endpoint with the new IP address.
The create-external-cluster-resources.py script now requires an extra
permission to query the manager service so Rook can discover the active
one and its IP.
Testing:
```
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
"epoch": 37,
"available": true,
"active_name": "b",
"num_standby": 1
}
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME ENDPOINTS AGE
rook-ceph-mgr-external 172.17.0.12:9283 3m10s
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] k scale --replicas=0 deployment rook-ceph-mgr-b
deployment.apps/rook-ceph-mgr-b scaled
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
"epoch": 40,
"available": true,
"active_name": "a",
"num_standby": 0
}
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME ENDPOINTS AGE
rook-ceph-mgr-external 172.17.0.13:9283 3m55s
```
Signed-off-by: Sébastien Han <seb@redhat.com>
We don't need to set log_file to empty string since "log_to_file" is set
to False. This is probably an old leftover/hack we were doing in the
past when "log_to_file" and other related options were not there.
Because of this line
https://github.com/ceph/ceph/blob/master/src/perfglue/heap_profiler.cc#L99,
Ceph looks for the log_file conf option. In our case it was empty, so
the logging would default to the current directory which points to the
container runtime root when not set and we don't have permission to
write there.
We now force it to Ceph's default so that existing cluster will get the
fix after upgrading.
Just removing the option allows us to get dump in /var/log/ceph.
Phew, what a bug!
Signed-off-by: Sébastien Han <seb@redhat.com>
Events like node drain can take more than 10-15 minutes. If the node is not update the default monTimeOut of 10 minutes, then rook will attempt to fail over the mon. This failover won't work as the node is still down.
This PR retries once before the mon failover if the mon pod is not scheduled
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
If the cluster is deleted and the mon canaries cannot be scheduled for
any reasons, let's not wait and return.
Signed-off-by: Sébastien Han <seb@redhat.com>
If the health spec of the CephCluster CRD has a timeout set to 0 like
so:
```
healthCheck:
daemonHealth:
mon:
disabled: false
interval: 45s
timeout: 0
```
And the mon goes out of quorum then Rook will not fail over the mon.
This is interesting when doing maintenance on a monitor and we don't
want to create a new one.
Signed-off-by: Sébastien Han <seb@redhat.com>
With simultaneous node drains, a mon can go down while another mon is failing over. This results in two mons down and ceph become inaccessible. The commit makes maxUnavailable=0 while a mon is failing over and updates it back to maxUnavailable=1 after the failover. This update action is best effort. Any errors while updating the maxUnavailable mon pdb is only logged
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
- Update mon PDB to use maxUnavailable=1 instead of minAvailable. The maxUnavailable will always be 1 irrespective of the number of mons in the cluster
- Move mon PDB reconcile logic from Disruption Controller to Cluster Controller
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.
Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
Previously, `interval` was just a string so no special validation was
done by the OpenAPI validator. Now that it is advertised as metav1.Duration
the server will introspect it correctly. Underneath the type is still a
string so no update issue to be worried about.
Signed-off-by: Sébastien Han <seb@redhat.com>
In the case of PVC,
We are giving lower priority to all placement.
We want deviceSet placement to applied and
override in case of overlapping settings and
we are merging nodeAffinity if applied in both
all placement and deviceSet.
In case of non-PVC,
we apply spec.placement
Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>