This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
The CI test for mirroring now runs on stable v16 tag. Internally our
code has a minimum version of 16.2.5 for cephfs mirroring since it's
available.
Signed-off-by: Sébastien Han <seb@redhat.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
Recently, the builds of `ceph/ceph` image moved to quay.io, see
https://github.com/ceph/ceph-build/pull/1883 for more details.
Current images will remain but new builds will happen on quay.io only.
This means that tags such as `v14.2`, `v15.2`,`v16.2` will need to
switch to quay.io to get updates.
Signed-off-by: Sébastien Han <seb@redhat.com>
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.
So the automatic configuration of Ceph Filesystem peers is now possible.
By editing the CephFilesystem CRD, you can now turn on mirroring:
```yaml
mirroring:
enabled: false
# list of Kubernetes Secrets containing the peer token
# for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
peers:
secretNames:
- secondary-cluster-peer
```
Also, the mirroring status is displayed in the CR status:
```
status:
info:
fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
mirroringStatus:
daemonsStatus:
- daemon_id: 4186
filesystems:
- filesystem_id: 2
name: myfs
lastChecked: "2021-07-01T14:16:29Z"
phase: Ready
snapshotScheduleStatus:
lastChecked: "2021-07-01T14:16:29Z"
snapshotSchedules:
- fs: myfs
path: /
rel_path: /
retention: {}
schedule: 24h
```
Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
We know append additional information to the rbd-mirror bootstrap peer
token. It is useful for disaster recovery scenario where the other
cluster is reading the peer token and needs to know the pool_id as well
as the namespace.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.
Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.
Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
In the case of PVC,
We are giving lower priority to all placement.
We want deviceSet placement to applied and
override in case of overlapping settings and
we are merging nodeAffinity if applied in both
all placement and deviceSet.
In case of non-PVC,
we apply spec.placement
Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
When deploying the ceph-fs-mirror daemon we don't need to name the
deployment like "rook-ceph-fs-mirror-a", we can simply go with
"rook-ceph-fs-mirror". Today, in Pacific, the daemon only supports a
single replica. In the future, the Ceph Quincy release should support
mulitiple concurrent daemons, at this point Rook will introduce a
"count" field in the CephFilesystemMirror CRD specification. Internally,
this field will map with the "replica" value of the Deployment.
Signed-off-by: Sébastien Han <seb@redhat.com>
we don't want to merge the node affinity if specified in both
storageClassDeviceSet and placement. Now, we are passing new
boolean argument in `ApplyToPodSpec` which will be false in
case of osd and prepare pods because we don't want to overlap
placement of osd on PVC's and non-PVC's.
Signed-off-by: subhamkrai <srai@redhat.com>
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.
The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050
The same goes for integration tests, they will get added later once we
start testing on Pacific.
Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>