This new integration test will deploy a cluster with multus enabled. It
will be comprised of two network interfaces for ceph public and cluster
communications.
For now, it only deploys a Ceph cluster up to the OSDs.
Closes: https://github.com/rook/rook/issues/9784
Signed-off-by: Sébastien Han <seb@redhat.com>
Data inconsistency might happen if the disk is accessed just after disk
zapping because direct I/O is not synchronous by itself.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This introduces a new CRD to add the ability to create subvolumegroup
for a given ceph filesystem volume. Typically the name of the volume is
the name of the filesystem created by rook.
Closes: https://github.com/rook/rook/issues/7036
Signed-off-by: Sébastien Han <seb@redhat.com>
This is handling a tricky scenario where the OSD deployment is manually
removed and the OSD never reconvers. This is unlikely to happen, but
still OSD should be able to run after that action. Essentially after a
manual deletion, we need to run the prepare job again to re-hydrate the
OSD information so that the OSD deployment can be deployed.
On encryption, it is a little bit tricky since ceph-volume list again
the main block won't return anything, so we need to target the encrypted
block to list.
There is another case this PR does not handle, which is the removal of
the OSD deployment and then the node is restarted. This means that the
encrypted container is not opened anymore. However, opening it requires
more work like writing the key on the filesystem (if not coming from the
Kubernete secret, eg,. KMS vault) and then run luksOpen. This is an
extreme corner case probably not worth worrying about for now.
Signed-off-by: Sébastien Han <seb@redhat.com>
Creating an EC pool is causing the CI to hang when the EC
pool is initialized since there aren't enough OSDs to
satisfy the EC parameters.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Adding cli agrument `--rbd-metadata-ec-pool-name` to read
rbd ec pool name to support ec pool in external cluster
and also updating the json blob.
Signed-off-by: subhamkrai <srai@redhat.com>
By default HPA can use details about memory or CPU consumption for
autoscaling, but also it can use custom metrics as well. There are alot
provides supports HPA via customer and one of them is KEDA project. Here
it is done with help of Prometheus Scaler
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
In the nightly job, run the test with the latest Ceph version so we can
detect if there are RGW changes in Ceph that might break multisite. Use
a reusable GitHub action workflow to duplicate as little code as
possible.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The yaml validation of the examples folder requires all the CRDs
to be created in advance of the dry-run command.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Add some commands to get more partition info when setting up the GH
action runner's disk for use in integration tests. This will both aid in
debugging and may "jog" the system such that it will no longer need to
reload the partition info when running the OSD prepare job.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The `head` command exits once it has output which can result in a
SIGPIPE error if the command piping its output to head hasn't yet
finished. Use `awk 'FNR <= 1'` instead, which waits on the input pipe to
close before it exits.
See here for more info:
https://unix.stackexchange.com/a/256047
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The CephObjectRealm controller would fail all subsequent reconciles if
the first reconcile created the Kubernetes Secret containing the access
keys for the realm but where the radosgw-admin command failed to create
the realm. This was the only idempotency issue found after reviewing the
CephObjectRealm controller.
Resolves#8954
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The daily arm test suite is failing due to the new local-build
image tag. Instead of waiting for the arm build to complete,
we can just pick up the latest tag from the same branch
that was already pushed to dockerhub and no need to build
again.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.
Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.
To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
of its operation that didn't cause failures in the past
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The canary tests were still picking up the tag from operator.yaml
and toolbox.yaml instead of the new test local-build tag.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 6f48dce3f5)
Try to avoid the following:
```
error pulling image configuration: received unexpected HTTP status: 500 Internal Server Error
```
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests must always be run against the local
build of rook, and an image should never be pulled from dockerhub.
To prevent pulling a release or master tag, the local build
will use a tag specific to the build and not ever published
elsewhere.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit a8a40428b0)
If the cluster where the rgw is started is secondary and not primary,
trying to create the admin ops user will fail with:
```
Please run the command on master zone.
Performing this operation on non-master zone
leads to inconsistent metadata between zones
```
So we need to force the creation regardless, it is fine the creation will
return UserAlreadyExist and then we just read the current user.
Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
We should not use .items[0].metadata.name if the array length is 0. This
is the case when nothing has been initialized yet. Instead, we should
use .items[*].metadata.name, the wildcard ensures to always return 0
even if nothing is present yet.
Fixes: https://github.com/rook/rook/issues/8676
Signed-off-by: Sébastien Han <seb@redhat.com>
we are adding new linter for shellcheck.
As we are writing more shell scripts this
will help maintain quality.
Also, doing all the changes required in
bash files to pass this shellcheck.
Closes: https://github.com/rook/rook/issues/8431
Signed-off-by: subhamkrai <srai@redhat.com>
This test:
- starts up 2 minikube clusters
- create 2 ceph clusters
- creates object multisite CRDs on each cluster and
syncs the clusters
- writes an object to cluster 1 and reads it on
cluster 2
This commit also adds new functions in
github-action-helper.sh that aid in the multisite
test and the multi cluster mirroring test.
Signed-off-by: Ali Maredia <amaredia@redhat.com>
No wonder why the build has been taking longer... We were building 3
times.
We just got at least 2 minutes back of build time thanks to this fix.
Signed-off-by: Sébastien Han <seb@redhat.com>
The build version should be detected with the git describe
command, but if the VERSION var is already set, it will be
used instead of being detected from git. During the official
builds or even local builds, we just want to let the makefile
detect the git version.
The full git history is added to enable proper detection of the
version. Without the full history only a hash is returned for
the version, which then results in an invalid semantic version
error for the helm chart.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.
So the automatic configuration of Ceph Filesystem peers is now possible.
By editing the CephFilesystem CRD, you can now turn on mirroring:
```yaml
mirroring:
enabled: false
# list of Kubernetes Secrets containing the peer token
# for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
peers:
secretNames:
- secondary-cluster-peer
```
Also, the mirroring status is displayed in the CR status:
```
status:
info:
fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
mirroringStatus:
daemonsStatus:
- daemon_id: 4186
filesystems:
- filesystem_id: 2
name: myfs
lastChecked: "2021-07-01T14:16:29Z"
phase: Ready
snapshotScheduleStatus:
lastChecked: "2021-07-01T14:16:29Z"
snapshotSchedules:
- fs: myfs
path: /
rel_path: /
retention: {}
schedule: 24h
```
Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
We have seen cases where the build might fail due to a small network
hiccup so let's retry up to 3 times if this is the case.
Signed-off-by: Sébastien Han <seb@redhat.com>
When configuring encrypted on-pvc clusters we now set the cluster fsid
in the LUKS header. If needed we can then determine if the encrypted
block is part of our cluster or not. We also attach the pvc_name in case
it might become useful.
Closes: https://github.com/rook/rook/issues/7991
Signed-off-by: Sébastien Han <seb@redhat.com>