The build process sometimes fails with intermittent problems. Some of them
are known problems and then we retry build process. However, there still
are unknown problems. In this case, it's hard to find the reason because
the build process exits immediately.
ref.
https://github.com/rook/rook/runs/8225144002?check_suite_focus=true#step:3:926
```
+ case "$o" in
+ exit 1
Error: Process completed with exit code 1.
```
To make debugging easier, let's print the output of `make`. This log won't be
too long since `make` prints most messages to stderr.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
these indenation where added by vscode extentions.
for yaml I have extention ` YAML` by Red Hat
for bash I have `shell-format` and `shellCheck`.
Signed-off-by: subhamkrai <srai@redhat.com>
The whereabouts manifests in the master branch
have moved around, so until the new approach
is investigated we pin to the latest release
version v0.5.3
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook supports raw mode OSD in host-based cluster. So we can also
support OSD on logical volume in this kind of cluster.
Logical volumes aren't picked by filters (i.e. `useAllDevices: true`
and `device{Path,}Filter` to avoid unwanted LV consumption on upgrade.
Closes: https://github.com/rook/rook/issues/2047
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
It's better to test the creation of encrypted devices in host-based clusters.
This test will reduce the potential risks of regression when modifying
the osd-creation code.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
If csi is configured not the use the
hostnetworking, deploy the holder
pod for executing the commands with nsenter.
The implementation is same as multus.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
OSD on PVC is tested in many patterns but OSD on device is not.
It's preferable to add the following patterns for OSD on device.
- An OSD on a raw disk
- Two OSDs on a device
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The following steps in CI scripts create corrupted GPT headers.
```console
$ sudo sgdisk --zap-all --clear --mbrtogpt -g -- "$DISK"
$ sudo dd if=/dev/zero of="$DISK" bs=1M count=10
```
It results in the failure of the succeeding `sgdisk --print "$DISK".
Here is an example.
```console
$ sudo parted /dev/sdb mklabel msdos
...
$ sudo sgdisk --zap-all --clear --mbrtogpt -g -- /dev/sdb
...
$ sudo dd if=/dev/zero of=/dev/sdb bs=1M count=10
...
$ sudo sgdisk --print /dev/sdb
Caution: invalid main GPT header, but valid backup; regenerating main header
from backup!
Warning: Invalid CRC on main header data; loaded backup partition table.
Warning! One or more CRCs don't match. You should repair the disk!
Main header: ERROR
Backup header: OK
Main partition table: OK
Backup partition table: OK
Invalid partition data!
$ echo $?
2
```
We can safely use the simple `sgdisk --zap-all "$DEVICE"` here.
It's not necessary to convert "mbr" to "gpt" because we'll make new GPT
labels just after this command.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Once we are done cleaning up the content of the encrypted osd, let's
close the main LUKS device.
This avoids having dm devices on the system after the cleanup.
Closes: https://github.com/rook/rook/issues/10181
Signed-off-by: Sébastien Han <seb@redhat.com>
Support is removed for k8s for various limitations such as
priority classes not working and csi driver feature
incompleteness and missing snapshots. Documentation is updated
with the new min version of k8s 1.17 and also the tests are
updated to run on the min version of 1.17.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This new integration test will deploy a cluster with multus enabled. It
will be comprised of two network interfaces for ceph public and cluster
communications.
For now, it only deploys a Ceph cluster up to the OSDs.
Closes: https://github.com/rook/rook/issues/9784
Signed-off-by: Sébastien Han <seb@redhat.com>
finally, admission controller will be enabled default
without any script/manual step. But it still requires cert-manager
to be installed which I believe is already installed in clusters.
**Note**
Code doesn't return error it just logs the error since
we don't want to stop reconciling if the admission controller fails.
We can work on this once the admission controller is stable.
Signed-off-by: subhamkrai <srai@redhat.com>
Data inconsistency might happen if the disk is accessed just after disk
zapping because direct I/O is not synchronous by itself.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This commit refactor webhook config by adding/removing
spaces. Also, moved the Issue and Certificate creation
from bash script to webhook config yaml.
Signed-off-by: subhamkrai <srai@redhat.com>
The backend path was added with additional `transit` to it.
Also added PR test case to check transit engine
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
This introduces a new CRD to add the ability to create subvolumegroup
for a given ceph filesystem volume. Typically the name of the volume is
the name of the filesystem created by rook.
Closes: https://github.com/rook/rook/issues/7036
Signed-off-by: Sébastien Han <seb@redhat.com>
This is handling a tricky scenario where the OSD deployment is manually
removed and the OSD never reconvers. This is unlikely to happen, but
still OSD should be able to run after that action. Essentially after a
manual deletion, we need to run the prepare job again to re-hydrate the
OSD information so that the OSD deployment can be deployed.
On encryption, it is a little bit tricky since ceph-volume list again
the main block won't return anything, so we need to target the encrypted
block to list.
There is another case this PR does not handle, which is the removal of
the OSD deployment and then the node is restarted. This means that the
encrypted container is not opened anymore. However, opening it requires
more work like writing the key on the filesystem (if not coming from the
Kubernete secret, eg,. KMS vault) and then run luksOpen. This is an
extreme corner case probably not worth worrying about for now.
Signed-off-by: Sébastien Han <seb@redhat.com>
Creating an EC pool is causing the CI to hang when the EC
pool is initialized since there aren't enough OSDs to
satisfy the EC parameters.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Adding cli agrument `--rbd-metadata-ec-pool-name` to read
rbd ec pool name to support ec pool in external cluster
and also updating the json blob.
Signed-off-by: subhamkrai <srai@redhat.com>
Do not use the cross build container when building, publishing, and
promoting rook/ceph images. It is no longer needed, and its complexity
can add flakiness.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>