Many of the canary tests have been failing much more
frequently in the past week or two. The test is typically
timing out pulling the image from quay.ceph.io since
it does not have as high bandwidth for the images.
A check is added to the test to wait specifically for the
first mon so it waits sufficiently for the image pull
before checking for other ceph daemons.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the release of K8s 1.32, we update the CI and docs
to support this new release, to maintain the most recent
six releases of K8s.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
This acceptance test demonstrates the creation of two CephObjectStore(s)
that share the same pool(s) manually managed by CephBlockPool(s).
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
This removes the execution of `sudo lsblk` three times for every single
invocation of the script. Usage of the BLOCK var is replaced with
functions which memoize the result of probing for block devices.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Factor out most of the CRs used by various canary tests to a new
deploy_cluster_full_of_cruft_please_stop_using_this() function.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Fix OSD isn't up.
As sdb device might change to vdb in the runner, let
find_extra_block_dev() exclude the nbd devices and find the proper
extra device for OSD.
Clean up the nbd devices after the test job is running.
Fix logs artifact upload twice and collect logs before clean up.
Signed-off-by: Xinliang Liu <xinliang.liu@linaro.org>
this commit upgrade the minikube, k8s, crictl versions
in CI and also fix permission error in the github runner.
Signed-off-by: subhamkrai <srai@redhat.com>
The docker.io image prefix is expected to be prepended
to the image names in the test images. This was missed
in 14550 related to some CI tests, which was now causing
the CI failures in the 1.15 branch where the search and
replace was missing the new docker.io prefix.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 3045076db8)
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.
Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
upgrading minimum kubernetes supported version to v1.24.17
and also upgrading other kubernetes version to their latest
respective version.
Signed-off-by: subhamkrai <srai@redhat.com>
Currently, when rook provisions OSDs(in the OSD prepare job), rook effectively run a
c-v command such as the following.
```console
ceph-volume lvm batch --prepare <deviceA> <deviceB> <deviceC> --db-devices <metadataDevice>
```
but c-v lvm batch only supports disk and lvm, instead of disk partitions.
We can resort to `ceph-volume lvm prepare` to implement it.
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
The 'extra' block device attached to GH actions runners has changed size
twice in 3 months. The previous strategy of detecting the disk by size
is becoming harder to maintain. Additionally, the block size with recent
changes (75G) is now the same as the boot device (also 75G), making the
method inexact.
The method can now be summarized as, "find the boot disk and choose the
disk that isn't the boot disk to be the 'extra' one used."
Prior to this, we used a one-liner based on `lsblk`. While we could
still make this a one-liner, the method is now updated to 2 effective
lines, plus debug text output to stderr to help if we need to debug
further in the future.
Of note, the 'extra' disk has a mount point of "/mnt", but it is unclear
whether this is a reliable heuristic for detecting the extra disk. For
years now, GH action runners have had only 2 disks. Therefore, it seems
slightly more likely that a heuristic to "choose the non-boot disk" will
be a more robust long-term solution.
If this strategy proves to be unreliable in the future, it may be wise
to consider whether "the device with a partition mounted to '/mnt'"
would be a good alternative.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
the disk size in the github action machine has
increased from 64G to 75G. Now, we detech the version
automatically not fetching hard coded value.
Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: subhamkrai <srai@redhat.com>
Currently, `wait_for_prepare_pod` only waits until 1 prepare pod and 1
OSD pod are in Running state.
So if `wait_for_ceph_to_be_ready` failed after `wait_for_prepare_pod`
succeeded, there are 4 possible cases:
1. some prepare pods didn't start correctly;
2. some prepare pods didn't finish correctly;
3. some OSD pods didn't start correctly; or
4. all prepare pods and OSD pods worked correctly, but some other
process failed.
As far as I understand, the number of prepare pods on the GitHub CI is
always 1, so cases 1 and 2 are not problematic. However, it is hard to
distinguish the other two cases from the CI log.
To solve the above problems, this patch makes wait_for_prepare_pod wait
for all OSD pods to become running state. If wait_for_prepare_pod
timeouts before the OSD pods become running state, this will be a strong
indication that the OSD pods didn't start properly (i.e., case 3).
Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
github ci runner is allocating both 14 and 64 Gb disks to the setup.
Rook CI currently hard codes the lsblk with a 14G filter.
This PR updates the filter to use either 14G or 64G disks.
Signed-off-by: sp98 <sapillai@redhat.com>
The issues are fixed from the clean disk action
repo and we shouldn't require any workaround. Also,
I'm not using git revert to revert the commit is earlier
clean disk action was mentioned in every CI which was
duplicate and now we have placed the action to composite yaml
Signed-off-by: subhamkrai <srai@redhat.com>
Since `google-cloud-sdk` is renamed with `google-cloud-cli`,
the github action is failing to remove the older name and this
is causing ci issues. Applying changes suggested by community to
manually remove some packages to resolve the issue.
Signed-off-by: subhamkrai <srai@redhat.com>
Update the multus canary test to reflect modern knowledge about how it
should be configured.
No longer test for the network device in OSD pods. Pods will utterly
fail to start if Multus is unable to attach interfaces.
Instead, look to the OSD map to test the connections more wholistically.
OSDs must have map IPs that include both public and cluster network.
This implicitly tests that the interfaces exist in the Pod, and it
additionally verifies other details, like Ceph `*_network` configs are
set propertly.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
using `rook/ceph:master` tags take longer time
to pull and also, we should be using `rook/ceph:local-build`
tag for our ci. This will help canary raw test to be more stable.
Signed-off-by: subhamkrai <srai@redhat.com>
add a new toolbox yaml manifest which will use the
rook image instead of ceph image
for running s5 cmd container needs to run with rook image
closes: https://github.com/rook/rook/issues/12227
Signed-off-by: parth-gr <paarora@redhat.com>
Let's use same ceph version(latest Reef) in both
cluster-test and toolbox.yaml so that we don't need
to pull image twice. Alos, github action helper script
was calling `deploy_manifest_with_local_build` which
is required for operator.yaml and not for toolbox.yaml.
Signed-off-by: subhamkrai <srai@redhat.com>
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
currently we just print the error is anything fails in
script for rgw, So by that there is no panic or
failiure of script if something wrong happened,
Added Explicittly forcing to catch the error from
the error message that is thrown
Closes: https://github.com/rook/rook/issues/12244
Signed-off-by: parth-gr <paarora@redhat.com>
The canary test is waiting for the prometheus module,
which is now disabled by default. For the canary test,
we need to enable the prometheus module for the external
cluster test.
Signed-off-by: travisn <tnielsen@redhat.com>
this commit bump minmum k8s version to 1.21.14 and
max k8s version to latest 1.26.0. Keeping support for
most recent 6 versions.
Signed-off-by: subhamkrai <srai@redhat.com>
The PSPs have long since been deprected. In K8s 1.21 the PSPs
were first deprecated, and support was completely removed
for them in 1.25. With Rook v1.11, the min supported version of
K8s is now 1.21. To reduce confusion in the documentation,
mention of the PSPs is now removed from the 1.11 docs.
For the corner case that users still require the PSPs,
the helm chart still contains the option for creating PSPs
or other users can still create the psp.yaml.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.
This feature is intended to be used for testing purposes only.
Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
Run the ganesha-rados-grace command in a remote pod when multus
networking is enabled.
Signed-off-by: parth-gr <paarora@redhat.com>
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The build process sometimes fails with intermittent problems. Some of them
are known problems and then we retry build process. However, there still
are unknown problems. In this case, it's hard to find the reason because
the build process exits immediately.
ref.
https://github.com/rook/rook/runs/8225144002?check_suite_focus=true#step:3:926
```
+ case "$o" in
+ exit 1
Error: Process completed with exit code 1.
```
To make debugging easier, let's print the output of `make`. This log won't be
too long since `make` prints most messages to stderr.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>