Commit Graph
119 Commits
Author SHA1 Message Date
Travis Nielsen 56c6659655 tests: canary tests to wait for first mon to start
Many of the canary tests have been failing much more
frequently in the past week or two. The test is typically
timing out pulling the image from quay.ceph.io since
it does not have as high bandwidth for the images.
A check is added to the test to wait specifically for the
first mon so it waits sufficiently for the image pull
before checking for other ceph daemons.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-12-20 13:53:21 -07:00
Travis Nielsen 7c2f41e72e build: add support for k8s 1.32
With the release of K8s 1.32, we update the CI and docs
to support this new release, to maintain the most recent
six releases of K8s.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-12-12 09:48:42 -07:00
Blaine Gardner 5539eedd1b multus: finish deprecating holder pods
Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-10-23 16:29:02 -06:00
Joshua Hoblitt 100c8bd7a9 test: require ceph tag param to replace_ceph_image
To prevent silent failures where the ceph image tag is not updated.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-18 12:08:57 -07:00
Joshua Hoblitt 11250013a0 test: add two-object-one-zone canary test
This acceptance test demonstrates the creation of two CephObjectStore(s)
that share the same pool(s) manually managed by CephBlockPool(s).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-14 16:54:39 -07:00
Joshua Hoblitt 92d9f994c2 test: improve reliability of canary rgw-multisite-testing
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-12 12:42:58 -07:00
Joshua Hoblitt d301114680 test: do not always run sudo lsblk in github-action-helper.sh
This removes the execution of `sudo lsblk` three times for every single
invocation of the script.  Usage of the BLOCK var is replaced with
functions which memoize the result of probing for block devices.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 16:40:10 -07:00
Joshua Hoblitt 581fd5c197 test: convert all github-action-helper functs to $REPO_DIR
This allows all functions to be called in any order without concern for
the CWD.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:18:17 -07:00
Joshua Hoblitt 31d55b90fe test: mv canary test specific resources out of deploy_cluster()
Factor out most of the CRs used by various canary tests to a new
deploy_cluster_full_of_cruft_please_stop_using_this() function.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:18:17 -07:00
Joshua Hoblitt ed8156017e test: add object-with-cephblockpool canary test
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:16:30 -07:00
Joshua Hoblitt cb1dd152ad test: add github-action-helper toolbox functions
Added these functions for running commands in the toolbox pod:

- toolbox()
- ceph()
- rbd()
- radosgw-admin()

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:16:30 -07:00
Xinliang Liu 05ac99c0c0 ci: fix canary-arm64 job
Fix OSD isn't up.
As sdb device might change to vdb in the runner, let
find_extra_block_dev() exclude the nbd devices and find the proper
extra device for OSD.

Clean up the nbd devices after the test job is running.

Fix logs artifact upload twice and collect logs before clean up.

Signed-off-by: Xinliang Liu <xinliang.liu@linaro.org>
2024-10-01 12:05:32 -06:00
subhamkrai f647444515 ci: fix ci permission issue with minikube start
this commit upgrade the minikube, k8s, crictl versions
in CI and also fix permission error in the github runner.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-09-19 15:55:18 +05:30
Travis Nielsen 8b60c52f31 build: generate the local build tag with docker io
The docker.io image prefix is expected to be prepended
to the image names in the test images. This was missed
in 14550 related to some CI tests, which was now causing
the CI failures in the 1.15 branch where the search and
replace was missing the new docker.io prefix.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 3045076db8)
2024-08-21 10:21:31 -06:00
subhamkrai 128ec16ee6 ci: add k8s 1.30 support in ci
adding support for kubernetes 1.30 in ci which released yesterday.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-04-19 07:41:05 +05:30
Travis Nielsen fdacfd51c5 object: create an object store based on shared pools
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.

Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-03-11 11:20:57 -06:00
Sunnatillo 3ad456c6a2 build: uplift prometheus operator version to v0.71.1
This commit uplifts prometheus operator version to v0.71.1.

Signed-off-by: Sunnatillo <sunnat.samadov@est.tech>
2024-02-27 20:00:25 +02:00
subhamkrai 135307a4df ci: upgrade min k8s supported version to 1.24.17
upgrading minimum kubernetes supported version to v1.24.17
and also upgrading other kubernetes version to their latest
respective version.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-02-13 22:11:44 +05:30
Liang Zheng ee05e82cea osd: support create osd with metadata partition
Currently, when rook provisions OSDs(in the OSD prepare job), rook effectively run a
c-v command such as the following.
```console
ceph-volume lvm batch --prepare <deviceA> <deviceB> <deviceC> --db-devices <metadataDevice>
```
but c-v lvm batch only supports disk and lvm, instead of disk partitions.

We can resort to `ceph-volume lvm prepare` to implement it.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2024-02-07 11:34:09 +08:00
sp98 45adae1931 ci: fix ci failure for multicluster tests
This PR fixes the failure while running multicluster mirroring CI tests

Signed-off-by: sp98 <sapillai@redhat.com>
2024-02-06 12:54:18 +05:30
Blaine Gardner feacb6424e ci: fix detection of GH actions extra disk
The 'extra' block device attached to GH actions runners has changed size
twice in 3 months. The previous strategy of detecting the disk by size
is becoming harder to maintain. Additionally, the block size with recent
changes (75G) is now the same as the boot device (also 75G), making the
method inexact.

The method can now be summarized as, "find the boot disk and choose the
disk that isn't the boot disk to be the 'extra' one used."

Prior to this, we used a one-liner based on `lsblk`. While we could
still make this a one-liner, the method is now updated to 2 effective
lines, plus debug text output to stderr to help if we need to debug
further in the future.

Of note, the 'extra' disk has a mount point of "/mnt", but it is unclear
whether this is a reliable heuristic for detecting the extra disk. For
years now, GH action runners have had only 2 disks. Therefore, it seems
slightly more likely that a heuristic to "choose the non-boot disk" will
be a more robust long-term solution.

If this strategy proves to be unreliable in the future, it may be wise
to consider whether "the device with a partition mounted to '/mnt'"
would be a good alternative.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-02-05 17:55:16 -07:00
subhamkraiandJan Klippel 6b0deb32d9 ci: disk in github action increased to 75G from 64G
the disk size in the github action machine has
increased from 64G to 75G. Now, we detech the version
automatically not fetching hard coded value.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2024-02-05 22:24:03 +05:30
Ryotaro Banno e9c1f81d79 ci: wait for all OSD pods to become Running in wait_for_prepare_pod
Currently, `wait_for_prepare_pod` only waits until 1 prepare pod and 1
OSD pod are in Running state.

So if `wait_for_ceph_to_be_ready` failed after `wait_for_prepare_pod`
succeeded, there are 4 possible cases:

1. some prepare pods didn't start correctly;
2. some prepare pods didn't finish correctly;
3. some OSD pods didn't start correctly; or
4. all prepare pods and OSD pods worked correctly, but some other
   process failed.

As far as I understand, the number of prepare pods on the GitHub CI is
always 1, so cases 1 and 2 are not problematic. However, it is hard to
distinguish the other two cases from the CI log.

To solve the above problems, this patch makes wait_for_prepare_pod wait
for all OSD pods to become running state. If wait_for_prepare_pod
timeouts before the OSD pods become running state, this will be a strong
indication that the OSD pods didn't start properly (i.e., case 3).

Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
2023-12-25 05:12:30 +00:00
Alexander Trost 4ed35d6bd6 operator: allow setting ceph config options via ceph cluster crd
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-12-02 00:05:48 +01:00
sp98 345f92961b ci: filter both 14 and 64 disks in CI
github ci runner is allocating both 14 and 64 Gb disks to the setup.
Rook CI currently hard codes the lsblk with a 14G filter.
This PR updates the filter to use either 14G or 64G disks.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-11-06 13:30:10 +05:30
subhamkrai e14f8c7710 ci: remove workaround for clean disk action
The issues are fixed from the clean disk action
repo and we shouldn't require any workaround. Also,
I'm not using git revert to revert the commit is earlier
clean disk action was mentioned in every CI which was
duplicate and now we have placed the action to composite yaml

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-29 19:17:01 +05:30
subhamkrai 3267576307 ci: fix github action free disk space
Since `google-cloud-sdk` is renamed with `google-cloud-cli`,
the github action is failing to remove the older name and this
is causing ci issues. Applying changes suggested by community to
manually remove some packages to resolve the issue.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-28 18:56:59 +05:30
Blaine Gardner 3e2e906ea7 test: update multus canary test
Update the multus canary test to reflect modern knowledge about how it
should be configured.

No longer test for the network device in OSD pods. Pods will utterly
fail to start if Multus is unable to attach interfaces.
Instead, look to the OSD map to test the connections more wholistically.
OSDs must have map IPs that include both public and cluster network.
This implicitly tests that the interfaces exist in the Pod, and it
additionally verifies other details, like Ceph `*_network` configs are
set propertly.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-09-18 17:15:23 -06:00
subhamkrai 90d81e11b6 ci: upgrade the kubernetes versions for ci
since, kubernetes v1.28.0 release upgrading the
latest k8s version to v.1.28.0 excepth objectSuite
tes.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-28 10:35:44 +05:30
subhamkrai ad50c8e780 ci: use local tag instead master image
using `rook/ceph:master` tags take longer time
to pull and also, we should be using `rook/ceph:local-build`
tag for our ci. This will help canary raw test to be more stable.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-25 17:26:54 +05:30
Travis Nielsen 8e1090b3be Merge pull request #12624 from parth-gr/s5-cmd
object: fix s5cmd for s3 endpoint verification
2023-08-04 08:19:28 -06:00
parth-gr 8a3f058329 object: fix s5cmd for s3 endpoint verification
add a new toolbox yaml manifest which will use the
rook image instead of ceph image
for running s5 cmd container needs to run with rook image

closes: https://github.com/rook/rook/issues/12227

Signed-off-by: parth-gr <paarora@redhat.com>
2023-08-04 18:29:05 +05:30
subhamkrai b9c41fd548 ci: same ceph version in toolbox and cluster-test
Let's use same ceph version(latest Reef) in both
cluster-test and toolbox.yaml so that we don't need
to pull image twice. Alos, github action helper script
was calling `deploy_manifest_with_local_build` which
is required for operator.yaml and not for toolbox.yaml.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:11:30 +05:30
Jiffin Tony Thottan b48dc8a335 object: intial cosi driver controller design
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-07-18 22:49:41 +05:30
parth-gr ef4fa96e26 ci: catch error in ci for external rgw server canary test
currently we just print the error is anything fails in
script for rgw, So by that there is no panic or
failiure of script if something wrong happened,

Added Explicittly forcing to catch the error from
the error message that is thrown

Closes: https://github.com/rook/rook/issues/12244

Signed-off-by: parth-gr <paarora@redhat.com>
2023-05-31 18:42:52 +05:30
travisn 558732d0ae ci: canary test should enable the prometheus module
The canary test is waiting for the prometheus module,
which is now disabled by default. For the canary test,
we need to enable the prometheus module for the external
cluster test.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-04-03 15:32:11 -06:00
Rakshith R b9fc9a643e ci: add test for key rotation
This commit adds test for key rotation.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-03-08 12:07:26 +05:30
subhamkrai d983f3620e ci: fix multus ci
Signed-off-by: subhamkrai <srai@redhat.com>
2023-03-02 12:20:59 +05:30
Travis Nielsen 11e80b7047 Merge pull request #11436 from subhamkrai/bump-ci-k8s-version
ci: Update min K8s support to 1.21 and add 1.26 support
2023-02-14 20:26:56 -07:00
subhamkrai 1effb11a0a ci: bump min and max k8s support
this commit bump minmum k8s version to 1.21.14 and
max k8s version to latest 1.26.0. Keeping support for
most recent 6 versions.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-02-15 07:08:02 +05:30
Travis Nielsen eb4a727658 docs: remove mention of psps from documentation
The PSPs have long since been deprected. In K8s 1.21 the PSPs
were first deprecated, and support was completely removed
for them in 1.25. With Rook v1.11, the min supported version of
K8s is now 1.21. To reduce confusion in the documentation,
mention of the PSPs is now removed from the 1.11 docs.
For the corner case that users still require the PSPs,
the helm chart still contains the option for creating PSPs
or other users can still create the psp.yaml.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-13 13:30:32 -07:00
subhamkrai 45a10dcb85 core: upgrade to latest operator-sdk v1.25.0
upgrading to latest operator-sdk later version
and changing the way of csv generation.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-02-13 17:08:09 +05:30
Joshua Hoblitt 468f3f6f35 test: add secondary->master zone repl test to rgw-multisite-test
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2023-01-19 16:12:30 -07:00
Joshua Hoblitt 97bcc0a349 test: remove naughty words
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2023-01-19 11:13:12 -07:00
Shinya Hayashi 05875a3f4f osd: support loop devices for test clusters
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.

This feature is intended to be used for testing purposes only.

Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
2022-11-09 06:45:41 +00:00
Blaine Gardner 51fac0a993 nfs: fix nfs grace period when multus is enabled
Run the ganesha-rados-grace command in a remote pod when multus
networking is enabled.

Signed-off-by: parth-gr <paarora@redhat.com>
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-05 16:28:36 -06:00
Satoru Takeuchi 5c2d325cc4 ci: move to ubuntu 2004 completely
github-hosted runner ubuntu-18.04 will be removed by Dec. 1st, 2022.

https://github.blog/changelog/2022-08-09-github-actions-the-ubuntu-18-04-actions-runner-image-is-being-deprecated-and-will-be-removed-by-12-1-22/

So, let's change the all runnners to ubuntu-20.04.

We need to uninstall `snapd` because `snapd` prevents kernel from
unmounting `/mnt` completely.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-09-27 04:32:20 +00:00
parth-gr ddea97ebde nfs: fix nfs if multus is enabled
Closes: https://github.com/rook/rook/issues/10812
Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-21 14:09:40 +05:30
Satoru Takeuchi 47f7aefa20 ci: improve the log on intermittent build failure
The build process sometimes fails with intermittent problems. Some of them
are known problems and then we retry build process. However, there still
are unknown problems. In this case, it's hard to find the reason because
the build process exits immediately.

ref.
https://github.com/rook/rook/runs/8225144002?check_suite_focus=true#step:3:926

```
+ case "$o" in
+ exit 1
Error: Process completed with exit code 1.
```

To make debugging easier, let's print the output of `make`. This log won't be
too long since `make` prints most messages to stderr.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-09-07 10:48:21 +00:00
Rakshith R 07aac106df ci: add e2e for csi-nfsplugin restart
This commit adds e2e for csi-nfsplugin restart
when it is not on host-networking.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-09-02 14:16:15 +05:30