Commit Graph
141 Commits
Author SHA1 Message Date
subhamkrai 4eefad42e8 csi: move csi management to admin
Going forward, admin will manage the csi operator
CR's and rook will only manage Ceph Connection cr
and client Profile cr.

The old csi driver is completely removed from Rook
and can no longer be used starting in Rook v1.20.

The upgrade guide will contain the needed transition steps
for managing the csi operator settings.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-04-29 14:24:26 -06:00
subhamkrai b50f1a3d60 ci: fix some minor canary test issues
The canary test was waiting for replicapool
instead of replicapool2, and a more reliable
wait for the toolbox pod start is added.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-04-29 14:01:42 -06:00
Oded Viner ebfbb49f15 nvmeof: add csi operator support for nvmeof canary test
update the nvmeof minikube canary test to use the
ceph-csi operator instead of manually deploying
the provisioner and node-plugin.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-03-24 12:18:45 +01:00
Travis Nielsen 24053d04d3 Merge pull request #17017 from parth-gr/test-mirror-ec
ci: add ec pool mirroring in our ci
2026-03-03 10:24:26 -07:00
parth-gr 50f5eec61d ci: add ec pool mirroring in our ci
we do support ec pool mirroring too
Add a test for that

Signed-off-by: parth-gr <partharora1010@gmail.com>
2026-03-03 11:35:05 +05:30
Oded Viner 2ebb5d968a nvmeof: add nvmeof minikube canary test without csi operator
adds a new nvmeof minikube canary job for ceph v20.
the test deploys rook with csi operator disabled for nvmeof flow.
it validates pvc and pod io, gateway restart, and data persistence.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-02-18 13:22:50 +02:00
Satoru Takeuchi 0137274b58 test: create iscsi disk if extra disk does not exist
We found a new github runner that doesn't have an extra
disk mounted on /mnt and it causes massive amounts
of CI failures. We can overcome this circumstance
by creating an iSCSI disk as an extra disk.

ref. https://github.com/rook/rook/issues/16978#issue-3862612835

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2026-02-02 16:29:44 +00:00
subhamkrai 6c4d8205b6 ci: use minikube github action instead script
using github action for minikube will avoid
unwanted errors and probably will be more stable
than manual install. But both script and action
both uses almost same time for installation so
there we don't have preference.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-01-13 15:24:01 +05:30
subhamkrai a62a1392f8 ci: run daily nightly job in with gh arm runner
let's run daily nightly canary job with github action
arm runner as self-hosted runners is shutting down which
was provided by upstream user.

Signed-off-by: subhamkrai <srai@redhat.com>
2025-12-08 21:20:37 +05:30
Erik Sundell e7a704ccf1 helm: remove legacy PodSecurityPolicy resource
The helm charts allowed rendering a PodSecurityPolicy resource via the
configuration `pspEnable`. This option is removed and all references to
psp, PodSecurityPolicy, and Pod Security Policy have been cleaned up.

The PSP resource was only rendered if k8s version was lower than 1.25
when it was still supported. It has been deprecated since k8s 1.21.

Signed-off-by: Erik Sundell <erik@sundellopensource.se>
2025-09-30 18:25:35 +02:00
Michael Adam 605e820fc8 ci: move minikube to a better location
In the deb package's location of /usr/bin/minilube, a wrong version seems to be
reported but from /usr/local/bin it reports correctly.

Signed-off-by: Michael Adam <obnox@samba.org>
2025-09-16 18:24:18 +02:00
Michael AdamandTravis Nielsen 40427d5c0d ci: update latest k8s version to 1.34
this change updates the k8s version to 1.34 and also updates
the  cri-ctl version for minikube.

Signed-off-by: Michael Adam <obnox@samba.org>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2025-09-16 17:57:28 +02:00
subhamkraiandTravis Nielsen c57a47f774 ci: run csi-operator only in canary and upgrade suite
this commit add check to only run the csi-operator in
all the canary tests and upgrade suite only, other suite
like smoke and object will still test csi-driver.

Also, adding changes to make CI happy.

Signed-off-by: subhamkrai <srai@redhat.com>
Co-Authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2025-08-19 11:23:14 +05:30
Travis Nielsen 3137244409 ci: update min k8s version to v1.29
For Rook v1.18 the min supported version of K8s is
v1.29. With the pending release of K8s 1.34, this
will be the typical six most recent releases that
Rook tests against.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-08-04 16:51:00 -06:00
parth-gr c592d36fc7 external: automate external users cephx key rotation
Ceph has a new ceph auth rotate command currently
present in ceph:main

Add a new flag `--cephx-key-rotate`  to rotate the
cephx keys genrated by external python script,
If we enable it, it will create a new user with suffix `.{x}`

Signed-off-by: parth-gr <partharora1010@gmail.com>
2025-07-24 14:41:51 +05:30
Vilius Puškunalis 27010f96f3 osd: fix cleanup job disk shredding
Fixes quick disk shredding by using dd to shred data at additional offsets where ceph metadata is duplicated. For full shred, the shred utility remains in use.

Signed-off-by: Vilius Puškunalis <47086537+puskunalis@users.noreply.github.com>
2025-06-23 21:46:41 +03:00
subhamkrai 8c2f724f6d ci: update latest k8s version to 1.33
this commit update k8s version to 1.33 and also update
other version like for minikube cri-ctl and so.

Signed-off-by: subhamkrai <srai@redhat.com>
2025-04-29 20:53:47 +05:30
Travis Nielsen ae895770e2 Merge pull request #15750 from OdedViner/update_prometheus_version
test: update Prometheus Operator to v0.82.0
2025-04-23 11:44:14 -06:00
Oded Viner 5d143d990c test: update Prometheus Operator to v0.82.0
this PR updates the Prometheus Operator URL references from
version v0.71.1 to the latest release v0.81.0 in documentation
and integration test scripts. This ensures we are aligned with
the latest features and improvements from
the Prometheus Operator project.

Signed-off-by: Oded Viner <oviner@redhat.com>
2025-04-22 15:37:56 +03:00
Travis Nielsen 7d908d0a5e build: stop publishing charts in master branch
The helm charts will now only be published when it is
an officially tagged release build.

The images will only be published to all repos for
dockerhub, quay, and ghcr when it is a tagged release.

The images will be published only to dockerhub for all
master and interim release branch builds.

Remove obsolete makefile option for images.
Ceph is the only image Rook ever expects to build.
Simplify the makefile by removing the legacy option
to select which image to build.

Also included are other small improvements to clean up
the release scripts.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-04-18 10:40:29 -06:00
parth-gr 1a296dccd5 ci: fix rgw flaky ci test
s3cmd can take more time to put the data
of 1M to the bucket as it waits for
connection to get established

currently ci fails with, Retrying failed
request: /test1-1mib-test.dat ([Errno 111]
Connection refused)

also increase the timeout for creating objectstore

Signed-off-by: parth-gr <partharora1010@gmail.com>
2025-04-08 21:15:59 +05:30
subhamkrai 0dda801488 ci: test rgw multisite test
Signed-off-by: subhamkrai <srai@redhat.com>
2025-02-27 23:19:13 +05:30
Travis Nielsen 56c6659655 tests: canary tests to wait for first mon to start
Many of the canary tests have been failing much more
frequently in the past week or two. The test is typically
timing out pulling the image from quay.ceph.io since
it does not have as high bandwidth for the images.
A check is added to the test to wait specifically for the
first mon so it waits sufficiently for the image pull
before checking for other ceph daemons.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-12-20 13:53:21 -07:00
Travis Nielsen 7c2f41e72e build: add support for k8s 1.32
With the release of K8s 1.32, we update the CI and docs
to support this new release, to maintain the most recent
six releases of K8s.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-12-12 09:48:42 -07:00
Blaine Gardner 5539eedd1b multus: finish deprecating holder pods
Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-10-23 16:29:02 -06:00
Joshua Hoblitt 100c8bd7a9 test: require ceph tag param to replace_ceph_image
To prevent silent failures where the ceph image tag is not updated.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-18 12:08:57 -07:00
Joshua Hoblitt 11250013a0 test: add two-object-one-zone canary test
This acceptance test demonstrates the creation of two CephObjectStore(s)
that share the same pool(s) manually managed by CephBlockPool(s).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-14 16:54:39 -07:00
Joshua Hoblitt 92d9f994c2 test: improve reliability of canary rgw-multisite-testing
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-12 12:42:58 -07:00
Joshua Hoblitt d301114680 test: do not always run sudo lsblk in github-action-helper.sh
This removes the execution of `sudo lsblk` three times for every single
invocation of the script.  Usage of the BLOCK var is replaced with
functions which memoize the result of probing for block devices.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 16:40:10 -07:00
Joshua Hoblitt 581fd5c197 test: convert all github-action-helper functs to $REPO_DIR
This allows all functions to be called in any order without concern for
the CWD.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:18:17 -07:00
Joshua Hoblitt 31d55b90fe test: mv canary test specific resources out of deploy_cluster()
Factor out most of the CRs used by various canary tests to a new
deploy_cluster_full_of_cruft_please_stop_using_this() function.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:18:17 -07:00
Joshua Hoblitt ed8156017e test: add object-with-cephblockpool canary test
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:16:30 -07:00
Joshua Hoblitt cb1dd152ad test: add github-action-helper toolbox functions
Added these functions for running commands in the toolbox pod:

- toolbox()
- ceph()
- rbd()
- radosgw-admin()

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:16:30 -07:00
Xinliang Liu 05ac99c0c0 ci: fix canary-arm64 job
Fix OSD isn't up.
As sdb device might change to vdb in the runner, let
find_extra_block_dev() exclude the nbd devices and find the proper
extra device for OSD.

Clean up the nbd devices after the test job is running.

Fix logs artifact upload twice and collect logs before clean up.

Signed-off-by: Xinliang Liu <xinliang.liu@linaro.org>
2024-10-01 12:05:32 -06:00
subhamkrai f647444515 ci: fix ci permission issue with minikube start
this commit upgrade the minikube, k8s, crictl versions
in CI and also fix permission error in the github runner.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-09-19 15:55:18 +05:30
Travis Nielsen 8b60c52f31 build: generate the local build tag with docker io
The docker.io image prefix is expected to be prepended
to the image names in the test images. This was missed
in 14550 related to some CI tests, which was now causing
the CI failures in the 1.15 branch where the search and
replace was missing the new docker.io prefix.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 3045076db8)
2024-08-21 10:21:31 -06:00
subhamkrai 128ec16ee6 ci: add k8s 1.30 support in ci
adding support for kubernetes 1.30 in ci which released yesterday.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-04-19 07:41:05 +05:30
Travis Nielsen fdacfd51c5 object: create an object store based on shared pools
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.

Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-03-11 11:20:57 -06:00
Sunnatillo 3ad456c6a2 build: uplift prometheus operator version to v0.71.1
This commit uplifts prometheus operator version to v0.71.1.

Signed-off-by: Sunnatillo <sunnat.samadov@est.tech>
2024-02-27 20:00:25 +02:00
subhamkrai 135307a4df ci: upgrade min k8s supported version to 1.24.17
upgrading minimum kubernetes supported version to v1.24.17
and also upgrading other kubernetes version to their latest
respective version.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-02-13 22:11:44 +05:30
Liang Zheng ee05e82cea osd: support create osd with metadata partition
Currently, when rook provisions OSDs(in the OSD prepare job), rook effectively run a
c-v command such as the following.
```console
ceph-volume lvm batch --prepare <deviceA> <deviceB> <deviceC> --db-devices <metadataDevice>
```
but c-v lvm batch only supports disk and lvm, instead of disk partitions.

We can resort to `ceph-volume lvm prepare` to implement it.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2024-02-07 11:34:09 +08:00
sp98 45adae1931 ci: fix ci failure for multicluster tests
This PR fixes the failure while running multicluster mirroring CI tests

Signed-off-by: sp98 <sapillai@redhat.com>
2024-02-06 12:54:18 +05:30
Blaine Gardner feacb6424e ci: fix detection of GH actions extra disk
The 'extra' block device attached to GH actions runners has changed size
twice in 3 months. The previous strategy of detecting the disk by size
is becoming harder to maintain. Additionally, the block size with recent
changes (75G) is now the same as the boot device (also 75G), making the
method inexact.

The method can now be summarized as, "find the boot disk and choose the
disk that isn't the boot disk to be the 'extra' one used."

Prior to this, we used a one-liner based on `lsblk`. While we could
still make this a one-liner, the method is now updated to 2 effective
lines, plus debug text output to stderr to help if we need to debug
further in the future.

Of note, the 'extra' disk has a mount point of "/mnt", but it is unclear
whether this is a reliable heuristic for detecting the extra disk. For
years now, GH action runners have had only 2 disks. Therefore, it seems
slightly more likely that a heuristic to "choose the non-boot disk" will
be a more robust long-term solution.

If this strategy proves to be unreliable in the future, it may be wise
to consider whether "the device with a partition mounted to '/mnt'"
would be a good alternative.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-02-05 17:55:16 -07:00
subhamkraiandJan Klippel 6b0deb32d9 ci: disk in github action increased to 75G from 64G
the disk size in the github action machine has
increased from 64G to 75G. Now, we detech the version
automatically not fetching hard coded value.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2024-02-05 22:24:03 +05:30
Ryotaro Banno e9c1f81d79 ci: wait for all OSD pods to become Running in wait_for_prepare_pod
Currently, `wait_for_prepare_pod` only waits until 1 prepare pod and 1
OSD pod are in Running state.

So if `wait_for_ceph_to_be_ready` failed after `wait_for_prepare_pod`
succeeded, there are 4 possible cases:

1. some prepare pods didn't start correctly;
2. some prepare pods didn't finish correctly;
3. some OSD pods didn't start correctly; or
4. all prepare pods and OSD pods worked correctly, but some other
   process failed.

As far as I understand, the number of prepare pods on the GitHub CI is
always 1, so cases 1 and 2 are not problematic. However, it is hard to
distinguish the other two cases from the CI log.

To solve the above problems, this patch makes wait_for_prepare_pod wait
for all OSD pods to become running state. If wait_for_prepare_pod
timeouts before the OSD pods become running state, this will be a strong
indication that the OSD pods didn't start properly (i.e., case 3).

Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
2023-12-25 05:12:30 +00:00
Alexander Trost 4ed35d6bd6 operator: allow setting ceph config options via ceph cluster crd
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-12-02 00:05:48 +01:00
sp98 345f92961b ci: filter both 14 and 64 disks in CI
github ci runner is allocating both 14 and 64 Gb disks to the setup.
Rook CI currently hard codes the lsblk with a 14G filter.
This PR updates the filter to use either 14G or 64G disks.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-11-06 13:30:10 +05:30
subhamkrai e14f8c7710 ci: remove workaround for clean disk action
The issues are fixed from the clean disk action
repo and we shouldn't require any workaround. Also,
I'm not using git revert to revert the commit is earlier
clean disk action was mentioned in every CI which was
duplicate and now we have placed the action to composite yaml

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-29 19:17:01 +05:30
subhamkrai 3267576307 ci: fix github action free disk space
Since `google-cloud-sdk` is renamed with `google-cloud-cli`,
the github action is failing to remove the older name and this
is causing ci issues. Applying changes suggested by community to
manually remove some packages to resolve the issue.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-28 18:56:59 +05:30
Blaine Gardner 3e2e906ea7 test: update multus canary test
Update the multus canary test to reflect modern knowledge about how it
should be configured.

No longer test for the network device in OSD pods. Pods will utterly
fail to start if Multus is unable to attach interfaces.
Instead, look to the OSD map to test the connections more wholistically.
OSDs must have map IPs that include both public and cluster network.
This implicitly tests that the interfaces exist in the Pod, and it
additionally verifies other details, like Ceph `*_network` configs are
set propertly.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-09-18 17:15:23 -06:00