Going forward, admin will manage the csi operator
CR's and rook will only manage Ceph Connection cr
and client Profile cr.
The old csi driver is completely removed from Rook
and can no longer be used starting in Rook v1.20.
The upgrade guide will contain the needed transition steps
for managing the csi operator settings.
Signed-off-by: subhamkrai <srai@redhat.com>
The canary test was waiting for replicapool
instead of replicapool2, and a more reliable
wait for the toolbox pod start is added.
Signed-off-by: subhamkrai <srai@redhat.com>
update the nvmeof minikube canary test to use the
ceph-csi operator instead of manually deploying
the provisioner and node-plugin.
Signed-off-by: Oded Viner <oviner@redhat.com>
adds a new nvmeof minikube canary job for ceph v20.
the test deploys rook with csi operator disabled for nvmeof flow.
it validates pvc and pod io, gateway restart, and data persistence.
Signed-off-by: Oded Viner <oviner@redhat.com>
using github action for minikube will avoid
unwanted errors and probably will be more stable
than manual install. But both script and action
both uses almost same time for installation so
there we don't have preference.
Signed-off-by: subhamkrai <srai@redhat.com>
let's run daily nightly canary job with github action
arm runner as self-hosted runners is shutting down which
was provided by upstream user.
Signed-off-by: subhamkrai <srai@redhat.com>
The helm charts allowed rendering a PodSecurityPolicy resource via the
configuration `pspEnable`. This option is removed and all references to
psp, PodSecurityPolicy, and Pod Security Policy have been cleaned up.
The PSP resource was only rendered if k8s version was lower than 1.25
when it was still supported. It has been deprecated since k8s 1.21.
Signed-off-by: Erik Sundell <erik@sundellopensource.se>
In the deb package's location of /usr/bin/minilube, a wrong version seems to be
reported but from /usr/local/bin it reports correctly.
Signed-off-by: Michael Adam <obnox@samba.org>
this change updates the k8s version to 1.34 and also updates
the cri-ctl version for minikube.
Signed-off-by: Michael Adam <obnox@samba.org>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
this commit add check to only run the csi-operator in
all the canary tests and upgrade suite only, other suite
like smoke and object will still test csi-driver.
Also, adding changes to make CI happy.
Signed-off-by: subhamkrai <srai@redhat.com>
Co-Authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
For Rook v1.18 the min supported version of K8s is
v1.29. With the pending release of K8s 1.34, this
will be the typical six most recent releases that
Rook tests against.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Ceph has a new ceph auth rotate command currently
present in ceph:main
Add a new flag `--cephx-key-rotate` to rotate the
cephx keys genrated by external python script,
If we enable it, it will create a new user with suffix `.{x}`
Signed-off-by: parth-gr <partharora1010@gmail.com>
Fixes quick disk shredding by using dd to shred data at additional offsets where ceph metadata is duplicated. For full shred, the shred utility remains in use.
Signed-off-by: Vilius Puškunalis <47086537+puskunalis@users.noreply.github.com>
this PR updates the Prometheus Operator URL references from
version v0.71.1 to the latest release v0.81.0 in documentation
and integration test scripts. This ensures we are aligned with
the latest features and improvements from
the Prometheus Operator project.
Signed-off-by: Oded Viner <oviner@redhat.com>
The helm charts will now only be published when it is
an officially tagged release build.
The images will only be published to all repos for
dockerhub, quay, and ghcr when it is a tagged release.
The images will be published only to dockerhub for all
master and interim release branch builds.
Remove obsolete makefile option for images.
Ceph is the only image Rook ever expects to build.
Simplify the makefile by removing the legacy option
to select which image to build.
Also included are other small improvements to clean up
the release scripts.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
s3cmd can take more time to put the data
of 1M to the bucket as it waits for
connection to get established
currently ci fails with, Retrying failed
request: /test1-1mib-test.dat ([Errno 111]
Connection refused)
also increase the timeout for creating objectstore
Signed-off-by: parth-gr <partharora1010@gmail.com>
Many of the canary tests have been failing much more
frequently in the past week or two. The test is typically
timing out pulling the image from quay.ceph.io since
it does not have as high bandwidth for the images.
A check is added to the test to wait specifically for the
first mon so it waits sufficiently for the image pull
before checking for other ceph daemons.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the release of K8s 1.32, we update the CI and docs
to support this new release, to maintain the most recent
six releases of K8s.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
This acceptance test demonstrates the creation of two CephObjectStore(s)
that share the same pool(s) manually managed by CephBlockPool(s).
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
This removes the execution of `sudo lsblk` three times for every single
invocation of the script. Usage of the BLOCK var is replaced with
functions which memoize the result of probing for block devices.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Factor out most of the CRs used by various canary tests to a new
deploy_cluster_full_of_cruft_please_stop_using_this() function.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Fix OSD isn't up.
As sdb device might change to vdb in the runner, let
find_extra_block_dev() exclude the nbd devices and find the proper
extra device for OSD.
Clean up the nbd devices after the test job is running.
Fix logs artifact upload twice and collect logs before clean up.
Signed-off-by: Xinliang Liu <xinliang.liu@linaro.org>
this commit upgrade the minikube, k8s, crictl versions
in CI and also fix permission error in the github runner.
Signed-off-by: subhamkrai <srai@redhat.com>
The docker.io image prefix is expected to be prepended
to the image names in the test images. This was missed
in 14550 related to some CI tests, which was now causing
the CI failures in the 1.15 branch where the search and
replace was missing the new docker.io prefix.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 3045076db8)
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.
Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
upgrading minimum kubernetes supported version to v1.24.17
and also upgrading other kubernetes version to their latest
respective version.
Signed-off-by: subhamkrai <srai@redhat.com>
Currently, when rook provisions OSDs(in the OSD prepare job), rook effectively run a
c-v command such as the following.
```console
ceph-volume lvm batch --prepare <deviceA> <deviceB> <deviceC> --db-devices <metadataDevice>
```
but c-v lvm batch only supports disk and lvm, instead of disk partitions.
We can resort to `ceph-volume lvm prepare` to implement it.
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
The 'extra' block device attached to GH actions runners has changed size
twice in 3 months. The previous strategy of detecting the disk by size
is becoming harder to maintain. Additionally, the block size with recent
changes (75G) is now the same as the boot device (also 75G), making the
method inexact.
The method can now be summarized as, "find the boot disk and choose the
disk that isn't the boot disk to be the 'extra' one used."
Prior to this, we used a one-liner based on `lsblk`. While we could
still make this a one-liner, the method is now updated to 2 effective
lines, plus debug text output to stderr to help if we need to debug
further in the future.
Of note, the 'extra' disk has a mount point of "/mnt", but it is unclear
whether this is a reliable heuristic for detecting the extra disk. For
years now, GH action runners have had only 2 disks. Therefore, it seems
slightly more likely that a heuristic to "choose the non-boot disk" will
be a more robust long-term solution.
If this strategy proves to be unreliable in the future, it may be wise
to consider whether "the device with a partition mounted to '/mnt'"
would be a good alternative.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
the disk size in the github action machine has
increased from 64G to 75G. Now, we detech the version
automatically not fetching hard coded value.
Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: subhamkrai <srai@redhat.com>
Currently, `wait_for_prepare_pod` only waits until 1 prepare pod and 1
OSD pod are in Running state.
So if `wait_for_ceph_to_be_ready` failed after `wait_for_prepare_pod`
succeeded, there are 4 possible cases:
1. some prepare pods didn't start correctly;
2. some prepare pods didn't finish correctly;
3. some OSD pods didn't start correctly; or
4. all prepare pods and OSD pods worked correctly, but some other
process failed.
As far as I understand, the number of prepare pods on the GitHub CI is
always 1, so cases 1 and 2 are not problematic. However, it is hard to
distinguish the other two cases from the CI log.
To solve the above problems, this patch makes wait_for_prepare_pod wait
for all OSD pods to become running state. If wait_for_prepare_pod
timeouts before the OSD pods become running state, this will be a strong
indication that the OSD pods didn't start properly (i.e., case 3).
Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
github ci runner is allocating both 14 and 64 Gb disks to the setup.
Rook CI currently hard codes the lsblk with a 14G filter.
This PR updates the filter to use either 14G or 64G disks.
Signed-off-by: sp98 <sapillai@redhat.com>
The issues are fixed from the clean disk action
repo and we shouldn't require any workaround. Also,
I'm not using git revert to revert the commit is earlier
clean disk action was mentioned in every CI which was
duplicate and now we have placed the action to composite yaml
Signed-off-by: subhamkrai <srai@redhat.com>
Since `google-cloud-sdk` is renamed with `google-cloud-cli`,
the github action is failing to remove the older name and this
is causing ci issues. Applying changes suggested by community to
manually remove some packages to resolve the issue.
Signed-off-by: subhamkrai <srai@redhat.com>
Update the multus canary test to reflect modern knowledge about how it
should be configured.
No longer test for the network device in OSD pods. Pods will utterly
fail to start if Multus is unable to attach interfaces.
Instead, look to the OSD map to test the connections more wholistically.
OSDs must have map IPs that include both public and cluster network.
This implicitly tests that the interfaces exist in the Pod, and it
additionally verifies other details, like Ceph `*_network` configs are
set propertly.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>