Commit Graph
400 Commits
Author SHA1 Message Date
Joshua Hoblitt d301114680 test: do not always run sudo lsblk in github-action-helper.sh
This removes the execution of `sudo lsblk` three times for every single
invocation of the script.  Usage of the BLOCK var is replaced with
functions which memoize the result of probing for block devices.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 16:40:10 -07:00
Joshua Hoblitt 581fd5c197 test: convert all github-action-helper functs to $REPO_DIR
This allows all functions to be called in any order without concern for
the CWD.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:18:17 -07:00
Joshua Hoblitt 31d55b90fe test: mv canary test specific resources out of deploy_cluster()
Factor out most of the CRs used by various canary tests to a new
deploy_cluster_full_of_cruft_please_stop_using_this() function.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:18:17 -07:00
Joshua Hoblitt ed8156017e test: add object-with-cephblockpool canary test
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:16:30 -07:00
Joshua Hoblitt cb1dd152ad test: add github-action-helper toolbox functions
Added these functions for running commands in the toolbox pod:

- toolbox()
- ceph()
- rbd()
- radosgw-admin()

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-10-03 10:16:30 -07:00
Xinliang Liu 05ac99c0c0 ci: fix canary-arm64 job
Fix OSD isn't up.
As sdb device might change to vdb in the runner, let
find_extra_block_dev() exclude the nbd devices and find the proper
extra device for OSD.

Clean up the nbd devices after the test job is running.

Fix logs artifact upload twice and collect logs before clean up.

Signed-off-by: Xinliang Liu <xinliang.liu@linaro.org>
2024-10-01 12:05:32 -06:00
subhamkrai f647444515 ci: fix ci permission issue with minikube start
this commit upgrade the minikube, k8s, crictl versions
in CI and also fix permission error in the github runner.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-09-19 15:55:18 +05:30
Praveen M afad40e404 csi: update csi-addons to v0.10.0
The csi-addons v0.10.0 release is now available.
Ref: https://github.com/csi-addons/kubernetes-csi-addons/releases/tag/v0.10.0

Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-09-18 12:19:46 +05:30
Michael Adam 93179a41f3 ci: slightly rework the docs-check workflow
This reworks the docs-check ci workflow in several ways:

* It renames the make target 'check-docs' to the  more systematic 'check.docs'.
* It adds a 'docs'mode to the  files validation script, and uses the script in `make check.docs`.

Overall, the workflow and local make targets are more systematic and
consistent with this change.

Signed-off-by: Michael Adam <obnox@samba.org>
2024-09-05 22:09:40 +02:00
Madhu Rajanna 05d579b607 csi: update csi-addons to v0.9.1
updating csi-addons to latest
v0.9.1 release.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-09-03 12:52:00 +02:00
Travis Nielsen 73391bbd89 Merge pull request #14629 from BlaineEXE/fix-multus-validation-test-default-service-account
multus: fix default service account handling
2024-08-22 10:31:20 -06:00
Blaine Gardner 58e3feaacf multus: fix default service account handling
The default service account isn't passed to the multus validation test
when no config file is used. Fix this.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-08-21 16:42:47 -06:00
Travis Nielsen 8b60c52f31 build: generate the local build tag with docker io
The docker.io image prefix is expected to be prepended
to the image names in the test images. This was missed
in 14550 related to some CI tests, which was now causing
the CI failures in the 1.15 branch where the search and
replace was missing the new docker.io prefix.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 3045076db8)
2024-08-21 10:21:31 -06:00
Madhu Rajanna 123025f22c csi: update csi-addons to v0.9.0
As we have new csi-addons v0.9.0
updating the same here as well.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-08-16 07:57:16 +02:00
Blaine Gardner 5773132d7f ci: fix failing multus validation tool test
Ceph image no longer has `ip` tool installed. Use a different container
image for the daemonset which sets host IPs and routes for multus hosts.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-08-14 12:25:44 -06:00
ee8bcad49d rgw: add support for keystone auth + swift/s3
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>

* extend the API object specs for swift and keystone integration

* adapt rgw to the new go-ceph version

  - The parameter lists of the API call have changes, as parameters
    ignored by the RGW Admin Ops API are no longer serialized, therefore
    the mock has to be adapted.

  - There is now validation for the user keys that are passed to the
    User get API, therefore things failed when we had empty keys in our
    User proxy object.

* expand the reconcile loop for the swift and keystone integration

* fix minor mistakes in design document

* add env var to pass extra args to minikube

  Minikube decides CPU cores and memory automatically based on the
  available resources on the machine which may be insufficient to
  run rook. This commit adds an environment variable to add arbitrary
  arguments to the minikube command, so both can be specified if
  desired.

* integration tests for swift and keystone

  The new integration of swift or s3 and keystone support by rook
  does not have any integration tests yet.

  This commit introduces integration tests for swift and keystone. The
  tests are done against a minimal keystone setup (keystone container
  image from Yaook-project (https://yaook.cloud), sqlite as database
  backend, cert-manager and trust-manager for test certificate setup).

  To prevent hardcoded credentials, passwords are generated
  by the tests. The integration tests use the openstack client
  (keystone- and swift-functionality) (https://docs.openstack.org/
  python-openstackclient/ latest/). This was a concious design decision
  to use client tooling as close as possible to the end user instead of
  using other go-libraries (such as gophercloud).

* add documentation on swift and keystone

  Currently there is no documentation on the use of Swift to access
  an object store as well as the use of OpenStack keystone for
  authentication.

  This commit adds documentation on the use of Swift and OpenStack
  keystone, as well as CRD-related documentation and an example setup.

* add integration tests for S3 via keystone

  This commit introduces integration tests for s3 and keystone. The
  tests are run against the same minimal keystone setup that the tests
  for swift and keystone use.

  The integration tests use the aws s3 client to use client tooling as
  close as possible to the end user instead of using other go-libraries.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
2024-08-08 14:26:21 +02:00
Blaine Gardner 33f5407dd4 multus: add host checking to validation tool
In order to help users check that they have implemented the newly-added
Multus host configuration prerequisites, add a check to the validation
tool to verify connectivity.

Because users who are already running clusters with Multus enabled, add
a flag that allows users to only check for host configuration
prerequisites. This mode will not start the large number of clients that
would normally be started because those clients could disrupt a running
Rook cluster negatively.

Host checking pods require host network access. Many Kubernetes
distributions have pod security features enabled. In order to allow
non-Vanilla distros to run this tool, allow specifying a service account
that pods will run as, which can be configured by the admin to allow
test pods.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-07-10 13:41:38 -06:00
Madhu Rajanna 266fd49c02 csi: update csi-addons repo link
updating csi-addons repo link
to pull the yamls for installation

closes: #14394

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-07-02 17:36:24 +02:00
Blaine Gardner 08a27940ed multus: add and test ipv6 support for validation tool
Add IPv6 support for multus validation tool. Also test that IPv6 support
works by specifying one of the NetAttachDefs with an IPv6 address range.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-06-03 17:30:36 -06:00
Blaine Gardner c9d99e01a0 ci: use markdownlint to enforce mkdocs compatibility
mkdocs uses a markdown renderer that is hardcoded to 4 spaces per tab
for detecting indentation levels, including ordered- and
unordered-lists. Since we cannot easily change the renderer, begin using
a markdown linter in CI that will fail if official docs do not adhere to
the spacing rules.

As a starting point, the markdownlint config does not begin with the
default set of checks, which might overwhelm attempts to fix them.
Instead, focus on list-tab-spacing rules and a few other highly useful
checks.

markdownlint also has some gaps in its abilities that allow common Rook
doc issues to pass acceptance. However, it allows creating custom
linting plugins. Create 2 such linting plugins to check 2 things:

- all doc lines (except code blocks) must be aligned to a 4-space
  boundary, without exception. This ensures that markdown will render
  correctly with mkdocs. This unfortunately makes it possible to create
  lists that are internally aligned strangely.
- admonitions must all follow the same format of
  ```
  !!! header
      body
  ```

For the strange lists, this is allowed and renders correctly, but it
looks strange:

```md
- first bullet
- second bullet
    still second bullet
- third bullet

    has a paragraph
    of text inside

- last bullet

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-04-29 17:25:11 -06:00
subhamkrai 128ec16ee6 ci: add k8s 1.30 support in ci
adding support for kubernetes 1.30 in ci which released yesterday.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-04-19 07:41:05 +05:30
Travis Nielsen fdacfd51c5 object: create an object store based on shared pools
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.

Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-03-11 11:20:57 -06:00
Sunnatillo 3ad456c6a2 build: uplift prometheus operator version to v0.71.1
This commit uplifts prometheus operator version to v0.71.1.

Signed-off-by: Sunnatillo <sunnat.samadov@est.tech>
2024-02-27 20:00:25 +02:00
Blaine Gardner 1f65d84a8e Merge pull request #13755 from BlaineEXE/ci-fix-nightly-ids
ci: allow canary jobs to have ids for nightly suite
2024-02-14 11:19:28 -07:00
Blaine Gardner 3e76de7868 ci: allow canary jobs to have ids for nightly suite
For nightly jobs, canary tests are all running with the same job IDs,
making the last-run jobs cancel previous runs. Add a workflow-id
parameter to canary jobs that can be used to give each job in a canary
suite a unique job ID to prevent this.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-02-14 11:04:40 -07:00
subhamkrai 135307a4df ci: upgrade min k8s supported version to 1.24.17
upgrading minimum kubernetes supported version to v1.24.17
and also upgrading other kubernetes version to their latest
respective version.

Signed-off-by: subhamkrai <srai@redhat.com>
2024-02-13 22:11:44 +05:30
Liang Zheng ee05e82cea osd: support create osd with metadata partition
Currently, when rook provisions OSDs(in the OSD prepare job), rook effectively run a
c-v command such as the following.
```console
ceph-volume lvm batch --prepare <deviceA> <deviceB> <deviceC> --db-devices <metadataDevice>
```
but c-v lvm batch only supports disk and lvm, instead of disk partitions.

We can resort to `ceph-volume lvm prepare` to implement it.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2024-02-07 11:34:09 +08:00
sp98 45adae1931 ci: fix ci failure for multicluster tests
This PR fixes the failure while running multicluster mirroring CI tests

Signed-off-by: sp98 <sapillai@redhat.com>
2024-02-06 12:54:18 +05:30
Blaine Gardner feacb6424e ci: fix detection of GH actions extra disk
The 'extra' block device attached to GH actions runners has changed size
twice in 3 months. The previous strategy of detecting the disk by size
is becoming harder to maintain. Additionally, the block size with recent
changes (75G) is now the same as the boot device (also 75G), making the
method inexact.

The method can now be summarized as, "find the boot disk and choose the
disk that isn't the boot disk to be the 'extra' one used."

Prior to this, we used a one-liner based on `lsblk`. While we could
still make this a one-liner, the method is now updated to 2 effective
lines, plus debug text output to stderr to help if we need to debug
further in the future.

Of note, the 'extra' disk has a mount point of "/mnt", but it is unclear
whether this is a reliable heuristic for detecting the extra disk. For
years now, GH action runners have had only 2 disks. Therefore, it seems
slightly more likely that a heuristic to "choose the non-boot disk" will
be a more robust long-term solution.

If this strategy proves to be unreliable in the future, it may be wise
to consider whether "the device with a partition mounted to '/mnt'"
would be a good alternative.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-02-05 17:55:16 -07:00
Travis Nielsen 577c7cc636 Merge pull request #13675 from subhamkrai/ci-fix-disk-size
ci: disk in github action increased to 75G from 64G
2024-02-05 10:28:59 -07:00
subhamkraiandJan Klippel 6b0deb32d9 ci: disk in github action increased to 75G from 64G
the disk size in the github action machine has
increased from 64G to 75G. Now, we detech the version
automatically not fetching hard coded value.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2024-02-05 22:24:03 +05:30
parth-gr 75fcffd08c ci: reformat the python scripts
currently there is a new release of black
formatter so updating it with new release

Closes: https://github.com/rook/rook/issues/13630

Signe-off-by: parth-gr <partharora1010@gmail.com>
Signed-off-by: parth-gr <paarora@redhat.com>
2024-01-30 20:17:49 +05:30
Madhu Rajanna c35a8532aa csi: option to customize csi driver name prefix
For now we are using the operator namespace name
as the prefix for the csi driver, This PR provides
an option for the users if someone wants to have
their own prefix for the csi driver, if someone tries
to change the prefix for existing csi driver rook
operator will fail to reconcile the csi driver.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-01-29 10:54:25 +01:00
Redouane Kachach 605710ee0f test: adding new variables for minikube parameters configuration
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2024-01-24 13:39:45 +01:00
Redouane Kachach f9ffa872d6 test: fixing examples directory path
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2024-01-12 11:52:40 +01:00
Redouane Kachach ddd852d767 test: improving the handling of arguments within the script
so far all the arguments were being handled as separate getopts
flags introducing a lot of boilerplate code for every new flag.
The new approach improves the arguments handling by using env
variables.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2024-01-12 11:47:37 +01:00
Blaine Gardner cd253d0218 multus: use nginx-unprivileged image from quay
Use Quay as the source for the nginx-unprivileged image used for the
Multus validation tool because Quay does not rate limit image pulls,
which are a common complaint for users of the tool.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-01-04 14:25:40 -07:00
Travis Nielsen c10fc802e3 Merge pull request #13453 from iPraveenParihar/test/csiaddons-validation
ci: add validation for csi-addons sidecar
2024-01-04 10:04:14 -07:00
Praveen M e9395d052f ci: add validation for csi-addons sidecar
Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-01-04 20:11:53 +05:30
Ryotaro Banno e9c1f81d79 ci: wait for all OSD pods to become Running in wait_for_prepare_pod
Currently, `wait_for_prepare_pod` only waits until 1 prepare pod and 1
OSD pod are in Running state.

So if `wait_for_ceph_to_be_ready` failed after `wait_for_prepare_pod`
succeeded, there are 4 possible cases:

1. some prepare pods didn't start correctly;
2. some prepare pods didn't finish correctly;
3. some OSD pods didn't start correctly; or
4. all prepare pods and OSD pods worked correctly, but some other
   process failed.

As far as I understand, the number of prepare pods on the GitHub CI is
always 1, so cases 1 and 2 are not problematic. However, it is hard to
distinguish the other two cases from the CI log.

To solve the above problems, this patch makes wait_for_prepare_pod wait
for all OSD pods to become running state. If wait_for_prepare_pod
timeouts before the OSD pods become running state, this will be a strong
indication that the OSD pods didn't start properly (i.e., case 3).

Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
2023-12-25 05:12:30 +00:00
subhamkrai 2d3e561a00 ci: improve collect-log script
I was debugging helm tests, I noticed, the operator namespace
content is empty and operator is created in same namespace as
other pods. Also, let's collect the 'kube-system' namespace logs
in multus test only as that is the only test where we need to debug
cluster networking and require kube-system logs.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-19 21:08:06 +05:30
subhamkrai cb602fe33c ci: add 'rook-ceph-seconday' ns to collect logs
In canary test `multi-cluster-mirroring` we also need to collect
logs of `rook-ceph-secondary` namespace to debug the ci.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-13 15:08:07 +05:30
Travis Nielsen 951d7ad6ec Merge pull request #13298 from rkachach/fix_issue_add_cluster_file_opt
Adding new option -i to specify the cluster spec file + some code refactoring
2023-12-07 08:44:55 -07:00
Alexander Trost 3097455d78 Merge pull request #13246 from koor-tech/ceph_config_via_cluster_crd_impl
operator: allow setting ceph config options via ceph cluster crd
2023-12-02 11:27:20 +01:00
Alexander Trost 4ed35d6bd6 operator: allow setting ceph config options via ceph cluster crd
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-12-02 00:05:48 +01:00
Blaine Gardner f4d67fd4db Merge pull request #13261 from subhamkrai/remove-controller-runtime
core: remove webhook & controller-runtime from apis
2023-12-01 09:55:34 -07:00
subhamkrai 28cc1ebc55 core: remove webhook & controller-runtime from apis
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-01 14:15:40 +05:30
Redouane Kachach e122c01708 test: fixing namespace to get dashboard endpoint
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-30 17:29:34 +01:00
Redouane Kachach 1d5b49eff6 test: adding a new option -i to specify which cluster spec file
at this moment we have two different cluster spec files for testing
cluster-test.yaml and cluster-on-pvc-minikube.yaml. With the new
option user can choose which one to use to bootstrap the cluster

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-30 17:27:36 +01:00
Redouane Kachach f576742775 test: moving parameters default to init_vars function
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-30 17:27:36 +01:00