Commit Graph
360 Commits
Author SHA1 Message Date
subhamkrai 2d3e561a00 ci: improve collect-log script
I was debugging helm tests, I noticed, the operator namespace
content is empty and operator is created in same namespace as
other pods. Also, let's collect the 'kube-system' namespace logs
in multus test only as that is the only test where we need to debug
cluster networking and require kube-system logs.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-19 21:08:06 +05:30
subhamkrai cb602fe33c ci: add 'rook-ceph-seconday' ns to collect logs
In canary test `multi-cluster-mirroring` we also need to collect
logs of `rook-ceph-secondary` namespace to debug the ci.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-13 15:08:07 +05:30
Travis Nielsen 951d7ad6ec Merge pull request #13298 from rkachach/fix_issue_add_cluster_file_opt
Adding new option -i to specify the cluster spec file + some code refactoring
2023-12-07 08:44:55 -07:00
Alexander Trost 3097455d78 Merge pull request #13246 from koor-tech/ceph_config_via_cluster_crd_impl
operator: allow setting ceph config options via ceph cluster crd
2023-12-02 11:27:20 +01:00
Alexander Trost 4ed35d6bd6 operator: allow setting ceph config options via ceph cluster crd
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-12-02 00:05:48 +01:00
Blaine Gardner f4d67fd4db Merge pull request #13261 from subhamkrai/remove-controller-runtime
core: remove webhook & controller-runtime from apis
2023-12-01 09:55:34 -07:00
subhamkrai 28cc1ebc55 core: remove webhook & controller-runtime from apis
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-01 14:15:40 +05:30
Redouane Kachach e122c01708 test: fixing namespace to get dashboard endpoint
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-30 17:29:34 +01:00
Redouane Kachach 1d5b49eff6 test: adding a new option -i to specify which cluster spec file
at this moment we have two different cluster spec files for testing
cluster-test.yaml and cluster-on-pvc-minikube.yaml. With the new
option user can choose which one to use to bootstrap the cluster

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-30 17:27:36 +01:00
Redouane Kachach f576742775 test: moving parameters default to init_vars function
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-30 17:27:36 +01:00
Redouane Kachach b02a2147bc test: improve examples directory checking
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-30 17:27:35 +01:00
Redouane Kachach b739f29bc8 test: fix how we obtain the dashboard endpoint
current code gets the ip:port for the dashboard by using
the ip of the mgr pods. This works great when there's only
one mgr but it fails in case of a multi-node cluster

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-30 11:26:48 +01:00
Travis Nielsen 94060b3433 Merge pull request #13238 from rkachach/fix_issue_adding_ns_support
test: adding support to specify the operator and cluster namepaces
2023-11-27 10:02:43 -07:00
Redouane Kachach 734d105790 test: applying adding rbac.yaml when enabling monitoring
applying adding rbac.yaml when enabling monitoring to avoid
permission issues when accessing servicemonitors

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-23 16:38:12 +01:00
Redouane Kachach f60cdde460 test: increasing rook-ceph-tools deployment timeout to 90s
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-23 16:37:01 +01:00
Redouane Kachach f93532a8e4 test: adding support to specify the operator and cluster namepaces
adding support to specify operator and cluster namespaces through new
command-line arguments: -c for cluster and -o for operator. The script
will automatically update the namespace in the corresponding YAML files

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-23 13:48:06 +01:00
Alexander Trost fcf52207f8 build: use /usr/bin/env to look up script interpreters
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-11-21 14:38:22 +01:00
Michael Adam 82c329c6af tests: create-dev-cluster: unify handling of invocation errors
This change to the create-dev-cluster script is intended to make the
handling of invalid invocations more uniform
examples are unknown options and options requiring an argument specified
without one.

The unification is achieved by encapsulating the corresponding code in a
function.

Signed-off-by: Michael Adam <obnox@samba.org>
2023-11-13 19:19:52 +01:00
Redouane Kachach b9501bdddb test: adding support for multi-rook clusters creation
This change introduce a new argument -p to specify the minikube
profile for the new cluster. This way we can have multiple rook
clusters running on the same machine.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-08 13:29:46 +01:00
Redouane Kachach 1d0fb2f5a1 test: adding a new variable for minikube command and its args
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-07 12:54:36 +01:00
Redouane Kachach 66f48e3777 test: using minikube profiles for rook cluster creation
Let's use the minikube profiles feature to set up a unique profile
exclusively for the Rook cluster. By doing this, we make sure that we
don't run into any issues with other minikube profiles that users
might already have in their local setup.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-07 00:55:38 +01:00
Travis Nielsen 68625bc737 Merge pull request #13155 from obnoxxx/improve-create-dev-cluster-script
Improve the create-dev-cluster script
2023-11-06 10:32:43 -07:00
Michael Adam f6068159af tests: improve create-dev-cluster script to find examples dir
The new create-dev-cluster script errored out for me, not finding the
deploy/examples directory.

This change fixes that problem by making the path relative to the
script's directory, so that it works in a rook code tree from where it
is expected to be run.

Signed-off-by: Michael Adam <obnox@samba.org>
2023-11-06 18:28:05 +01:00
sp98 345f92961b ci: filter both 14 and 64 disks in CI
github ci runner is allocating both 14 and 64 Gb disks to the setup.
Rook CI currently hard codes the lsblk with a 14G filter.
This PR updates the filter to use either 14G or 64G disks.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-11-06 13:30:10 +05:30
Redouane Kachach 5c7951f92b test: adding support for monitoring in dev-create-cluster script
Adding a new '-m' option for activating monitoring during
cluster setup. With this option turned on, the script will handle the
installation of the monitoring stack and configure the dashboard to
connect to the newly installed Prometheus server automatically.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-03 15:14:21 +01:00
Sheetal Pamecha 06c176524a multus: improve the multus validation test's flakiness metric
Allow the flakiness threshold window to be tuned from the cli

Signed-off-by: Sheetal Pamecha <spamecha@redhat.com>
2023-10-31 15:23:10 +05:30
Redouane Kachach 73b7e649d7 test: adding a script to simplify test cluster creation
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-10-30 16:48:49 +01:00
Blaine Gardner 0c721e05d5 multus: allow node profiles in validation test
Add the ability to specify node profiles in the multus validation test.

This addresses a few points of early feedback on the validation tool.
Statements below critique the tool's behavior before this patch.
1. The tool assumes all daemons are on public and cluster network, which
   means users who have a significantly smaller cluster net (a
   design choice) cannot run a single test to determine if Rook is
   likely to install correctly.
2. The tool does not have placement options to select only a subset of
   Kubernetes nodes to run validation on.
3. Users of multus seem to have a dedicated pool of storage nodes more
   often than the average Rook install. This makes sense for security-
   and perforance-minded users. The tool cannot run a single test to
   verify storage-only and general-workload nodes at one time.

These points are addressed by allowing users to specify configurations
for different "NodeTypes."

Each NodeType config has options for selecting the number of OSDs as
well as the number of other (non-OSD) Ceph daemons. This limits the
unnecessary exhaustion of cluster network addresses from critique 1.

Each NodeType config has its own placement (critique 2).

Users can define as many NodeTypes as needed to test the network for
their planned CephCluster. Specifically, this allows the tool to test
storage-only nodes and generalized-workload nodes at the same time. An
arbitrary number of NodeTypes are allowed to support even more highly
specialized cluster setups, such as multiple tiers of storage nodes
where some storage-only nodes may run more OSDs than others.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-10-04 15:26:35 -06:00
subhamkrai e14f8c7710 ci: remove workaround for clean disk action
The issues are fixed from the clean disk action
repo and we shouldn't require any workaround. Also,
I'm not using git revert to revert the commit is earlier
clean disk action was mentioned in every CI which was
duplicate and now we have placed the action to composite yaml

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-29 19:17:01 +05:30
subhamkrai 3267576307 ci: fix github action free disk space
Since `google-cloud-sdk` is renamed with `google-cloud-cli`,
the github action is failing to remove the older name and this
is causing ci issues. Applying changes suggested by community to
manually remove some packages to resolve the issue.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-28 18:56:59 +05:30
Blaine Gardner 3e2e906ea7 test: update multus canary test
Update the multus canary test to reflect modern knowledge about how it
should be configured.

No longer test for the network device in OSD pods. Pods will utterly
fail to start if Multus is unable to attach interfaces.
Instead, look to the OSD map to test the connections more wholistically.
OSDs must have map IPs that include both public and cluster network.
This implicitly tests that the interfaces exist in the Pod, and it
additionally verifies other details, like Ceph `*_network` configs are
set propertly.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-09-18 17:15:23 -06:00
subhamkrai 90d81e11b6 ci: upgrade the kubernetes versions for ci
since, kubernetes v1.28.0 release upgrading the
latest k8s version to v.1.28.0 excepth objectSuite
tes.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-28 10:35:44 +05:30
subhamkrai ad50c8e780 ci: use local tag instead master image
using `rook/ceph:master` tags take longer time
to pull and also, we should be using `rook/ceph:local-build`
tag for our ci. This will help canary raw test to be more stable.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-25 17:26:54 +05:30
Travis Nielsen 8e1090b3be Merge pull request #12624 from parth-gr/s5-cmd
object: fix s5cmd for s3 endpoint verification
2023-08-04 08:19:28 -06:00
parth-gr 8a3f058329 object: fix s5cmd for s3 endpoint verification
add a new toolbox yaml manifest which will use the
rook image instead of ceph image
for running s5 cmd container needs to run with rook image

closes: https://github.com/rook/rook/issues/12227

Signed-off-by: parth-gr <paarora@redhat.com>
2023-08-04 18:29:05 +05:30
subhamkrai 4502ea8ee1 multus: use right interface in ci validation
runner version `2.306` had the interface `net`
but somehow version `2.307.1` which is latest
doesn't have `net` it has `eth0*` so using that.

```
cat /proc/net/dev
Inter-|   Receive                                                |  Transmit
 face |bytes    packets errs drop fifo frame compressed multicast|bytes    packets errs drop fifo colls carrier compressed
    lo:       0       0    0    0    0     0          0         0        0       0    0    0    0     0       0          0
 tunl0:       0       0    0    0    0     0          0         0        0       0    0    0    0     0       0          0
  eth0:     446       5    0    0    0     0          0         0        0       0    0    0    0     0       0          0
sh-4.4# grep etho /proc/net/dev
sh-4.4# grep eth0 /proc/net/dev
  eth0:     446       5    0    0    0     0          0         0        0       0    0    0    0     0       0          0
sh-4.4# exit
```

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:11:34 +05:30
subhamkrai b9c41fd548 ci: same ceph version in toolbox and cluster-test
Let's use same ceph version(latest Reef) in both
cluster-test and toolbox.yaml so that we don't need
to pull image twice. Alos, github action helper script
was calling `deploy_manifest_with_local_build` which
is required for operator.yaml and not for toolbox.yaml.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:11:30 +05:30
subhamkrai 0ed91458dc test: collect kube-system logs for debugging
added kube-system namespace also to collect
logs from to debug the smoke suite issue

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-26 14:09:15 +05:30
Jiffin Tony Thottan b48dc8a335 object: intial cosi driver controller design
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-07-18 22:49:41 +05:30
subhamkrai 70f29be748 ci: fix ci test encryption-pvc-kms-vault-token-auth
we need to wait for the rgw pod to be delete and not
only the cephobjectsore, sometime the pod could be in
terminating state. Also, in some place it require proper
command to wait for pod to be ready/delete and get the
pod name only.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-05 14:49:18 +05:30
Travis Nielsen 3b4ba859fd Merge pull request #12267 from parth-gr/fix-rgw-ci
ci: catch error in ci for external rgw server canary test
2023-05-31 12:42:12 -06:00
parth-gr ef4fa96e26 ci: catch error in ci for external rgw server canary test
currently we just print the error is anything fails in
script for rgw, So by that there is no panic or
failiure of script if something wrong happened,

Added Explicittly forcing to catch the error from
the error message that is thrown

Closes: https://github.com/rook/rook/issues/12244

Signed-off-by: parth-gr <paarora@redhat.com>
2023-05-31 18:42:52 +05:30
Blaine Gardner 2eb5a9a3f6 test: add CI e2e test for multus validation test
Add a CI e2e test for the multus validation routine that runs whenever
the multus validation test is modified and on master/releases.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2023-05-26 12:19:54 -06:00
Blaine Gardner e58b09c01c test: add tmate pod manifest
Add a tmate manifest that can be added to CI tests to allow manually
debugging them in real-time while the test is ongoing.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2023-05-19 16:59:23 -06:00
Blaine Gardner 0f6e7ee921 test: add multus validation test routine to rook binary
Add a more involved multus validation test to the Rook binary. Because
this is intended to be end-user runnable, make sure operator-only
commands are hidden.

Build this into the rook binary instead of creating a separate binary
for ease, and because any binary built with the kube api becomes 40+
megabytes. We save quite a bit of space by including this in the Rook
binary, which is good for keeping container layers as small as possible.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2023-05-02 10:23:00 -06:00
Travis Nielsen 8ea26a72ed Merge pull request #11789 from thotz/test-failure-encryption-pvc-kms-vault-token
test: check rgw pod is running for canary github workflow
2023-04-06 10:55:55 -06:00
Jiffin Tony Thottan 715c89bcd8 test: check rgw pod is running for canary github workflow
For the RGW daemon validation please check whether pod is Running than
the exisitng checks

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-04-06 16:28:10 +05:30
Alexander Trost 99f5c9c713 build: fix publish docs step
The `DOCS_GIT_REPO` var needs to be exported when not set through a `.env`
file, as otherwise it is not propagated to commands run in the
`build-release.sh` script.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-04-04 18:50:02 +02:00
Alexander Trost 76ddb32025 Merge pull request #11968 from koor-tech/feature/ksd-124 2023-04-04 17:48:06 +02:00
travisn 558732d0ae ci: canary test should enable the prometheus module
The canary test is waiting for the prometheus module,
which is now disabled by default. For the canary test,
we need to enable the prometheus module for the external
cluster test.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-04-03 15:32:11 -06:00