Commit Graph
96 Commits
Author SHA1 Message Date
Alexander Trost 4ed35d6bd6 operator: allow setting ceph config options via ceph cluster crd
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-12-02 00:05:48 +01:00
sp98 345f92961b ci: filter both 14 and 64 disks in CI
github ci runner is allocating both 14 and 64 Gb disks to the setup.
Rook CI currently hard codes the lsblk with a 14G filter.
This PR updates the filter to use either 14G or 64G disks.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-11-06 13:30:10 +05:30
subhamkrai e14f8c7710 ci: remove workaround for clean disk action
The issues are fixed from the clean disk action
repo and we shouldn't require any workaround. Also,
I'm not using git revert to revert the commit is earlier
clean disk action was mentioned in every CI which was
duplicate and now we have placed the action to composite yaml

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-29 19:17:01 +05:30
subhamkrai 3267576307 ci: fix github action free disk space
Since `google-cloud-sdk` is renamed with `google-cloud-cli`,
the github action is failing to remove the older name and this
is causing ci issues. Applying changes suggested by community to
manually remove some packages to resolve the issue.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-28 18:56:59 +05:30
Blaine Gardner 3e2e906ea7 test: update multus canary test
Update the multus canary test to reflect modern knowledge about how it
should be configured.

No longer test for the network device in OSD pods. Pods will utterly
fail to start if Multus is unable to attach interfaces.
Instead, look to the OSD map to test the connections more wholistically.
OSDs must have map IPs that include both public and cluster network.
This implicitly tests that the interfaces exist in the Pod, and it
additionally verifies other details, like Ceph `*_network` configs are
set propertly.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-09-18 17:15:23 -06:00
subhamkrai 90d81e11b6 ci: upgrade the kubernetes versions for ci
since, kubernetes v1.28.0 release upgrading the
latest k8s version to v.1.28.0 excepth objectSuite
tes.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-28 10:35:44 +05:30
subhamkrai ad50c8e780 ci: use local tag instead master image
using `rook/ceph:master` tags take longer time
to pull and also, we should be using `rook/ceph:local-build`
tag for our ci. This will help canary raw test to be more stable.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-25 17:26:54 +05:30
Travis Nielsen 8e1090b3be Merge pull request #12624 from parth-gr/s5-cmd
object: fix s5cmd for s3 endpoint verification
2023-08-04 08:19:28 -06:00
parth-gr 8a3f058329 object: fix s5cmd for s3 endpoint verification
add a new toolbox yaml manifest which will use the
rook image instead of ceph image
for running s5 cmd container needs to run with rook image

closes: https://github.com/rook/rook/issues/12227

Signed-off-by: parth-gr <paarora@redhat.com>
2023-08-04 18:29:05 +05:30
subhamkrai b9c41fd548 ci: same ceph version in toolbox and cluster-test
Let's use same ceph version(latest Reef) in both
cluster-test and toolbox.yaml so that we don't need
to pull image twice. Alos, github action helper script
was calling `deploy_manifest_with_local_build` which
is required for operator.yaml and not for toolbox.yaml.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:11:30 +05:30
Jiffin Tony Thottan b48dc8a335 object: intial cosi driver controller design
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-07-18 22:49:41 +05:30
parth-gr ef4fa96e26 ci: catch error in ci for external rgw server canary test
currently we just print the error is anything fails in
script for rgw, So by that there is no panic or
failiure of script if something wrong happened,

Added Explicittly forcing to catch the error from
the error message that is thrown

Closes: https://github.com/rook/rook/issues/12244

Signed-off-by: parth-gr <paarora@redhat.com>
2023-05-31 18:42:52 +05:30
travisn 558732d0ae ci: canary test should enable the prometheus module
The canary test is waiting for the prometheus module,
which is now disabled by default. For the canary test,
we need to enable the prometheus module for the external
cluster test.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-04-03 15:32:11 -06:00
Rakshith R b9fc9a643e ci: add test for key rotation
This commit adds test for key rotation.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-03-08 12:07:26 +05:30
subhamkrai d983f3620e ci: fix multus ci
Signed-off-by: subhamkrai <srai@redhat.com>
2023-03-02 12:20:59 +05:30
Travis Nielsen 11e80b7047 Merge pull request #11436 from subhamkrai/bump-ci-k8s-version
ci: Update min K8s support to 1.21 and add 1.26 support
2023-02-14 20:26:56 -07:00
subhamkrai 1effb11a0a ci: bump min and max k8s support
this commit bump minmum k8s version to 1.21.14 and
max k8s version to latest 1.26.0. Keeping support for
most recent 6 versions.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-02-15 07:08:02 +05:30
Travis Nielsen eb4a727658 docs: remove mention of psps from documentation
The PSPs have long since been deprected. In K8s 1.21 the PSPs
were first deprecated, and support was completely removed
for them in 1.25. With Rook v1.11, the min supported version of
K8s is now 1.21. To reduce confusion in the documentation,
mention of the PSPs is now removed from the 1.11 docs.
For the corner case that users still require the PSPs,
the helm chart still contains the option for creating PSPs
or other users can still create the psp.yaml.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-13 13:30:32 -07:00
subhamkrai 45a10dcb85 core: upgrade to latest operator-sdk v1.25.0
upgrading to latest operator-sdk later version
and changing the way of csv generation.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-02-13 17:08:09 +05:30
Joshua Hoblitt 468f3f6f35 test: add secondary->master zone repl test to rgw-multisite-test
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2023-01-19 16:12:30 -07:00
Joshua Hoblitt 97bcc0a349 test: remove naughty words
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2023-01-19 11:13:12 -07:00
Shinya Hayashi 05875a3f4f osd: support loop devices for test clusters
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.

This feature is intended to be used for testing purposes only.

Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
2022-11-09 06:45:41 +00:00
Blaine Gardner 51fac0a993 nfs: fix nfs grace period when multus is enabled
Run the ganesha-rados-grace command in a remote pod when multus
networking is enabled.

Signed-off-by: parth-gr <paarora@redhat.com>
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-05 16:28:36 -06:00
Satoru Takeuchi 5c2d325cc4 ci: move to ubuntu 2004 completely
github-hosted runner ubuntu-18.04 will be removed by Dec. 1st, 2022.

https://github.blog/changelog/2022-08-09-github-actions-the-ubuntu-18-04-actions-runner-image-is-being-deprecated-and-will-be-removed-by-12-1-22/

So, let's change the all runnners to ubuntu-20.04.

We need to uninstall `snapd` because `snapd` prevents kernel from
unmounting `/mnt` completely.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-09-27 04:32:20 +00:00
parth-gr ddea97ebde nfs: fix nfs if multus is enabled
Closes: https://github.com/rook/rook/issues/10812
Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-21 14:09:40 +05:30
Satoru Takeuchi 47f7aefa20 ci: improve the log on intermittent build failure
The build process sometimes fails with intermittent problems. Some of them
are known problems and then we retry build process. However, there still
are unknown problems. In this case, it's hard to find the reason because
the build process exits immediately.

ref.
https://github.com/rook/rook/runs/8225144002?check_suite_focus=true#step:3:926

```
+ case "$o" in
+ exit 1
Error: Process completed with exit code 1.
```

To make debugging easier, let's print the output of `make`. This log won't be
too long since `make` prints most messages to stderr.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-09-07 10:48:21 +00:00
Rakshith R 07aac106df ci: add e2e for csi-nfsplugin restart
This commit adds e2e for csi-nfsplugin restart
when it is not on host-networking.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-09-02 14:16:15 +05:30
Travis Nielsen fb13df1d33 Merge pull request #10784 from subhamkrai/update-log-collector-script
test: update logcollector to get secrets -oyaml
2022-08-26 09:39:21 -06:00
subhamkrai d1e3023439 ci: collect operator logs in debug mode
collect operator logs in debug mode for both
canary tests and integration tests.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-08-26 19:32:59 +05:30
Blaine Gardner 1e9bbae583 docs: move PSPs from common.yaml to psp.yaml
Also update docs. Use `make gen-rbac` to generate psp.yaml.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-08-25 16:39:52 -06:00
subhamkrai 4a05115aef ci: fix multus depment cr link
Signed-off-by: subhamkrai <srai@redhat.com>
2022-08-22 16:58:55 +05:30
Travis Nielsen fc541764f1 Merge pull request #10484 from jsoref/spelling
core: fix spelling
2022-07-15 08:31:41 -06:00
Travis Nielsen 53955790f3 ci: pin the whereabouts version to v0.5.3
The whereabouts manifests in the master branch
have moved around, so until the new approach
is investigated we pin to the latest release
version v0.5.3

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-13 10:37:15 -06:00
Satoru Takeuchi 7e571f6114 osd: support OSD on logical volume in host-based cluster
Rook supports raw mode OSD in host-based cluster. So we can also
support OSD on logical volume in this kind of cluster.

Logical volumes aren't picked by filters (i.e. `useAllDevices: true`
and `device{Path,}Filter` to avoid unwanted LV consumption on upgrade.

Closes: https://github.com/rook/rook/issues/2047

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-07-11 21:08:50 +00:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Satoru Takeuchi 6b0f539f61 test: add canary test for encrypted osd in host-based clusters
It's better to test the creation of encrypted devices in host-based clusters.
This test will reduce the potential risks of regression when modifying
the osd-creation code.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-06-08 22:34:01 +00:00
Madhu Rajanna 276005cfd5 csi: add holder pod if csi hostnetworking is disabled
If csi is configured not the use the
hostnetworking, deploy the holder
pod for executing the commands with nsenter.
The implementation is same as multus.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-06-07 13:09:12 +05:30
Satoru Takeuchi 5b3633bfd6 test: add canary integration test for osd with metadata device
"metadataDevice" field in host based cluster is not tested yet.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-25 01:36:33 +00:00
Satoru Takeuchi a9465efa07 Merge pull request #9931 from cybozu-go/test-add-canary-integration-tests-for-osd-on-device
test: add canary integration tests for osd on device
2022-05-06 05:25:57 +09:00
Satoru Takeuchi 805163e2e2 test: add canary integration tests for osd on device
OSD on PVC is tested in many patterns but OSD on device is not.
It's preferable to add the following patterns for OSD on device.

- An OSD on a raw disk
- Two OSDs on a device

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-05 13:15:18 +00:00
Satoru Takeuchi ab38c10c2d test: avoid to create a corrupted GPT headers
The following steps in CI scripts create corrupted GPT headers.

```console
$ sudo sgdisk --zap-all --clear --mbrtogpt -g -- "$DISK"
$ sudo dd if=/dev/zero of="$DISK" bs=1M count=10
```

It results in the failure of the succeeding `sgdisk --print "$DISK".

Here is an example.

```console
$ sudo parted /dev/sdb mklabel msdos
...
$ sudo sgdisk --zap-all --clear --mbrtogpt -g -- /dev/sdb
...
$ sudo dd if=/dev/zero of=/dev/sdb bs=1M count=10
...
$ sudo sgdisk --print /dev/sdb
Caution: invalid main GPT header, but valid backup; regenerating main header
from backup!

Warning: Invalid CRC on main header data; loaded backup partition table.
Warning! One or more CRCs don't match. You should repair the disk!
Main header: ERROR
Backup header: OK
Main partition table: OK
Backup partition table: OK

Invalid partition data!
$ echo $?
2
```

We can safely use the simple `sgdisk --zap-all "$DEVICE"` here.
It's not necessary to convert "mbr" to "gpt" because we'll make new GPT
labels just after this command.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-05 13:11:21 +00:00
Sébastien Han 6ef491d9fc osd: close the encrypted disk after cleanup is done
Once we are done cleaning up the content of the encrypted osd, let's
close the main LUKS device.
This avoids having dm devices on the system after the cleanup.

Closes: https://github.com/rook/rook/issues/10181
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-05-03 16:47:59 +02:00
Sébastien Han 73b1347675 core: fix csi-cephfsplugin pod restart on non-hostnetworking env
Implementation of the design proposed in
https://github.com/rook/rook/pull/9903.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 17:53:01 +02:00
Sébastien Han 2c09bbb91b ci: add multus integration test
This new integration test will deploy a cluster with multus enabled. It
will be comprised of two network interfaces for ceph public and cluster
communications.
For now, it only deploys a Ceph cluster up to the OSDs.

Closes: https://github.com/rook/rook/issues/9784
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-12 17:30:47 +02:00
Travis Nielsen 4f9a72c1db Merge pull request #9894 from yuvalman/resources
helm: enable resource defaults on rook components
2022-03-21 13:58:17 -06:00
Yuval Manor ff7a5c2d2d helm: enable resource defaults on rook components
This PR assign default values to the resources of all Rook and Ceph components

Closes: https://github.com/rook/rook/issues/9858
Signed-off-by: Yuval Manor <yuvalman958@gmail.com>
2022-03-21 21:23:35 +02:00
Satoru Takeuchi 3bbffd1e4d test: avoid potential data inconsistency on zapping disk
Data inconsistency might happen if the disk is accessed just after disk
zapping because direct I/O is not synchronous by itself.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-03-18 11:50:04 +00:00
Travis Nielsen 75a385e187 csi: update volume replication to v0.3.0
Update to the latest volume replication crd version of
v0.3.0

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-03-09 15:53:45 -07:00
Blaine Gardner bb0d0d39f6 Merge pull request #9384 from leseb/fix-7036
subvolumegroup: add new crd
2021-12-21 10:24:23 -07:00
Sébastien Han 6e9eb33782 subvolumegroup: add new crd
This introduces a new CRD to add the ability to create subvolumegroup
for a given ceph filesystem volume. Typically the name of the volume is
the name of the filesystem created by rook.

Closes: https://github.com/rook/rook/issues/7036
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 16:28:38 +01:00