Commit Graph
244 Commits
Author SHA1 Message Date
Blaine Gardner 243a47bfc4 build: do not use cross build container for ceph
Do not use the cross build container when building, publishing, and
promoting rook/ceph images. It is no longer needed, and its complexity
can add flakiness.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-03 11:06:21 -07:00
Sébastien Han c890710b63 core: change directory layout
As per discussion, proposing a new layout for the charts/yaml/olm files.

./deploy
├── charts
│   ├── rook-ceph
│   │   └── templates
│   └── rook-ceph-cluster
│       └── templates
├── examples
│   ├── csi
│   │   ├── cephfs
│   │   └── rbd
│   ├── flex
│   ├── monitoring
│   ├── pre-k8s-1.16
└── olm
    └── assemble

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-30 09:12:53 +01:00
Sébastien Han 7a223adde4 Merge pull request #9230 from leseb/osd-rm-check
osd: check if osd is safe-to-destroy before removal
2021-11-25 10:55:34 +01:00
Sébastien Han ad2c3c2ae6 Merge pull request #8931 from BlaineEXE/test-rgw-multisite-in-nightly-tests
test: test rgw multisite nightly
2021-11-25 10:25:57 +01:00
Sébastien Han 7402c2cce6 osd: check if osd is ok-to-stop before removal
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.

So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 10:21:32 +01:00
Travis Nielsen d340bccd44 Merge pull request #9202 from thotz/hpakedadoc
docs: add details about HPA via KEDA
2021-11-23 10:42:03 -07:00
Jiffin Tony Thottan d32c837e35 docs: add details about HPA via KEDA
By default HPA can use details about memory or CPU consumption for
autoscaling, but also it can use custom metrics as well. There are alot
provides supports HPA via customer and one of them is KEDA project. Here
it is done with help of Prometheus Scaler

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-23 16:50:18 +05:30
Blaine Gardner e80368674d test: run RGW multisite test in nightly job
In the nightly job, run the test with the latest Ceph version so we can
detect if there are RGW changes in Ceph that might break multisite. Use
a reusable GitHub action workflow to duplicate as little code as
possible.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-19 09:56:41 -07:00
Jiffin Tony Thottan aba50d3ca9 object: add support in RGW to communicate vault with TLS
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-17 10:19:28 +05:30
Sébastien Han 0e26176b54 ci: wait for kubeproxy to be ready
When requesting issuer, we run a local kubectl proxy command which spawn
a proxy server. However, we must wait for the proxy to be ready before
we actually start making requests to it.
Now the CI waits up to 10sec to retrieve the issuer.

Closes: https://github.com/rook/rook/issues/9090
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-08 17:45:53 +01:00
Sébastien Han 16729e0e38 osd: use multiple service account for vault role
We can pass bound_service_account_names with a comma separated list of
service accounts. Let's do this instead of remapping new values.
Earlier, we thought a single service account could be added per Vault
role and we were using other variables like
`VAULT_AUTH_KUBERNETES_ROOK_OPERATOR_ROLE` that we were remapping to
`VAULT_AUTH_KUBERNETES_ROLE` internal for the API calls to Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-04 15:15:46 +01:00
Travis Nielsen b3b8172e13 Merge pull request #9056 from leseb/fix-8851
docs: simplify dev guide
2021-11-03 12:34:41 -06:00
Sébastien Han 85f8fdec5a docs: simplify dev guide
We removed the obsolete `minikube.sh` script for a cleaner dev
procedure.

Closes: https://github.com/rook/rook/issues/8851
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-03 18:56:33 +01:00
Yuzuki Mimura 536b59ef0f rgw: change the way to livenessProbe and introduce readinessProbe
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.

Closes: #8407

Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-10-29 15:07:29 +00:00
Sébastien Han e0145b9643 nfs: add pool setting CR option
Ths NFS spec now supports the CephBlockPool spec which means that it can
take advantage of all the known settings like compression, size, failure
domain etc.

Closes: https://github.com/rook/rook/issues/9034
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-29 09:04:31 +02:00
Travis Nielsen 2c61ea2fc3 test: create volume replication crds for yaml validation
The yaml validation of the examples folder requires all the CRDs
to be created in advance of the dry-run command.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-26 23:28:40 -06:00
Sébastien Han 871932cc8e ci: log and describe all the pod/deploys on errors
When the CI fails, we want to collect more logs and describe from the
entire environment.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-26 08:29:51 +00:00
Sébastien Han 18a4047679 osd: add support for k8s with vault kms
Rook cluster-wide encryption can now use the native Kubernetes
authentication to interact with vault KMS instead of using the token
method.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-21 13:59:28 +02:00
Travis Nielsen 8f639c61de Merge pull request #8938 from subhamkrai/update-minikube
ci: update minikube and kubernetes version
2021-10-20 07:52:06 -06:00
subhamkrai 6be9071a22 ci: update minikube.sh script to use minikube command
using `docker-env` command to copy image giving error
`X Exiting due to ENV_DRIVER_CONFLICT: 'none' driver does not support 'minikube docker-env' command`
so, using `minikube image load <image>` commmand to copy
image.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:26:50 +05:30
subhamkrai e49cdf0a8d ci: clean validate_cluster.sh script
removing `trap display_status SIGINT ERR` command
from the file as `display_status` func has been removed.

Closes: https://github.com/rook/rook/issues/9004
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 10:43:50 +05:30
Blaine Gardner 65619c2677 test: get more partition info setting up ci disk
Add some commands to get more partition info when setting up the GH
action runner's disk for use in integration tests. This will both aid in
debugging and may "jog" the system such that it will no longer need to
reload the partition info when running the OSD prepare job.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-18 11:29:13 -06:00
Blaine Gardner ba7745d4af test: do not use head because of pipe errors
The `head` command exits once it has output which can result in a
SIGPIPE error if the command piping its output to head hasn't yet
finished. Use `awk 'FNR <= 1'` instead, which waits on the input pipe to
close before it exits.

See here for more info:
https://unix.stackexchange.com/a/256047

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-15 10:14:41 -06:00
Sébastien Han 73970a56c5 ci: fix prepare pod wrong selector
We missed the `app=` in the selector.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-15 10:36:15 +02:00
Blaine Gardner f287df70ee test: fix prepare pod log collection in CI
In the CI tests that use `validate_cluster.sh display_status` to gather
logs, the prepare pod log collection failed. Fix this.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-14 11:15:59 -06:00
Blaine Gardner 34a8b42e97 test: try to un-flake multi-cluster-mirror test
Try to un-flake the multi-cluster-mirror test that keeps failing on this
PR.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-13 11:06:51 -06:00
Blaine Gardner 65c4972038 rgw: make CephObjectRealm controller idempotent
The CephObjectRealm controller would fail all subsequent reconciles if
the first reconcile created the Kubernetes Secret containing the access
keys for the realm but where the radosgw-admin command failed to create
the realm. This was the only idempotency issue found after reviewing the
CephObjectRealm controller.

Resolves #8954

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-13 11:06:51 -06:00
Travis Nielsen 74dac725a4 test: run the daily arm tests against official builds
The daily arm test suite is failing due to the new local-build
image tag. Instead of waiting for the arm build to complete,
we can just pick up the latest tag from the same branch
that was already pushed to dockerhub and no need to build
again.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-12 15:26:18 -06:00
Blaine Gardner a1dd256d4b Merge pull request #8911 from BlaineEXE/rgw-commands-use-staging-flag
rgw: replace period update --commit with function
2021-10-12 08:22:41 -06:00
Blaine Gardner 956430826c rgw: add integration test for committing period
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.

Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.

To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
   of its operation that didn't cause failures in the past

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 15:24:59 -06:00
Sébastien Han abc5c635a6 ci: fix osd disk permission on provisioning
When the OSD is prepared, systemd-udev kicks in since the device has
been exclusively opened and thus reverts the permissions to root:disk.
We need a udev rule to force re-applying the correct ceph permission so
we can consume the disk.

Closes: https://github.com/rook/rook/issues/8942
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-11 16:50:28 +02:00
Sébastien Han 4f9c31fa8b ci: clarify the wait for csi to be ready
The wait is now more comprehensive.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-11 16:47:02 +02:00
Travis Nielsen af29d1d801 build: run canary tests against the local-build tag
The canary tests were still picking up the tag from operator.yaml
and toolbox.yaml instead of the new test local-build tag.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 6f48dce3f5)
2021-09-30 15:33:48 -06:00
Blaine Gardner e2ba16f710 build: start tracking rbac generated from helm chart
This is a starting step to be able to generate common.yaml from Helm
charts. For right now, we merely want to be able to determine when the
rendered output of the Helm chart changes.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-30 12:20:24 -06:00
Sébastien Han e3f7d2f046 ci: retry on image pull error
Try to avoid the following:

```
error pulling image configuration: received unexpected HTTP status: 500 Internal Server Error
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-29 16:03:38 +02:00
Satoru Takeuchi fc59b76afc Merge pull request #8837 from leseb/fix-8825
ci: wait longer for csi to be available
2021-09-27 19:56:13 +09:00
Sébastien Han 85216a266c ci: wait longer for csi to be available
Sometimes the CI needs more time...

Closes: https://github.com/rook/rook/issues/8825
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 12:12:46 +02:00
Travis Nielsen 76748f0419 test: remove obsolete minishift and kubeadm scripts
The test scripts to start minishift and kubeadm are no
longer in use by Jenkins or anyone else.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 22:16:53 +00:00
Travis Nielsen 4583a80001 Merge pull request #8790 from travisn/test-local-build
test: Run all integration tests against the local build
2021-09-22 14:31:18 -06:00
Travis Nielsen 9ff0753514 test: run all integration tests against the local build
The integration tests must always be run against the local
build of rook, and an image should never be pulled from dockerhub.
To prevent pulling a release or master tag, the local build
will use a tag specific to the build and not ever published
elsewhere.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit a8a40428b0)
2021-09-22 07:39:42 -06:00
Sébastien Han 7c8dc4bc1a ci: fix multisite test
We just need to wait a little for the object to be replicated to the
other gateway. A simple retry solves this.

Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-22 10:23:02 +02:00
Sébastien Han 8786b40d64 rgw: do not create the rgw ops user on the secondary cluster
If the cluster where the rgw is started is secondary and not primary,
trying to create the admin ops user will fail with:

```
Please run the command on master zone.
Performing this operation on non-master zone
leads to inconsistent metadata between zones
```

So we need to force the creation regardless, it is fine the creation will
return UserAlreadyExist and then we just read the current user.

Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 16:58:03 +02:00
subhamkrai 038031ddbd ci: dry-run is deprecated and replaced with --dry-run=client
--dry-run is deprecated and replaced with --dry-run=client

Signed-off-by: subhamkrai <srai@redhat.com>
2021-09-14 16:13:27 +05:30
Sébastien Han 0e961c4bdc ci: stop build 3 times
This time for real.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-09 17:51:03 +02:00
Sébastien Han 73c340ded7 ci: fix pod list
We should not use .items[0].metadata.name if the array length is 0. This
is the case when nothing has been initialized yet. Instead, we should
use .items[*].metadata.name, the wildcard ensures to always return 0
even if nothing is present yet.

Fixes: https://github.com/rook/rook/issues/8676
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-09 17:24:56 +02:00
Blaine Gardner 7febd4e303 ci: add debugging to release build script
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-01 13:58:21 -06:00
parth-gr 827ee1d879 ceph: auto grow OSDs size on PVCs
When an OSD reaches OSD_NEARFULL state,
we have to manually increase the PVC volume claim
or manually increase the count of OSDs in the device set

Added a script auto-grow-storage.sh which will
i)automatically increase claim volume
ii)automatically add number of OSDs

Closes: https://github.com/rook/rook/issues/6101
Signed-off-by: parth-gr <paarora@redhat.com>
2021-09-01 12:57:42 +05:30
subhamkrai ebe21c8f68 ci: add action for shellcheck linter
we are adding new linter for shellcheck.
As we are writing more shell scripts this
will help maintain quality.

Also, doing all the changes required in
bash files to pass this shellcheck.

Closes: https://github.com/rook/rook/issues/8431
Signed-off-by: subhamkrai <srai@redhat.com>
2021-08-26 10:03:32 +05:30
parth-gr 33bf21ba69 ceph: update CSR from v1beta1 to v1
This commit update the certificates.k8s.io to use version v1
Updated to v1 as v1beta1 certificates.k8s.io is deprecated in v1.19+

Closes: https://github.com/rook/rook/issues/8308
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-12 20:27:05 +05:30
Ali Maredia b6404ae2ab ceph: rgw multisite intgration testing
This test:
- starts up 2 minikube clusters
- create 2 ceph clusters
- creates object multisite CRDs on each cluster and
syncs the clusters
- writes an object to cluster 1 and reads it on
cluster 2

This commit also adds new functions in
github-action-helper.sh that aid in the multisite
test and the multi cluster mirroring test.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2021-08-11 13:25:43 -04:00