The generation of a long node name in the integration tests was
being done based on the k8s version. In the past, older K8s versions
did not support the changing name. Now it's more maintainable if
we generate the long name depending on the test suite.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Do not use the cross build container when building, publishing, and
promoting rook/ceph images. It is no longer needed, and its complexity
can add flakiness.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The upgrade integration test was from rook v1.6 to the latest master.
This was necessary until we are ready for the v1.8 release, from which
time we want to focus the upgrade testing from v1.7 to the latest
master.
The duplication in the test CRs and other resources is now reduced
by the upgrade calling a thin wrapper to forward a call to the
master version of the resource. When a new feature is added that
needs to be differentiated from the previous version, the method
then can be implemented instead of wrapping the master implementation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.
So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.
Signed-off-by: Sébastien Han <seb@redhat.com>
The instructions to deploy a developer environment are useful and should
be part of the official documentation instead of being hidden in the
repo's README.md. Also, remove ancient documentation to deploy with
kubespray a multi-node env.
Also, add a warning to always run tests env in a virtual machine.
Closes: https://github.com/rook/rook/issues/9166
Signed-off-by: Sébastien Han <seb@redhat.com>
By default HPA can use details about memory or CPU consumption for
autoscaling, but also it can use custom metrics as well. There are alot
provides supports HPA via customer and one of them is KEDA project. Here
it is done with help of Prometheus Scaler
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
The ingress api version changed when it went to v1, and this has caused some upheaval
throughout the kubernetes ecosystem. This commit uses a common method of deciding which
ingress api to use, and allows the optional override of the kubernetes version
presented to helm using the helm build-in capabilities.
also add an ingress into the helm integration tests so any regressions to how ingresses
are handled in the future are caught easier.
Closes rook#9174
Signed-off-by: Tom Hellier <me@tomhellier.com>
In the nightly job, run the test with the latest Ceph version so we can
detect if there are RGW changes in Ceph that might break multisite. Use
a reusable GitHub action workflow to duplicate as little code as
possible.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
When requesting issuer, we run a local kubectl proxy command which spawn
a proxy server. However, we must wait for the proxy to be ready before
we actually start making requests to it.
Now the CI waits up to 10sec to retrieve the issuer.
Closes: https://github.com/rook/rook/issues/9090
Signed-off-by: Sébastien Han <seb@redhat.com>
We can pass bound_service_account_names with a comma separated list of
service accounts. Let's do this instead of remapping new values.
Earlier, we thought a single service account could be added per Vault
role and we were using other variables like
`VAULT_AUTH_KUBERNETES_ROOK_OPERATOR_ROLE` that we were remapping to
`VAULT_AUTH_KUBERNETES_ROLE` internal for the API calls to Vault.
Signed-off-by: Sébastien Han <seb@redhat.com>
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Ths NFS spec now supports the CephBlockPool spec which means that it can
take advantage of all the known settings like compression, size, failure
domain etc.
Closes: https://github.com/rook/rook/issues/9034
Signed-off-by: Sébastien Han <seb@redhat.com>
The yaml validation of the examples folder requires all the CRDs
to be created in advance of the dry-run command.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook cluster-wide encryption can now use the native Kubernetes
authentication to interact with vault KMS instead of using the token
method.
Signed-off-by: Sébastien Han <seb@redhat.com>
using `docker-env` command to copy image giving error
`X Exiting due to ENV_DRIVER_CONFLICT: 'none' driver does not support 'minikube docker-env' command`
so, using `minikube image load <image>` commmand to copy
image.
Signed-off-by: subhamkrai <srai@redhat.com>
Add some commands to get more partition info when setting up the GH
action runner's disk for use in integration tests. This will both aid in
debugging and may "jog" the system such that it will no longer need to
reload the partition info when running the OSD prepare job.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The `head` command exits once it has output which can result in a
SIGPIPE error if the command piping its output to head hasn't yet
finished. Use `awk 'FNR <= 1'` instead, which waits on the input pipe to
close before it exits.
See here for more info:
https://unix.stackexchange.com/a/256047
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
In the CI tests that use `validate_cluster.sh display_status` to gather
logs, the prepare pod log collection failed. Fix this.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The CephObjectRealm controller would fail all subsequent reconciles if
the first reconcile created the Kubernetes Secret containing the access
keys for the realm but where the radosgw-admin command failed to create
the realm. This was the only idempotency issue found after reviewing the
CephObjectRealm controller.
Resolves#8954
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
When TLS is used and includes a caert, client key/cert, we need to copy
the content of the secret to a file in the operator's container
filesystem so that we can build the TLS config and thus the HTTP Client,
which reads those files.
Also, removing the files after each API call so they don't persist on
the filesystem forever.
Signed-off-by: Sébastien Han <seb@redhat.com>
Using a default value for CompressionMode to none effectively overrides
any values for Parameters. It is deprecated but still takes precedence.
Which means that in its previous form, Parameters was always ignored
since CompressionMode was always set to none when empty.
Signed-off-by: Sébastien Han <seb@redhat.com>
The daily arm test suite is failing due to the new local-build
image tag. Instead of waiting for the arm build to complete,
we can just pick up the latest tag from the same branch
that was already pushed to dockerhub and no need to build
again.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.
Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.
To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
of its operation that didn't cause failures in the past
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>