Do not use the cross build container when building, publishing, and
promoting rook/ceph images. It is no longer needed, and its complexity
can add flakiness.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Use yq instead of Python for parsing RBAC from the Helm chart. We need
to use yq v4.14.1 or higher to fix yq's handling of the yaml header
markers ('---'). Update the Makefile's yq version to v4, which also
requires updating the script to update the CRDs. This was quite easy.
It is very difficult, however, to change the version of yq used by the
CSV generating/parsing scripts, which already used their own yq
download. Continue using yq v3 for this.
In order to make sure the scripts are using the right version of yq, add
basic validation to them to verify they are running v3 or v4 as required
for their operation.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.
So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.
Signed-off-by: Sébastien Han <seb@redhat.com>
The ingress api version changed when it went to v1, and this has caused some upheaval
throughout the kubernetes ecosystem. This commit uses a common method of deciding which
ingress api to use, and allows the optional override of the kubernetes version
presented to helm using the helm build-in capabilities.
also add an ingress into the helm integration tests so any regressions to how ingresses
are handled in the future are caught easier.
Closes rook#9174
Signed-off-by: Tom Hellier <me@tomhellier.com>
In the nightly job, run the test with the latest Ceph version so we can
detect if there are RGW changes in Ceph that might break multisite. Use
a reusable GitHub action workflow to duplicate as little code as
possible.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
When requesting issuer, we run a local kubectl proxy command which spawn
a proxy server. However, we must wait for the proxy to be ready before
we actually start making requests to it.
Now the CI waits up to 10sec to retrieve the issuer.
Closes: https://github.com/rook/rook/issues/9090
Signed-off-by: Sébastien Han <seb@redhat.com>
The rook operator as well as the toolbox pod run with the "rook" user
with UID 2016. The UID was chosen based on the year of the initial
commit in the rook/rook repository.
No more root user running.
Closes: https://github.com/rook/rook/issues/8734
Signed-off-by: Sébastien Han <seb@redhat.com>
This reverts commit c53304c3b4.
After experimenting with this release, we have found that the
`workflow_run` action source doesn't work for branches. The
documentation was updated with this PR and has more detail.
https://github.com/github/docs/pull/531
For now, we will revert back to using on.push.tags = ["v*"]
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The deployment of the OSD is already confirmed at the end of
`deploy first cluster rook' step.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
We cannot use the secret when PRs are pushed from forks so let's run the
security scan only after PRs are merged.
Signed-off-by: Sébastien Han <seb@redhat.com>
Rook cluster-wide encryption can now use the native Kubernetes
authentication to interact with vault KMS instead of using the token
method.
Signed-off-by: Sébastien Han <seb@redhat.com>
When we copy the peer token secret from a namespace to another we also
copy the ownerref, however the creation succeeds but the controller
removes the secret since the uid of the owner do not exist in that
namespace.
Signed-off-by: Sébastien Han <seb@redhat.com>
Add some commands to get more partition info when setting up the GH
action runner's disk for use in integration tests. This will both aid in
debugging and may "jog" the system such that it will no longer need to
reload the partition info when running the OSD prepare job.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The `head` command exits once it has output which can result in a
SIGPIPE error if the command piping its output to head hasn't yet
finished. Use `awk 'FNR <= 1'` instead, which waits on the input pipe to
close before it exits.
See here for more info:
https://unix.stackexchange.com/a/256047
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The Filesystem is deployed later in the run so at this point we only
need to validate the OSDs are running. This was failing since we were
validating the cephfs pool which is not present yet.
Signed-off-by: Sébastien Han <seb@redhat.com>
The daily arm test suite is failing due to the new local-build
image tag. Instead of waiting for the arm build to complete,
we can just pick up the latest tag from the same branch
that was already pushed to dockerhub and no need to build
again.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
earlier push build action was not triggered due
to github action limitation.
```
An action in a workflow run can’t trigger a new workflow run.
```
so now, we'll use user personal toke to create tag so that push
build action pick the user not github action who created the tag.
Closes: https://github.com/rook/rook/issues/8580
Signed-off-by: subhamkrai <srai@redhat.com>
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.
Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.
To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
of its operation that didn't cause failures in the past
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The canary tests were still picking up the tag from operator.yaml
and toolbox.yaml instead of the new test local-build tag.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit 6f48dce3f5)
This is a starting step to be able to generate common.yaml from Helm
charts. For right now, we merely want to be able to determine when the
rendered output of the Helm chart changes.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Let's run the test every day at midnight since the code is moving
fast and tends to fail. This gives us some time to fix the issues.
It can be triggered on PR too if the "run-mgr-suite" label is set.
Signed-off-by: Sébastien Han <seb@redhat.com>
Reading from a file, processing the file in a pipe, and outputting the
piped content to the same file can have undefined results, often leading
to an empty file. Use an intermediate temp file for generating the
offline image list.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Since we were pointing at CephUpgradeSuite test only this was running
all the tests and thus some resources were created by other tests.
Let's just run the test we need.
Closes: https://github.com/rook/rook/issues/8843
Signed-off-by: Sébastien Han <seb@redhat.com>
run tmate only for `pull_request` event and
remove for other events. As during release
when tests fail it takes extra 30 min to fail the
test or someone manually has to end the test
which is painfully especially during release time.
Signed-off-by: subhamkrai <srai@redhat.com>
We have new jobs now:
* one that runs both smoke and object on the next Pacific version
* one that runs both smoke and object on Ceph master
* one that tests the upgrade from the current pacific stable to the
pacific devel
* one that tests the upgrade from the current octopus stable to the
octopus devel
Signed-off-by: Sébastien Han <seb@redhat.com>
In Rook v1.8 the min version of K8s supported is updated to 1.16.
Users running on older versions of K8s are recommended to update
to 1.16 or newer before updating to Rook v1.8.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>