we have changed the --cluster-name flag to --k8s-cluster-name
but to support automation while upgrade we should support
the legacy flag
Signed-off-by: parth-gr <partharora1010@gmail.com>
Currently, `wait_for_prepare_pod` only waits until 1 prepare pod and 1
OSD pod are in Running state.
So if `wait_for_ceph_to_be_ready` failed after `wait_for_prepare_pod`
succeeded, there are 4 possible cases:
1. some prepare pods didn't start correctly;
2. some prepare pods didn't finish correctly;
3. some OSD pods didn't start correctly; or
4. all prepare pods and OSD pods worked correctly, but some other
process failed.
As far as I understand, the number of prepare pods on the GitHub CI is
always 1, so cases 1 and 2 are not problematic. However, it is hard to
distinguish the other two cases from the CI log.
To solve the above problems, this patch makes wait_for_prepare_pod wait
for all OSD pods to become running state. If wait_for_prepare_pod
timeouts before the OSD pods become running state, this will be a strong
indication that the OSD pods didn't start properly (i.e., case 3).
Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
currently, in the cluster-test yaml, file we have both cephCluster
CR and cephBlockPool CR and we're trying to update the cephCluster
CR to disable the livnessProbe of mon, osd, mgr but since we had `yq write -d1`
it was updating the second instance of `spec` which is cephBlockPool
instead we have to do `yq write -d0` to update the cephCluster `spec`.
Signed-off-by: subhamkrai <srai@redhat.com>
I was debugging helm tests, I noticed, the operator namespace
content is empty and operator is created in same namespace as
other pods. Also, let's collect the 'kube-system' namespace logs
in multus test only as that is the only test where we need to debug
cluster networking and require kube-system logs.
Signed-off-by: subhamkrai <srai@redhat.com>
use `sed` to update all the namespace in filesystem-test.yaml file.
Since with subvolumegroup CR added there are more than one `namespace: *`
present and only one was updating by yq command.
Signed-off-by: subhamkrai <srai@redhat.com>
In canary test `multi-cluster-mirroring` we also need to collect
logs of `rook-ceph-secondary` namespace to debug the ci.
Signed-off-by: subhamkrai <srai@redhat.com>
we are passing inputs to composite action, like
`github_token` which is not being used anymore. Its'
better to clean those lines.
Signed-off-by: subhamkrai <srai@redhat.com>
Loop variables cannot be reliably uses since they will
change with each iteration. Update these loop variable
uses to be safe by indexing the slice rather than
using the loop variable directly.
Also suppress the linter issues for passwords used
in tests.
Signed-off-by: travisn <tnielsen@redhat.com>
The K8s v1.25.17 release does not exist, which was mistakenly
set in the CI for the master and release tests. Now we revert
it back to v1.25.16.
Signed-off-by: travisn <tnielsen@redhat.com>
change the json output of storageclass to also inclusde rados namespace,
and also added the changes in the csi secret
Signed-off-by: parth-gr <paarora@redhat.com>
Upgrade minimum kubernetes support to 1.23.17 from 1.22.17.
Also, upgrade kubernetes version latest minor release.
Signed-off-by: subhamkrai <srai@redhat.com>
In, go 1.20 govulncheck ci is failing due to
some vulnerabilities which affects windows
machine and rook is not used in windows machine.
The solution is to update go version to 1.21.4
Signed-off-by: subhamkrai <srai@redhat.com>
Since nobody seems to see the instructions for opening a PR,
make another attempt to clarify what someone should do
when they are opening a PR.
Signed-off-by: travisn <tnielsen@redhat.com>
github ci runner is allocating both 14 and 64 Gb disks to the setup.
Rook CI currently hard codes the lsblk with a 14G filter.
This PR updates the filter to use either 14G or 64G disks.
Signed-off-by: sp98 <sapillai@redhat.com>
Automatic PR from app.stepsecurity.io (initiated by BlaineEXE) to limit
permissions in github actions to only those needed.
Signed-off-by: StepSecurity Bot <bot@stepsecurity.io>
govulncheck is a static code analyzer that scans more deeply than Snyk.
Additionally, because govulncheck checks the code itself, it can rule
out vulnerabilities that are present in dependencies but that are not
used in Rook's code path.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Add the ability to specify node profiles in the multus validation test.
This addresses a few points of early feedback on the validation tool.
Statements below critique the tool's behavior before this patch.
1. The tool assumes all daemons are on public and cluster network, which
means users who have a significantly smaller cluster net (a
design choice) cannot run a single test to determine if Rook is
likely to install correctly.
2. The tool does not have placement options to select only a subset of
Kubernetes nodes to run validation on.
3. Users of multus seem to have a dedicated pool of storage nodes more
often than the average Rook install. This makes sense for security-
and perforance-minded users. The tool cannot run a single test to
verify storage-only and general-workload nodes at one time.
These points are addressed by allowing users to specify configurations
for different "NodeTypes."
Each NodeType config has options for selecting the number of OSDs as
well as the number of other (non-OSD) Ceph daemons. This limits the
unnecessary exhaustion of cluster network addresses from critique 1.
Each NodeType config has its own placement (critique 2).
Users can define as many NodeTypes as needed to test the network for
their planned CephCluster. Specifically, this allows the tool to test
storage-only nodes and generalized-workload nodes at the same time. An
arbitrary number of NodeTypes are allowed to support even more highly
specialized cluster setups, such as multiple tiers of storage nodes
where some storage-only nodes may run more OSDs than others.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The GH action to set up Golang (actions/setup-go) has cache
functionality. For cache hits, this saves ~2 minutes for each run. In
order for the cache to hit, caching must not be disabled, and the code
must be checked out before setting up Golang.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The issues are fixed from the clean disk action
repo and we shouldn't require any workaround. Also,
I'm not using git revert to revert the commit is earlier
clean disk action was mentioned in every CI which was
duplicate and now we have placed the action to composite yaml
Signed-off-by: subhamkrai <srai@redhat.com>