Use Quay as the source for the nginx-unprivileged image used for the
Multus validation tool because Quay does not rate limit image pulls,
which are a common complaint for users of the tool.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Since v18.2.1 is the default recommendation for the ceph version,
update all the docs and examples to that version.
Signed-off-by: travisn <tnielsen@redhat.com>
Currently, `wait_for_prepare_pod` only waits until 1 prepare pod and 1
OSD pod are in Running state.
So if `wait_for_ceph_to_be_ready` failed after `wait_for_prepare_pod`
succeeded, there are 4 possible cases:
1. some prepare pods didn't start correctly;
2. some prepare pods didn't finish correctly;
3. some OSD pods didn't start correctly; or
4. all prepare pods and OSD pods worked correctly, but some other
process failed.
As far as I understand, the number of prepare pods on the GitHub CI is
always 1, so cases 1 and 2 are not problematic. However, it is hard to
distinguish the other two cases from the CI log.
To solve the above problems, this patch makes wait_for_prepare_pod wait
for all OSD pods to become running state. If wait_for_prepare_pod
timeouts before the OSD pods become running state, this will be a strong
indication that the OSD pods didn't start properly (i.e., case 3).
Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
I was debugging helm tests, I noticed, the operator namespace
content is empty and operator is created in same namespace as
other pods. Also, let's collect the 'kube-system' namespace logs
in multus test only as that is the only test where we need to debug
cluster networking and require kube-system logs.
Signed-off-by: subhamkrai <srai@redhat.com>
Both the helm tests are failing because,
```
2023-12-19 08:47:06.136640 E | ceph-file-controller: failed to reconcile CephFilesystem "helm-ns/ceph-filesystem-test". CephFilesystem "helm-ns/ceph-filesystem-test" will not be deleted until all dependents are removed: CephFilesystemSubVolumeGroups: [ceph-filesystem-test-csi]
```
So, let's remove the ceph SVGs before removing Ceph filesystem.
Signed-off-by: subhamkrai <srai@redhat.com>
In canary test `multi-cluster-mirroring` we also need to collect
logs of `rook-ceph-secondary` namespace to debug the ci.
Signed-off-by: subhamkrai <srai@redhat.com>
The filesystem is not being deleted because of the existing svg.
Add a cliet call to delete the deafult csi svg,
for helm test
Signed-off-by: parth-gr <partharora1010@gmail.com>
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.
Signed-off-by: subhamkrai <srai@redhat.com>
at this moment we have two different cluster spec files for testing
cluster-test.yaml and cluster-on-pvc-minikube.yaml. With the new
option user can choose which one to use to bootstrap the cluster
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
current code gets the ip:port for the dashboard by using
the ip of the mgr pods. This works great when there's only
one mgr but it fails in case of a multi-node cluster
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
adding support to specify operator and cluster namespaces through new
command-line arguments: -c for cluster and -o for operator. The script
will automatically update the namespace in the corresponding YAML files
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
The upgrade test should always upgrade from the previous
minor release to the latest master. With 1.13 releasing
soon, now we uprade from 1.12 to master, to confirm
if there are any upgrade issues to 1.13.
Signed-off-by: travisn <tnielsen@redhat.com>
This change to the create-dev-cluster script is intended to make the
handling of invalid invocations more uniform
examples are unknown options and options requiring an argument specified
without one.
The unification is achieved by encapsulating the corresponding code in a
function.
Signed-off-by: Michael Adam <obnox@samba.org>
This change introduce a new argument -p to specify the minikube
profile for the new cluster. This way we can have multiple rook
clusters running on the same machine.
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition
Signed-off-by: parth-gr <paarora@redhat.com>
Let's use the minikube profiles feature to set up a unique profile
exclusively for the Rook cluster. By doing this, we make sure that we
don't run into any issues with other minikube profiles that users
might already have in their local setup.
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
The new create-dev-cluster script errored out for me, not finding the
deploy/examples directory.
This change fixes that problem by making the path relative to the
script's directory, so that it works in a rook code tree from where it
is expected to be run.
Signed-off-by: Michael Adam <obnox@samba.org>
github ci runner is allocating both 14 and 64 Gb disks to the setup.
Rook CI currently hard codes the lsblk with a 14G filter.
This PR updates the filter to use either 14G or 64G disks.
Signed-off-by: sp98 <sapillai@redhat.com>
Adding a new '-m' option for activating monitoring during
cluster setup. With this option turned on, the script will handle the
installation of the monitoring stack and configure the dashboard to
connect to the newly installed Prometheus server automatically.
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
From the Go specification [1]:
"1. For a nil slice, the number of iterations is 0."
"3. If the map is nil, the number of iterations is 0."
`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.
[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len
Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
Add the ability to specify node profiles in the multus validation test.
This addresses a few points of early feedback on the validation tool.
Statements below critique the tool's behavior before this patch.
1. The tool assumes all daemons are on public and cluster network, which
means users who have a significantly smaller cluster net (a
design choice) cannot run a single test to determine if Rook is
likely to install correctly.
2. The tool does not have placement options to select only a subset of
Kubernetes nodes to run validation on.
3. Users of multus seem to have a dedicated pool of storage nodes more
often than the average Rook install. This makes sense for security-
and perforance-minded users. The tool cannot run a single test to
verify storage-only and general-workload nodes at one time.
These points are addressed by allowing users to specify configurations
for different "NodeTypes."
Each NodeType config has options for selecting the number of OSDs as
well as the number of other (non-OSD) Ceph daemons. This limits the
unnecessary exhaustion of cluster network addresses from critique 1.
Each NodeType config has its own placement (critique 2).
Users can define as many NodeTypes as needed to test the network for
their planned CephCluster. Specifically, this allows the tool to test
storage-only nodes and generalized-workload nodes at the same time. An
arbitrary number of NodeTypes are allowed to support even more highly
specialized cluster setups, such as multiple tiers of storage nodes
where some storage-only nodes may run more OSDs than others.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The issues are fixed from the clean disk action
repo and we shouldn't require any workaround. Also,
I'm not using git revert to revert the commit is earlier
clean disk action was mentioned in every CI which was
duplicate and now we have placed the action to composite yaml
Signed-off-by: subhamkrai <srai@redhat.com>
Since `google-cloud-sdk` is renamed with `google-cloud-cli`,
the github action is failing to remove the older name and this
is causing ci issues. Applying changes suggested by community to
manually remove some packages to resolve the issue.
Signed-off-by: subhamkrai <srai@redhat.com>