13 Commits
Author SHA1 Message Date
Joshua Hoblitt d59abb1956 ci: convert integration and canary CI from minikube to kind
The integration and canary suites ran on a single-node minikube
`driver: none` cluster, where the kubelet ran directly on the GitHub
runner, so the host docker daemon doubled as the cluster runtime and
host block devices and host paths were directly visible to pods. Replace
that with a kind cluster.

Every suite creates its cluster through the shared
integration-test-setup-cluster-resources composite action, so the
conversion is centralized there and converts the smoke, object, helm,
keystone, multi-cluster, upgrade, on-release, nightly, encryption-KMS and
all canary jobs at once.

- Replace the setup-minikube step with helm/kind-action, selecting the
  kubernetes version via the kindest/node image tag and creating a
  single-node cluster from a new kind config (kind pinned to v0.32.0 for
  reproducibility).
- Drop the cri-dockerd install; kind nodes use their built-in containerd.
- Add a kind config that bind-mounts the host /dev, /var/lib/rook and
  /run/udev into the node so the existing host-based disk-prep helpers
  (use_local_disk*, create_partitions_for_osds, blockDevicePV.sh,
  localPathPV.sh, ...) keep working unchanged: devices and partitions
  created on the host appear in the node and in the OSD pods that
  hostPath-mount the node /dev, and ceph-volume can read the host udev
  database it needs to inventory disks.
- Prepare the kind node for the host-level operations rook runs against the
  underlying host: remount /sys read-write so CSI's kernel RBD mapping
  (`rbd map --device-type krbd`, which writes /sys/bus/rbd) works, and install
  lvm2 and cryptsetup, which rook runs in the node's mount namespace to
  provision LVM- and encryption-backed OSDs. kindest/node images provide none
  of this; the minikube driver:none runner host did.
- Route the Service and pod CIDRs from the runner to the kind node so
  host-side tests (the `go test` process runs on the runner) can reach
  in-cluster ClusterIPs, e.g. an S3 request to the RGW service. With minikube
  driver:none the runner already shared the cluster network.
- Load locally built images into the cluster. Under minikube `driver: none`
  the built image was already in the cluster runtime; under kind it must be
  imported, so build_rook and create_helm_tag now import their images into
  each node's containerd through a new load_image_into_cluster helper (via the
  node's ctr, which avoids the kind/kindest-node containerd-config version
  skew that breaks `kind load docker-image`).
- Point Vault's kubernetes-auth at the in-cluster API endpoint
  (kubernetes.default.svc) instead of the kubeconfig server URL: kind exposes
  that as https://127.0.0.1:<port>, unreachable from the in-cluster Vault pod,
  so OSD encryption-key retrieval via k8s-auth failed.
- Replace the remaining direct minikube references in the canary workflow: a
  `minikube kubectl` call and the external-cluster topology values.
- Adapt host-name assumptions that only held under driver:none: resolve the
  disk-cleanup job by the k8s node name rather than the runner hostname, and
  let kind-action ignore post-job cluster-teardown failures (the runner is
  ephemeral; nvme/multus devices can wedge `docker rm` of the node).
- Update stale comments that described the CI environment as minikube.
- Move the multus integration test's kind config under tests/config too, so
  both kind cluster configs live in one place.

create-dev-cluster.sh and other local-dev tooling are intentionally left on
minikube.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-06 08:42:39 -07:00
Joshua Hoblitt 4da654bae2 ci: fix multus integration test setup flakes
Multus CI failures on master pushes from 2026-04-08 to 2026-06-10 were
dominated by a race in setup-multus.sh: 'kubectl wait' was invoked on a
pod label selector immediately after 'kubectl create' of a daemonset,
and it exits with "error: no matching resources found" when the
daemonset controller has not created any pods yet. This caused 9 of the
10 genuine multus flakes across the standalone multus workflow and the
canary multus-public-and-cluster job (the two share setup-multus.sh).

- setup-multus.sh: wait with 'kubectl rollout status' on the daemonsets
  instead of 'kubectl wait' on pod label selectors. The daemonset
  object exists as soon as 'kubectl create' returns, and rollout status
  correctly waits for all desired pods to be created and become ready.
  The old selector wait also silently under-waited: pods created after
  its initial LIST were never waited on at all. The stricter wait
  revealed that full multi-node convergence can exceed 2 minutes on
  busy runners, so the wait timeout is raised to 5 minutes (the waits
  return as soon as the rollouts are ready, so this costs nothing on
  healthy runs).

- Wait for the host-net-config daemonset rollout in both workflows
  before proceeding; it configures the host routing that the multus
  public network depends on and was previously not waited on at all.
  Pin its jonlabelle/network-tools image.

- test_multus_connections: retry the osd dump / fs dump network checks
  for up to 2 minutes. Daemons register their addresses in the mon maps
  asynchronously after the cluster reports ready; one canary failure
  (2026-04-29) ran the MDS check at fsmap epoch 1, before any MDS had
  registered.

- Dump cluster state (pods, daemonsets, events, multus logs) when the
  multus workflow job fails; the workflow previously had no failure
  diagnostics at all.

Also remove the unused NUMBER_OF_COMPUTE_NODES env var from the multus
workflow (kind-config.yaml creates 3 workers; nothing consumes it).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-10 21:12:46 -07:00
Joshua Hoblitt f52f48c477 ci: codespell: s/re-using/reusing/
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2025-01-21 15:11:33 -07:00
Blaine Gardner 5539eedd1b multus: finish deprecating holder pods
Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-10-23 16:29:02 -06:00
Blaine Gardner 58e3feaacf multus: fix default service account handling
The default service account isn't passed to the multus validation test
when no config file is used. Fix this.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-08-21 16:42:47 -06:00
Blaine Gardner 5773132d7f ci: fix failing multus validation tool test
Ceph image no longer has `ip` tool installed. Use a different container
image for the daemonset which sets host IPs and routes for multus hosts.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-08-14 12:25:44 -06:00
Blaine Gardner 33f5407dd4 multus: add host checking to validation tool
In order to help users check that they have implemented the newly-added
Multus host configuration prerequisites, add a check to the validation
tool to verify connectivity.

Because users who are already running clusters with Multus enabled, add
a flag that allows users to only check for host configuration
prerequisites. This mode will not start the large number of clients that
would normally be started because those clients could disrupt a running
Rook cluster negatively.

Host checking pods require host network access. Many Kubernetes
distributions have pod security features enabled. In order to allow
non-Vanilla distros to run this tool, allow specifying a service account
that pods will run as, which can be configured by the admin to allow
test pods.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-07-10 13:41:38 -06:00
Blaine Gardner 08a27940ed multus: add and test ipv6 support for validation tool
Add IPv6 support for multus validation tool. Also test that IPv6 support
works by specifying one of the NetAttachDefs with an IPv6 address range.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-06-03 17:30:36 -06:00
Blaine Gardner cd253d0218 multus: use nginx-unprivileged image from quay
Use Quay as the source for the nginx-unprivileged image used for the
Multus validation tool because Quay does not rate limit image pulls,
which are a common complaint for users of the tool.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-01-04 14:25:40 -07:00
Alexander Trost fcf52207f8 build: use /usr/bin/env to look up script interpreters
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-11-21 14:38:22 +01:00
Sheetal Pamecha 06c176524a multus: improve the multus validation test's flakiness metric
Allow the flakiness threshold window to be tuned from the cli

Signed-off-by: Sheetal Pamecha <spamecha@redhat.com>
2023-10-31 15:23:10 +05:30
Blaine Gardner 0c721e05d5 multus: allow node profiles in validation test
Add the ability to specify node profiles in the multus validation test.

This addresses a few points of early feedback on the validation tool.
Statements below critique the tool's behavior before this patch.
1. The tool assumes all daemons are on public and cluster network, which
   means users who have a significantly smaller cluster net (a
   design choice) cannot run a single test to determine if Rook is
   likely to install correctly.
2. The tool does not have placement options to select only a subset of
   Kubernetes nodes to run validation on.
3. Users of multus seem to have a dedicated pool of storage nodes more
   often than the average Rook install. This makes sense for security-
   and perforance-minded users. The tool cannot run a single test to
   verify storage-only and general-workload nodes at one time.

These points are addressed by allowing users to specify configurations
for different "NodeTypes."

Each NodeType config has options for selecting the number of OSDs as
well as the number of other (non-OSD) Ceph daemons. This limits the
unnecessary exhaustion of cluster network addresses from critique 1.

Each NodeType config has its own placement (critique 2).

Users can define as many NodeTypes as needed to test the network for
their planned CephCluster. Specifically, this allows the tool to test
storage-only nodes and generalized-workload nodes at the same time. An
arbitrary number of NodeTypes are allowed to support even more highly
specialized cluster setups, such as multiple tiers of storage nodes
where some storage-only nodes may run more OSDs than others.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-10-04 15:26:35 -06:00
Blaine Gardner 2eb5a9a3f6 test: add CI e2e test for multus validation test
Add a CI e2e test for the multus validation routine that runs whenever
the multus validation test is modified and on master/releases.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2023-05-26 12:19:54 -06:00