The integration and canary suites ran on a single-node minikube
`driver: none` cluster, where the kubelet ran directly on the GitHub
runner, so the host docker daemon doubled as the cluster runtime and
host block devices and host paths were directly visible to pods. Replace
that with a kind cluster.
Every suite creates its cluster through the shared
integration-test-setup-cluster-resources composite action, so the
conversion is centralized there and converts the smoke, object, helm,
keystone, multi-cluster, upgrade, on-release, nightly, encryption-KMS and
all canary jobs at once.
- Replace the setup-minikube step with helm/kind-action, selecting the
kubernetes version via the kindest/node image tag and creating a
single-node cluster from a new kind config (kind pinned to v0.32.0 for
reproducibility).
- Drop the cri-dockerd install; kind nodes use their built-in containerd.
- Add a kind config that bind-mounts the host /dev, /var/lib/rook and
/run/udev into the node so the existing host-based disk-prep helpers
(use_local_disk*, create_partitions_for_osds, blockDevicePV.sh,
localPathPV.sh, ...) keep working unchanged: devices and partitions
created on the host appear in the node and in the OSD pods that
hostPath-mount the node /dev, and ceph-volume can read the host udev
database it needs to inventory disks.
- Prepare the kind node for the host-level operations rook runs against the
underlying host: remount /sys read-write so CSI's kernel RBD mapping
(`rbd map --device-type krbd`, which writes /sys/bus/rbd) works, and install
lvm2 and cryptsetup, which rook runs in the node's mount namespace to
provision LVM- and encryption-backed OSDs. kindest/node images provide none
of this; the minikube driver:none runner host did.
- Route the Service and pod CIDRs from the runner to the kind node so
host-side tests (the `go test` process runs on the runner) can reach
in-cluster ClusterIPs, e.g. an S3 request to the RGW service. With minikube
driver:none the runner already shared the cluster network.
- Load locally built images into the cluster. Under minikube `driver: none`
the built image was already in the cluster runtime; under kind it must be
imported, so build_rook and create_helm_tag now import their images into
each node's containerd through a new load_image_into_cluster helper (via the
node's ctr, which avoids the kind/kindest-node containerd-config version
skew that breaks `kind load docker-image`).
- Point Vault's kubernetes-auth at the in-cluster API endpoint
(kubernetes.default.svc) instead of the kubeconfig server URL: kind exposes
that as https://127.0.0.1:<port>, unreachable from the in-cluster Vault pod,
so OSD encryption-key retrieval via k8s-auth failed.
- Replace the remaining direct minikube references in the canary workflow: a
`minikube kubectl` call and the external-cluster topology values.
- Adapt host-name assumptions that only held under driver:none: resolve the
disk-cleanup job by the k8s node name rather than the runner hostname, and
let kind-action ignore post-job cluster-teardown failures (the runner is
ephemeral; nvme/multus devices can wedge `docker rm` of the node).
- Update stale comments that described the CI environment as minikube.
- Move the multus integration test's kind config under tests/config too, so
both kind cluster configs live in one place.
create-dev-cluster.sh and other local-dev tooling are intentionally left on
minikube.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The default service account isn't passed to the multus validation test
when no config file is used. Fix this.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Add the ability to specify node profiles in the multus validation test.
This addresses a few points of early feedback on the validation tool.
Statements below critique the tool's behavior before this patch.
1. The tool assumes all daemons are on public and cluster network, which
means users who have a significantly smaller cluster net (a
design choice) cannot run a single test to determine if Rook is
likely to install correctly.
2. The tool does not have placement options to select only a subset of
Kubernetes nodes to run validation on.
3. Users of multus seem to have a dedicated pool of storage nodes more
often than the average Rook install. This makes sense for security-
and perforance-minded users. The tool cannot run a single test to
verify storage-only and general-workload nodes at one time.
These points are addressed by allowing users to specify configurations
for different "NodeTypes."
Each NodeType config has options for selecting the number of OSDs as
well as the number of other (non-OSD) Ceph daemons. This limits the
unnecessary exhaustion of cluster network addresses from critique 1.
Each NodeType config has its own placement (critique 2).
Users can define as many NodeTypes as needed to test the network for
their planned CephCluster. Specifically, this allows the tool to test
storage-only nodes and generalized-workload nodes at the same time. An
arbitrary number of NodeTypes are allowed to support even more highly
specialized cluster setups, such as multiple tiers of storage nodes
where some storage-only nodes may run more OSDs than others.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>