forked from rook/rook
Multus CI failures on master pushes from 2026-04-08 to 2026-06-10 were dominated by a race in setup-multus.sh: 'kubectl wait' was invoked on a pod label selector immediately after 'kubectl create' of a daemonset, and it exits with "error: no matching resources found" when the daemonset controller has not created any pods yet. This caused 9 of the 10 genuine multus flakes across the standalone multus workflow and the canary multus-public-and-cluster job (the two share setup-multus.sh). - setup-multus.sh: wait with 'kubectl rollout status' on the daemonsets instead of 'kubectl wait' on pod label selectors. The daemonset object exists as soon as 'kubectl create' returns, and rollout status correctly waits for all desired pods to be created and become ready. The old selector wait also silently under-waited: pods created after its initial LIST were never waited on at all. The stricter wait revealed that full multi-node convergence can exceed 2 minutes on busy runners, so the wait timeout is raised to 5 minutes (the waits return as soon as the rollouts are ready, so this costs nothing on healthy runs). - Wait for the host-net-config daemonset rollout in both workflows before proceeding; it configures the host routing that the multus public network depends on and was previously not waited on at all. Pin its jonlabelle/network-tools image. - test_multus_connections: retry the osd dump / fs dump network checks for up to 2 minutes. Daemons register their addresses in the mon maps asynchronously after the cluster reports ready; one canary failure (2026-04-29) ran the MDS check at fsmap epoch 1, before any MDS had registered. - Dump cluster state (pods, daemonsets, events, multus logs) when the multus workflow job fails; the workflow previously had no failure diagnostics at all. Also remove the unused NUMBER_OF_COMPUTE_NODES env var from the multus workflow (kind-config.yaml creates 3 workers; nothing consumes it). Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
59 lines
2.2 KiB
Bash
Executable File
59 lines
2.2 KiB
Bash
Executable File
#!/usr/bin/env -S bash
|
|
set -xeo pipefail
|
|
|
|
# initially copied from https://github.com/k8snetworkplumbingwg/whereabouts
|
|
|
|
MULTUS_DAEMONSET_URL="https://raw.githubusercontent.com/k8snetworkplumbingwg/multus-cni/master/deployments/multus-daemonset.yml"
|
|
RETRY_MAX=10
|
|
INTERVAL=10
|
|
TIMEOUT=60
|
|
# image pulls plus CNI bootstrap across all nodes can take a few minutes on
|
|
# busy CI runners; the waits below return as soon as the rollouts are ready
|
|
TIMEOUT_K8="300s"
|
|
|
|
retry() {
|
|
local status=0
|
|
local retries=${RETRY_MAX:=5}
|
|
local delay=${INTERVAL:=5}
|
|
local to=${TIMEOUT:=20}
|
|
cmd="$*"
|
|
|
|
while [ $retries -gt 0 ]; do
|
|
status=0
|
|
timeout $to bash -c "echo $cmd && $cmd" || status=$?
|
|
if [ $status -eq 0 ]; then
|
|
break
|
|
fi
|
|
echo "Exit code: '$status'. Sleeping '$delay' seconds before retrying"
|
|
sleep $delay
|
|
let retries--
|
|
done
|
|
return $status
|
|
}
|
|
|
|
echo "#### set up multus ####"
|
|
|
|
echo " ## wait for coreDNS"
|
|
kubectl -n kube-system wait --for=condition=available deploy/coredns --timeout=$TIMEOUT_K8
|
|
|
|
echo "## install multus"
|
|
retry kubectl create -f "${MULTUS_DAEMONSET_URL}"
|
|
# use 'rollout status' on the daemonset instead of 'kubectl wait' on a pod
|
|
# label selector: right after 'kubectl create' the daemonset controller may
|
|
# not have created any pods yet, and 'kubectl wait' exits immediately with
|
|
# "no matching resources found" when the selector matches nothing
|
|
kubectl -n kube-system rollout status daemonset/kube-multus-ds --timeout=$TIMEOUT_K8
|
|
|
|
echo "## install CNIs"
|
|
retry kubectl create -f "https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/master/hack/cni-install.yml"
|
|
kubectl -n kube-system rollout status daemonset/install-cni-plugins --timeout=$TIMEOUT_K8
|
|
|
|
echo "## install whereabouts"
|
|
kubectl create \
|
|
-f https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/master/doc/crds/daemonset-install.yaml \
|
|
-f https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/master/doc/crds/whereabouts.cni.cncf.io_ippools.yaml \
|
|
-f https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/master/doc/crds/whereabouts.cni.cncf.io_overlappingrangeipreservations.yaml
|
|
kubectl -n kube-system rollout status daemonset/whereabouts --timeout=$TIMEOUT_K8
|
|
|
|
echo "#### set up multus done ####"
|