Files
Joshua Hoblitt 4da654bae2 ci: fix multus integration test setup flakes
Multus CI failures on master pushes from 2026-04-08 to 2026-06-10 were
dominated by a race in setup-multus.sh: 'kubectl wait' was invoked on a
pod label selector immediately after 'kubectl create' of a daemonset,
and it exits with "error: no matching resources found" when the
daemonset controller has not created any pods yet. This caused 9 of the
10 genuine multus flakes across the standalone multus workflow and the
canary multus-public-and-cluster job (the two share setup-multus.sh).

- setup-multus.sh: wait with 'kubectl rollout status' on the daemonsets
  instead of 'kubectl wait' on pod label selectors. The daemonset
  object exists as soon as 'kubectl create' returns, and rollout status
  correctly waits for all desired pods to be created and become ready.
  The old selector wait also silently under-waited: pods created after
  its initial LIST were never waited on at all. The stricter wait
  revealed that full multi-node convergence can exceed 2 minutes on
  busy runners, so the wait timeout is raised to 5 minutes (the waits
  return as soon as the rollouts are ready, so this costs nothing on
  healthy runs).

- Wait for the host-net-config daemonset rollout in both workflows
  before proceeding; it configures the host routing that the multus
  public network depends on and was previously not waited on at all.
  Pin its jonlabelle/network-tools image.

- test_multus_connections: retry the osd dump / fs dump network checks
  for up to 2 minutes. Daemons register their addresses in the mon maps
  asynchronously after the cluster reports ready; one canary failure
  (2026-04-29) ran the MDS check at fsmap epoch 1, before any MDS had
  registered.

- Dump cluster state (pods, daemonsets, events, multus logs) when the
  multus workflow job fails; the workflow previously had no failure
  diagnostics at all.

Also remove the unused NUMBER_OF_COMPUTE_NODES env var from the multus
workflow (kind-config.yaml creates 3 workers; nothing consumes it).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-10 21:12:46 -07:00

59 lines
2.2 KiB
Bash
Executable File

#!/usr/bin/env -S bash
set -xeo pipefail
# initially copied from https://github.com/k8snetworkplumbingwg/whereabouts
MULTUS_DAEMONSET_URL="https://raw.githubusercontent.com/k8snetworkplumbingwg/multus-cni/master/deployments/multus-daemonset.yml"
RETRY_MAX=10
INTERVAL=10
TIMEOUT=60
# image pulls plus CNI bootstrap across all nodes can take a few minutes on
# busy CI runners; the waits below return as soon as the rollouts are ready
TIMEOUT_K8="300s"
retry() {
local status=0
local retries=${RETRY_MAX:=5}
local delay=${INTERVAL:=5}
local to=${TIMEOUT:=20}
cmd="$*"
while [ $retries -gt 0 ]; do
status=0
timeout $to bash -c "echo $cmd && $cmd" || status=$?
if [ $status -eq 0 ]; then
break
fi
echo "Exit code: '$status'. Sleeping '$delay' seconds before retrying"
sleep $delay
let retries--
done
return $status
}
echo "#### set up multus ####"
echo " ## wait for coreDNS"
kubectl -n kube-system wait --for=condition=available deploy/coredns --timeout=$TIMEOUT_K8
echo "## install multus"
retry kubectl create -f "${MULTUS_DAEMONSET_URL}"
# use 'rollout status' on the daemonset instead of 'kubectl wait' on a pod
# label selector: right after 'kubectl create' the daemonset controller may
# not have created any pods yet, and 'kubectl wait' exits immediately with
# "no matching resources found" when the selector matches nothing
kubectl -n kube-system rollout status daemonset/kube-multus-ds --timeout=$TIMEOUT_K8
echo "## install CNIs"
retry kubectl create -f "https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/master/hack/cni-install.yml"
kubectl -n kube-system rollout status daemonset/install-cni-plugins --timeout=$TIMEOUT_K8
echo "## install whereabouts"
kubectl create \
-f https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/master/doc/crds/daemonset-install.yaml \
-f https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/master/doc/crds/whereabouts.cni.cncf.io_ippools.yaml \
-f https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/master/doc/crds/whereabouts.cni.cncf.io_overlappingrangeipreservations.yaml
kubectl -n kube-system rollout status daemonset/whereabouts --timeout=$TIMEOUT_K8
echo "#### set up multus done ####"