forked from rook/rook
The CephUpgradeSuite workflow failed ~31% of PR runs vs the ~19%
broken-PR baseline of the other integration suites. Classifying the
failing step of every failed run from the last ~400 PR runs and the
logs of every failure on non-dependabot branches shows the excess
comes from a handful of too-tight waits, unretried network fetches,
and environment races rather than from the upgrade logic itself:
- The biggest class (21 of 44 genuine test-phase failure runs):
"giving up waiting for deployment(s) with label
app=rook-ceph-{rgw,osd},ceph-version!=..." during the
Squid->Tentacle upgrade. The operator updates daemons sequentially
(mons, mgr, osds with PG health gates between them, mds, rgw), so
each later daemon's fixed 275s wait also absorbs the time spent on
every daemon before it; rgw, last in the sequence, failed most
often. The mon wait was already extended for slow image pulls in
98bd03a42; extend the same retry count to all the daemon waits.
- ~7 runs: "snapshot controller is not ready" in the Helm upgrade
path. WaitForSnapshotController(30) allows only 150s, and the logs
show the deployment still converging (readyreplicas 1 < replicas 2)
on the final poll. Raise to 90 retries.
- InstallOrUpgradeHelmRepoChart ran helm with no retry; an observed
failure fetched the ceph-csi-drivers chart tarball from GitHub
release assets and got a 504. Retry up to 5 times, as
InstallLocalHelmChart already does.
- Raise the go test timeouts (2400s->3600s rook, 1800s->2700s helm).
Failed runs ended in "panic: test timed out" during teardown, which
aborts cleanup and log collection; the longer daemon waits above
also need the headroom.
- The "setup cluster resources" composite action (shared by all
suites) accounted for half of all failed jobs. The one steady class
there (8 distinct runs across 7 different days): minikube exits with
K8S_FAIL_CONNECT (code 40) when the requested kubernetes version is
missing from its bundled version list, because it then validates the
version with an anonymous GitHub API request (GITHUB_TOKEN is not
honored on that code path), and anonymous requests from shared
runner IPs are regularly rate-limited. minikube 1.38.x predates
v1.35.5, so only the v1.35.5 jobs hit this class. Pass --force to
minikube start to skip that check: the kubernetes versions used in
CI are pinned constants already validated by the PRs that bump
them, and an invalid version would still fail fast at the kubeadm
download. Also seen twice: dpkg failing on a corrupt cri-dockerd .deb
because curl ran without --fail and saved an HTTP error page as the
package; add --fail and retries.
- On runners without the /mnt resource disk, the fallback OSD disk is
an iSCSI (LIO) device and use_local_disk_for_integration_test
returned early on those runners, skipping the udev nowatch
workaround for the device re-probe storms of rook issue 8975. A burn-in
failure on this PR captured the consequence with full kernel
forensics: ~100 udev change events on the OSD disk, and ceph-volume
activate wedged in uninterruptible sleep on the block device lock
(blkdev_llseek -> rwsem_down_write_slowpath) for 18+ minutes during
the ceph version upgrade while the cluster's only OSD stayed down.
Install the nowatch rule before the early return so it applies to
every runner type, and before the disk is first written rather than
after.
- One burn-in round failed before the upgrade even began: the
pre-upgrade PVC create on the Helm path expired WaitUntilPVCIsBound
(RetryLoop, 275s) while the CSI provisioner pods were still
starting. Set RETRY_MAX=110 for this workflow to double the
framework's base wait budget, instead of extending individual waits
one flake at a time.
- minikube also exits with GUEST_START (code 80) when its internal 6
minute node-ready wait expires on slow runners (5 runs in the
dataset, two more observed while burning in this PR, one of them on
another suite). Pass --wait-timeout=15m to minikube start.
Not addressed (small or episodic classes): OSDs never coming up on
initial deploy (3 runs, possibly the same udev/iSCSI wedge), the
filesystem not becoming active (2 runs), and several day-clustered
minikube incidents (exit codes 65/67/90).
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>