Files
my-rook-config/tests/scripts
Joshua Hoblitt cff8b7f311 ci: reduce CephUpgradeSuite flakes
The CephUpgradeSuite workflow failed ~31% of PR runs vs the ~19%
broken-PR baseline of the other integration suites. Classifying the
failing step of every failed run from the last ~400 PR runs and the
logs of every failure on non-dependabot branches shows the excess
comes from a handful of too-tight waits, unretried network fetches,
and environment races rather than from the upgrade logic itself:

- The biggest class (21 of 44 genuine test-phase failure runs):
  "giving up waiting for deployment(s) with label
  app=rook-ceph-{rgw,osd},ceph-version!=..." during the
  Squid->Tentacle upgrade. The operator updates daemons sequentially
  (mons, mgr, osds with PG health gates between them, mds, rgw), so
  each later daemon's fixed 275s wait also absorbs the time spent on
  every daemon before it; rgw, last in the sequence, failed most
  often. The mon wait was already extended for slow image pulls in
  98bd03a42; extend the same retry count to all the daemon waits.

- ~7 runs: "snapshot controller is not ready" in the Helm upgrade
  path. WaitForSnapshotController(30) allows only 150s, and the logs
  show the deployment still converging (readyreplicas 1 < replicas 2)
  on the final poll. Raise to 90 retries.

- InstallOrUpgradeHelmRepoChart ran helm with no retry; an observed
  failure fetched the ceph-csi-drivers chart tarball from GitHub
  release assets and got a 504. Retry up to 5 times, as
  InstallLocalHelmChart already does.

- Raise the go test timeouts (2400s->3600s rook, 1800s->2700s helm).
  Failed runs ended in "panic: test timed out" during teardown, which
  aborts cleanup and log collection; the longer daemon waits above
  also need the headroom.

- The "setup cluster resources" composite action (shared by all
  suites) accounted for half of all failed jobs. The one steady class
  there (8 distinct runs across 7 different days): minikube exits with
  K8S_FAIL_CONNECT (code 40) when the requested kubernetes version is
  missing from its bundled version list, because it then validates the
  version with an anonymous GitHub API request (GITHUB_TOKEN is not
  honored on that code path), and anonymous requests from shared
  runner IPs are regularly rate-limited. minikube 1.38.x predates
  v1.35.5, so only the v1.35.5 jobs hit this class. Pass --force to
  minikube start to skip that check: the kubernetes versions used in
  CI are pinned constants already validated by the PRs that bump
  them, and an invalid version would still fail fast at the kubeadm
  download. Also seen twice: dpkg failing on a corrupt cri-dockerd .deb
  because curl ran without --fail and saved an HTTP error page as the
  package; add --fail and retries.

- On runners without the /mnt resource disk, the fallback OSD disk is
  an iSCSI (LIO) device and use_local_disk_for_integration_test
  returned early on those runners, skipping the udev nowatch
  workaround for the device re-probe storms of rook issue 8975. A burn-in
  failure on this PR captured the consequence with full kernel
  forensics: ~100 udev change events on the OSD disk, and ceph-volume
  activate wedged in uninterruptible sleep on the block device lock
  (blkdev_llseek -> rwsem_down_write_slowpath) for 18+ minutes during
  the ceph version upgrade while the cluster's only OSD stayed down.
  Install the nowatch rule before the early return so it applies to
  every runner type, and before the disk is first written rather than
  after.

- One burn-in round failed before the upgrade even began: the
  pre-upgrade PVC create on the Helm path expired WaitUntilPVCIsBound
  (RetryLoop, 275s) while the CSI provisioner pods were still
  starting. Set RETRY_MAX=110 for this workflow to double the
  framework's base wait budget, instead of extending individual waits
  one flake at a time.

- minikube also exits with GUEST_START (code 80) when its internal 6
  minute node-ready wait expires on slow runners (5 runs in the
  dataset, two more observed while burning in this PR, one of them on
  another suite). Pass --wait-timeout=15m to minikube start.

Not addressed (small or episodic classes): OSDs never coming up on
initial deploy (3 runs, possibly the same udev/iSCSI wedge), the
filesystem not becoming active (2 runs), and several day-clustered
minikube incidents (exit codes 65/67/90).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 15:27:02 -07:00
..
2025-01-21 15:11:33 -07:00
2026-01-16 11:50:18 +05:30
2023-05-19 16:59:23 -06:00