Add a single example YAML with NetworkPolicy definitions
for all Ceph operand pods (mon, osd, mgr, mds, exporter,
osd-prepare, crashcollector, tools).
Users can apply these policies to restrict egress/ingress
traffic for Rook-Ceph daemon pods.
Signed-off-by: Oded Viner <oviner@redhat.com>
The CephUpgradeSuite workflow failed ~31% of PR runs vs the ~19%
broken-PR baseline of the other integration suites. Classifying the
failing step of every failed run from the last ~400 PR runs and the
logs of every failure on non-dependabot branches shows the excess
comes from a handful of too-tight waits, unretried network fetches,
and environment races rather than from the upgrade logic itself:
- The biggest class (21 of 44 genuine test-phase failure runs):
"giving up waiting for deployment(s) with label
app=rook-ceph-{rgw,osd},ceph-version!=..." during the
Squid->Tentacle upgrade. The operator updates daemons sequentially
(mons, mgr, osds with PG health gates between them, mds, rgw), so
each later daemon's fixed 275s wait also absorbs the time spent on
every daemon before it; rgw, last in the sequence, failed most
often. The mon wait was already extended for slow image pulls in
98bd03a42; extend the same retry count to all the daemon waits.
- ~7 runs: "snapshot controller is not ready" in the Helm upgrade
path. WaitForSnapshotController(30) allows only 150s, and the logs
show the deployment still converging (readyreplicas 1 < replicas 2)
on the final poll. Raise to 90 retries.
- InstallOrUpgradeHelmRepoChart ran helm with no retry; an observed
failure fetched the ceph-csi-drivers chart tarball from GitHub
release assets and got a 504. Retry up to 5 times, as
InstallLocalHelmChart already does.
- Raise the go test timeouts (2400s->3600s rook, 1800s->2700s helm).
Failed runs ended in "panic: test timed out" during teardown, which
aborts cleanup and log collection; the longer daemon waits above
also need the headroom.
- The "setup cluster resources" composite action (shared by all
suites) accounted for half of all failed jobs. The one steady class
there (8 distinct runs across 7 different days): minikube exits with
K8S_FAIL_CONNECT (code 40) when the requested kubernetes version is
missing from its bundled version list, because it then validates the
version with an anonymous GitHub API request (GITHUB_TOKEN is not
honored on that code path), and anonymous requests from shared
runner IPs are regularly rate-limited. minikube 1.38.x predates
v1.35.5, so only the v1.35.5 jobs hit this class. Pass --force to
minikube start to skip that check: the kubernetes versions used in
CI are pinned constants already validated by the PRs that bump
them, and an invalid version would still fail fast at the kubeadm
download. Also seen twice: dpkg failing on a corrupt cri-dockerd .deb
because curl ran without --fail and saved an HTTP error page as the
package; add --fail and retries.
- On runners without the /mnt resource disk, the fallback OSD disk is
an iSCSI (LIO) device and use_local_disk_for_integration_test
returned early on those runners, skipping the udev nowatch
workaround for the device re-probe storms of rook issue 8975. A burn-in
failure on this PR captured the consequence with full kernel
forensics: ~100 udev change events on the OSD disk, and ceph-volume
activate wedged in uninterruptible sleep on the block device lock
(blkdev_llseek -> rwsem_down_write_slowpath) for 18+ minutes during
the ceph version upgrade while the cluster's only OSD stayed down.
Install the nowatch rule before the early return so it applies to
every runner type, and before the disk is first written rather than
after.
- One burn-in round failed before the upgrade even began: the
pre-upgrade PVC create on the Helm path expired WaitUntilPVCIsBound
(RetryLoop, 275s) while the CSI provisioner pods were still
starting. Set RETRY_MAX=110 for this workflow to double the
framework's base wait budget, instead of extending individual waits
one flake at a time.
- minikube also exits with GUEST_START (code 80) when its internal 6
minute node-ready wait expires on slow runners (5 runs in the
dataset, two more observed while burning in this PR, one of them on
another suite). Pass --wait-timeout=15m to minikube start.
Not addressed (small or episodic classes): OSDs never coming up on
initial deploy (3 runs, possibly the same udev/iSCSI wedge), the
filesystem not becoming active (2 runs), and several day-clustered
minikube incidents (exit codes 65/67/90).
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
validate_cluster.sh built its ceph exec command around a toolbox pod
name captured once at script start, without waiting for the toolbox to
be ready. When the toolbox container was still being created, or the
pod had been replaced, every check failed with 'unable to upgrade
connection: container not found' for the entire wait window and the
mon quorum check timed out without ever observing the cluster. This
intermittently failed the 'wait for ceph cluster N to be ready' steps
of canary jobs (observed on master pushes of multi-cluster-mirroring
and the canary job between 2026-03-10 and 2026-04-24, always with the
container-not-found signature filling the whole window).
Wait for the toolbox deployment rollout up front, and exec through
deploy/rook-ceph-tools so each call resolves a currently-ready pod,
matching how the other CI scripts invoke the toolbox.
Also call wait_for_daemon directly instead of 'return $(...)'. The
command substitution ran wait_for_daemon in a subshell and captured
its output, so the 'current status' diagnostics printed on timeout
were never displayed; worse, the captured text became the argument of
'return', which fails with 'numeric argument required' and aborts the
script with a misleading exit code 2. The osd variant's captured
'Return value' debug echo had the same problem and is removed.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Multus CI failures on master pushes from 2026-04-08 to 2026-06-10 were
dominated by a race in setup-multus.sh: 'kubectl wait' was invoked on a
pod label selector immediately after 'kubectl create' of a daemonset,
and it exits with "error: no matching resources found" when the
daemonset controller has not created any pods yet. This caused 9 of the
10 genuine multus flakes across the standalone multus workflow and the
canary multus-public-and-cluster job (the two share setup-multus.sh).
- setup-multus.sh: wait with 'kubectl rollout status' on the daemonsets
instead of 'kubectl wait' on pod label selectors. The daemonset
object exists as soon as 'kubectl create' returns, and rollout status
correctly waits for all desired pods to be created and become ready.
The old selector wait also silently under-waited: pods created after
its initial LIST were never waited on at all. The stricter wait
revealed that full multi-node convergence can exceed 2 minutes on
busy runners, so the wait timeout is raised to 5 minutes (the waits
return as soon as the rollouts are ready, so this costs nothing on
healthy runs).
- Wait for the host-net-config daemonset rollout in both workflows
before proceeding; it configures the host routing that the multus
public network depends on and was previously not waited on at all.
Pin its jonlabelle/network-tools image.
- test_multus_connections: retry the osd dump / fs dump network checks
for up to 2 minutes. Daemons register their addresses in the mon maps
asynchronously after the cluster reports ready; one canary failure
(2026-04-29) ran the MDS check at fsmap epoch 1, before any MDS had
registered.
- Dump cluster state (pods, daemonsets, events, multus logs) when the
multus workflow job fails; the workflow previously had no failure
diagnostics at all.
Also remove the unused NUMBER_OF_COMPUTE_NODES env var from the multus
workflow (kind-config.yaml creates 3 workers; nothing consumes it).
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The rgw-multisite-testing canary job has been chronically flaky
(roughly a third of recent canary runs, including on master and
release branches). Job logs and collected artifacts show two recurring
mechanisms, neither of which is addressed by bumping timeouts:
1. Every recent failure of the "write an object" step died within
~16s: the very first "s3cmd mb" against the primary zone got
ECONNREFUSED three times across the 10s retry window. The RGW
process never crashed; RGWs pause their HTTP frontends to reload
the realm whenever the RGW configuration period changes ("rgw
realm reloader: Pausing frontends for realm update..."), which
keeps happening shortly after the second zone joins the zonegroup,
exactly when the test starts writing. While paused, the readiness
probe also times out, the endpoint is removed from the Service, and
connections to the ClusterIP are refused.
2. The "kubectl wait --for=condition=available deployment/..."
barrier added previously is a no-op for these deployments: with one
replica and maxUnavailable=1, the availability threshold is zero
ready pods, so the condition is always true.
Changes:
- Wait on RGW pod readiness instead of deployment availability.
- Add a "wait for multisite sync to be established" step that polls
"radosgw-admin sync status" on both clusters until metadata and
data sync report caught up, mirroring the checkpoints ceph's own
multisite QA performs before asserting replication.
- Probe both RGW endpoints over HTTP until several consecutive
requests succeed before writing, to ride out realm-reload pauses.
- Replace the fixed 3-attempt s3cmd retries with time-budgeted
retries.
- Replace the s3cmd read polling loop and its SIGUSR1 timer (which was
never cancelled on success, silently shrinking the second
direction's time budget) with "radosgw-admin bucket sync
checkpoint", the marker-based barrier used by ceph QA, followed by
a single get and diff.
- Dump multisite diagnostics (pod state, RGW logs, sync status,
period) when the job fails.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
having csi-addons enabled by default is causing random
pod restart on non-openshift cluster. Let's disable
it by default.
Signed-off-by: subhamkrai <srai@redhat.com>
Two csi controllers cannot start on minikube, so the CI
was failing in the canary and nvmeof tests when they validated
that the controllers were running. For the tests that validate
the csi controllers, ensure only a single controller is configured.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The rgw-multisite job is failing frequently. Observed
"connectionRefusedERror: connection Refused" on the s3cmd mb call. It
might be happening because CephObjectSTore CR is ready but RGW gateway
is not fully accepting TCP connections.
This PR wraps s3cmd mb and s3cmd put calls in a retry loop that
will attempt each command up to 3 times with a 5 second sleep between
attemps. This might help fix the flaky CI test.
Signed-off-by: Santosh <sapillai@redhat.com>
Going forward, admin will manage the csi operator
CR's and rook will only manage Ceph Connection cr
and client Profile cr.
The old csi driver is completely removed from Rook
and can no longer be used starting in Rook v1.20.
The upgrade guide will contain the needed transition steps
for managing the csi operator settings.
Signed-off-by: subhamkrai <srai@redhat.com>
The canary test was waiting for replicapool
instead of replicapool2, and a more reliable
wait for the toolbox pod start is added.
Signed-off-by: subhamkrai <srai@redhat.com>
initialize both aws sdk v1 and v2 clients in s3agent.
update createbucket to use sdk v2 while keeping all
other methods on sdk v1.
Signed-off-by: Oded Viner <oviner@redhat.com>
update the nvmeof minikube canary test to use the
ceph-csi operator instead of manually deploying
the provisioner and node-plugin.
Signed-off-by: Oded Viner <oviner@redhat.com>
adds a new nvmeof minikube canary job for ceph v20.
the test deploys rook with csi operator disabled for nvmeof flow.
it validates pvc and pod io, gateway restart, and data persistence.
Signed-off-by: Oded Viner <oviner@redhat.com>
Kubernetes has begun releasing a kube-api-linter application for linting
Kubernetes APIs to match against common best-practices. (a.k.a., KAL).
Of note: KAL is still new and has no semver releases.
I believe this linter is helpful in encouraging Rook devs to think
deeply about how end-users consume the API. It ensures that new APIs
account for set/unset status, zero values, upper/lower bounds, and
possibly others in the future.
This commit adds initial support to Rook, without making the results a
requirement. Maintainers still need to evaluate rule configuration and
determine how tightly Rook wants to hold to the results.
As of this commit, I observe around 650 "issues" with the current Rook APIs.
This doesn't mean that Rook is in danger of breaking. Issues with API
types are often just indications that upper/lower bounds are missing, or
that there is ambiguity between set-empty and unset-empty values.
This commit uses `--new` in CI to ensure only new API issues are
flagged.
Also of note, KAL is a golangci-lint plugin. It is possible to locally
build KAL (from source) into golangci-lint to run both at the same time.
However, this is time-consuming, and it's harder to coordinate tool
binary caching this way when it's related to 2 different versions.
Treating KAL as a standalone binary allows for simpler make scripting
and maintenance.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
using github action for minikube will avoid
unwanted errors and probably will be more stable
than manual install. But both script and action
both uses almost same time for installation so
there we don't have preference.
Signed-off-by: subhamkrai <srai@redhat.com>
let's run daily nightly canary job with github action
arm runner as self-hosted runners is shutting down which
was provided by upstream user.
Signed-off-by: subhamkrai <srai@redhat.com>
The helm charts allowed rendering a PodSecurityPolicy resource via the
configuration `pspEnable`. This option is removed and all references to
psp, PodSecurityPolicy, and Pod Security Policy have been cleaned up.
The PSP resource was only rendered if k8s version was lower than 1.25
when it was still supported. It has been deprecated since k8s 1.21.
Signed-off-by: Erik Sundell <erik@sundellopensource.se>
In the deb package's location of /usr/bin/minilube, a wrong version seems to be
reported but from /usr/local/bin it reports correctly.
Signed-off-by: Michael Adam <obnox@samba.org>
this change updates the k8s version to 1.34 and also updates
the cri-ctl version for minikube.
Signed-off-by: Michael Adam <obnox@samba.org>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Currently create-dev-cluster.sh will hang as the rook operator gets stuck like so:
```
2025-09-08 20:29:58.093909 E | ceph-cluster-controller: failed to reconcile CephCluster "rook-ceph/my-cluster". failed to reconcile cluster "my-cluster": failed to configure local ceph cluster: failed to create cluster: failed to start ceph monitors: failed to initialize ceph cluster info: failed to save mons: failed to create/update cephConnection: failed to get ceph connection CR: no matches for kind "CephConnection" in version "csi.ceph.io/v1"
```
Reconciling the mons is blocked until the new CephConnection CRD is available, and that CRD is new as of v1.18.
The simple fix is to simply install those CRDs in create-dev-cluster.sh. After doing this, the dev cluster bootstraps fine.
Signed-off-by: Elias Carter <elias@dropbox.com>
this commit add check to only run the csi-operator in
all the canary tests and upgrade suite only, other suite
like smoke and object will still test csi-driver.
Also, adding changes to make CI happy.
Signed-off-by: subhamkrai <srai@redhat.com>
Co-Authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
For Rook v1.18 the min supported version of K8s is
v1.29. With the pending release of K8s 1.34, this
will be the typical six most recent releases that
Rook tests against.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This PR implements the CSI cephx key rotation feature,
which follows overlapping rotation for non-daemon keys.
Refer to the cephx rotation API from design PR 15915 for details.
Signed-off-by: subhamkrai <srai@redhat.com>
Ceph has a new ceph auth rotate command currently
present in ceph:main
Add a new flag `--cephx-key-rotate` to rotate the
cephx keys genrated by external python script,
If we enable it, it will create a new user with suffix `.{x}`
Signed-off-by: parth-gr <partharora1010@gmail.com>
The create-dev-cluster script used the kvm2 minikube driver on Linux.
On newer Linux flavors like Fedora 41+, the kvm2 driver can have problems
and the qemu2 driver is recommended.
This changes the script to use the qemu2 driver for Linux
Signed-off-by: Michael Adam <obnox@samba.org>
Fixes quick disk shredding by using dd to shred data at additional offsets where ceph metadata is duplicated. For full shred, the shred utility remains in use.
Signed-off-by: Vilius Puškunalis <47086537+puskunalis@users.noreply.github.com>
The script tests/scripts/helm.sh was previously used by some ci
workflows to install helm.
Now that this is not used anymore, this change removes the script.
Signed-off-by: Michael Adam <obnox@samba.org>
tests/scripts/helm.sh clean is not implemented.
So remove frpom the script's help text
and from the development guilde
Signed-off-by: Michael Adam <obnox@samba.org>
The script sel-release-ver.sh was added in a somewhat unnatural place
tests/scripts .
This moves it to a more natural location build/release .
Signed-off-by: Michael Adam <obnox@samba.org>
This script updates examples and docs for a new release.
Invocation in principle: set-release-ver.sh NEW_VER
For example: set-release-ver.sh v1.16.8
Fixes: #15749
Signed-off-by: Michael Adam <obnox@samba.org>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
this PR updates the Prometheus Operator URL references from
version v0.71.1 to the latest release v0.81.0 in documentation
and integration test scripts. This ensures we are aligned with
the latest features and improvements from
the Prometheus Operator project.
Signed-off-by: Oded Viner <oviner@redhat.com>
The helm charts will now only be published when it is
an officially tagged release build.
The images will only be published to all repos for
dockerhub, quay, and ghcr when it is a tagged release.
The images will be published only to dockerhub for all
master and interim release branch builds.
Remove obsolete makefile option for images.
Ceph is the only image Rook ever expects to build.
Simplify the makefile by removing the legacy option
to select which image to build.
Also included are other small improvements to clean up
the release scripts.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
s3cmd can take more time to put the data
of 1M to the bucket as it waits for
connection to get established
currently ci fails with, Retrying failed
request: /test1-1mib-test.dat ([Errno 111]
Connection refused)
also increase the timeout for creating objectstore
Signed-off-by: parth-gr <partharora1010@gmail.com>