485 Commits
Author SHA1 Message Date
Joshua Hoblitt 16cf150a9e Merge pull request #17975 from jhoblitt/docs-ci-gates-and-agents
docs: document the CI checks that gate a pull request, and add AGENTS.md
2026-08-06 12:54:52 -07:00
Travis Nielsen 2bfee41b0b Merge pull request #18063 from malayparida2000/fix/modcheck-error-message
ci: fix mod.check validation error message
2026-07-28 12:30:37 -06:00
Malay Kumar Parida 408dfdebd2 ci: fix mod.check validation error message
Advise running make mod.check and committing results instead of make clean.

Signed-off-by: Malay Kumar Parida <mparida@redhat.com>
2026-07-28 23:35:10 +05:30
Joshua Hoblitt 9277239590 ci: check the links in AGENTS.md
`AGENTS.md` is almost entirely links into `Documentation/`, which is what keeps
it from duplicating the documentation and drifting away from it. That only works
while the links resolve, and nothing checked them: the link checker walks
`Documentation/` alone, so a moved page would leave `AGENTS.md` pointing at
nothing and silently misdirect its reader.

Include it in the files the link checker walks.

The other markdown files in the repository root are left out on purpose. Several
of them already contain dead external links, and adding them here would fail the
nightly job for reasons unrelated to this change.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-27 16:25:27 -07:00
Oded Viner ec502c3f80 nvmeof: use .nvmeof pool for gateway and remove pool field from CRD
Changes:
- Use a dedicated .nvmeof pool (via CephBlockPool CR named
  builtin-nvmeof) for the NVMe-oF gateway internal state.
- Use a separate nvmeof pool for the StorageClass data.
- Remove the pool field from the CephNVMeOFGateway CRD.
  The gateway now always uses the .nvmeof pool, hardcoded
  in the operator.
- Add .nvmeof to the allowed CephBlockPool name overrides.
- Create a production example nvmeof.yaml (instances: 2,
  replicas: 3) and a CI-only nvmeof-test.yaml (instances: 1,
  replicas: 1).
- Update documentation and CI test script accordingly.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-07-21 20:19:44 +03:00
Michael Adam 031830d418 tests: add make lint.markdown-links
This adds a make target lint.markdown-links
for checking forbroken links in the markdown
documentation sources.

Thid is intended to be used in the CI as well as
locally by developers

The link checking is implemented using the markdown-link-check tool
but it is self-contained in that it does not require the
tool to be installed on the system.
The only prerequisite for using this target is to have
a working docker or podman available.

Assisted-by GitHub copilot
Assisted-by: Google Antigravity/Gemini
Assisted-by: IBM Bob
Signed-off-by: Michael Adam <obnox@samba.org>
2026-07-09 17:59:15 +02:00
Joshua Hoblitt 0a7d64bc69 Merge pull request #17877 from jhoblitt/ci/fix-find-block-dev-broken-pipe
ci: fix broken-pipe race in find_extra_block_dev
2026-07-07 08:38:39 -07:00
Joshua Hoblitt d59abb1956 ci: convert integration and canary CI from minikube to kind
The integration and canary suites ran on a single-node minikube
`driver: none` cluster, where the kubelet ran directly on the GitHub
runner, so the host docker daemon doubled as the cluster runtime and
host block devices and host paths were directly visible to pods. Replace
that with a kind cluster.

Every suite creates its cluster through the shared
integration-test-setup-cluster-resources composite action, so the
conversion is centralized there and converts the smoke, object, helm,
keystone, multi-cluster, upgrade, on-release, nightly, encryption-KMS and
all canary jobs at once.

- Replace the setup-minikube step with helm/kind-action, selecting the
  kubernetes version via the kindest/node image tag and creating a
  single-node cluster from a new kind config (kind pinned to v0.32.0 for
  reproducibility).
- Drop the cri-dockerd install; kind nodes use their built-in containerd.
- Add a kind config that bind-mounts the host /dev, /var/lib/rook and
  /run/udev into the node so the existing host-based disk-prep helpers
  (use_local_disk*, create_partitions_for_osds, blockDevicePV.sh,
  localPathPV.sh, ...) keep working unchanged: devices and partitions
  created on the host appear in the node and in the OSD pods that
  hostPath-mount the node /dev, and ceph-volume can read the host udev
  database it needs to inventory disks.
- Prepare the kind node for the host-level operations rook runs against the
  underlying host: remount /sys read-write so CSI's kernel RBD mapping
  (`rbd map --device-type krbd`, which writes /sys/bus/rbd) works, and install
  lvm2 and cryptsetup, which rook runs in the node's mount namespace to
  provision LVM- and encryption-backed OSDs. kindest/node images provide none
  of this; the minikube driver:none runner host did.
- Route the Service and pod CIDRs from the runner to the kind node so
  host-side tests (the `go test` process runs on the runner) can reach
  in-cluster ClusterIPs, e.g. an S3 request to the RGW service. With minikube
  driver:none the runner already shared the cluster network.
- Load locally built images into the cluster. Under minikube `driver: none`
  the built image was already in the cluster runtime; under kind it must be
  imported, so build_rook and create_helm_tag now import their images into
  each node's containerd through a new load_image_into_cluster helper (via the
  node's ctr, which avoids the kind/kindest-node containerd-config version
  skew that breaks `kind load docker-image`).
- Point Vault's kubernetes-auth at the in-cluster API endpoint
  (kubernetes.default.svc) instead of the kubeconfig server URL: kind exposes
  that as https://127.0.0.1:<port>, unreachable from the in-cluster Vault pod,
  so OSD encryption-key retrieval via k8s-auth failed.
- Replace the remaining direct minikube references in the canary workflow: a
  `minikube kubectl` call and the external-cluster topology values.
- Adapt host-name assumptions that only held under driver:none: resolve the
  disk-cleanup job by the k8s node name rather than the runner hostname, and
  let kind-action ignore post-job cluster-teardown failures (the runner is
  ephemeral; nvme/multus devices can wedge `docker rm` of the node).
- Update stale comments that described the CI environment as minikube.
- Move the multus integration test's kind config under tests/config too, so
  both kind cluster configs live in one place.

create-dev-cluster.sh and other local-dev tooling are intentionally left on
minikube.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-06 08:42:39 -07:00
Joshua Hoblitt 7079c9bf62 ci: fix broken-pipe race in find_extra_block_dev
find_extra_block_dev piped find_extra_block_devs into head -1. Since
find_extra_block_devs prints every eligible disk and ends with echo
"$devs", head closes the pipe after the first line and the producer's
echo races into SIGPIPE (echo: write error: Broken pipe), failing the
command substitution under pipefail/errexit.

The race only fires reliably now that the runners present two data disks
(sdb and sdc), so find_extra_block_devs emits multiple lines; with a
single disk the one-line output almost never triggered it.

Read the producer to completion with mapfile via process substitution
and return the first element, mirroring find_second_block_dev, so there
is no downstream head to close the pipe early.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 09:17:20 -07:00
Joshua Hoblitt c0eb360041 docs: fix typos and grammar in code comments
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:34:11 -07:00
Joshua Hoblitt 029b13b000 ci: provision extra disks when the runner starts with none
find_extra_block_devs reads the runner's current extra disks with
`lsblk ... | egrep -v "($boot_dev|loop|nbd)"` to decide how many more to
provision. When the runner has no extra disk, that egrep matches nothing
and exits 1; under the script's `set -eo pipefail` the assignment aborts
the function before it ever calls create_extra_disk, so callers get zero
disks and fail with "expected >= N extra disks, found 0" (or a bare /dev/
path) on those runners.

Guard that first read with `|| true` (as the adjacent device-count line
already is) so an empty result means "no extra disks yet" and the
provisioning path runs. The second read, after create_extra_disk, is left
strict on purpose: if provisioning produced no disk, egrep's non-zero exit
should abort there rather than silently return nothing.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-26 15:00:37 -07:00
Joshua Hoblitt 04734d7150 ci: capture build_rook stderr so the network-retry can see it
build_rook() captures the make output into $o and inspects it for
transient network errors (connection reset, INTERNAL_ERROR, 503, 500)
to decide whether to retry the build. The capture was stdout-only,
but go and make write those errors to stderr, so $o never matched any
retry case. A transient module-proxy failure (e.g. proxy.golang.org
returning INTERNAL_ERROR while downloading modules) therefore fell
through to the "valid failure" branch and failed the job in the build
step, before the integration test even ran.

Redirect stderr into the captured output (2>&1) so the existing retry
logic actually triggers on network failures.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-25 12:36:33 -07:00
Joshua Hoblitt c1594a82ea Merge pull request #17801 from jhoblitt/ci-loop-to-iscsi
ci: replace loop devices with iscsi in canary osd tests
2026-06-24 22:25:16 -07:00
Joshua Hoblitt 343462981a Merge pull request #17798 from jhoblitt/ci-multisite-sync-converge
ci: converge wedged multisite secondary before asserting sync
2026-06-24 08:56:49 -07:00
Joshua Hoblitt 4b7d769102 ci: replace loop devices with iscsi in canary osd tests
The raw-disk-with-object, osd-with-metadata-device, and pvc canary jobs
used loopback devices to supply an extra block device beyond the single
spare disk a GitHub runner provides. Loop devices behave differently
from real block devices -- Rook gates them behind
ROOK_CEPH_ALLOW_LOOP_DEVICES and special-cases them in the OSD daemon --
so exercising core scenarios on them is less realistic than using real
SCSI disks.

Provision the extra device(s) over iSCSI (LIO) instead, reusing the
existing create_extra_disk helper: it now creates N LUNs, the new
find_extra_block_devs and find_second_block_dev helpers enumerate them,
and deploy_cluster gains a two_raw_disks mode. The three jobs now run
entirely on real disks, and the canary CI no longer creates loop devices
or relies on the loop-device test helpers.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-24 08:54:54 -07:00
Joshua Hoblitt 335e840079 test: remove orphaned legacy multi-node and kubeadm scripts
These shell scripts under tests/scripts are referenced nowhere in the
repository -- no workflow, Makefile, Go test, documentation, or
.mergify.yml invokes them:

- tests/scripts/multi-node/ (build-rook.sh, rpm-system-prerequisites.sh,
  config.rb): the abandoned Vagrant multi-node VM harness, pinned to the
  long-EOL centos/7 box.
- tests/scripts/kubeadm-install.sh: an old kubeadm/apt installer for a
  pinned Kubernetes version, no longer wired into any test path.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-23 15:37:11 -07:00
Joshua Hoblitt 6f1e7a7dd5 test: remove unused functions from tests/scripts helpers
Three shell functions defined under tests/scripts are never invoked --
not internally, not from another script, and not dispatched by name from
any CI workflow:

- deploy_csi_hostnetwork_disabled_cluster (github-action-helper.sh)
- wait_for_ceph_csi_configmap_to_be_updated (github-action-helper.sh)
- test_demo_pool (validate_cluster.sh): the daemon validation loop never
  passes a "pool" daemon, so this branch is unreachable.

Removing them orphans no other helpers; each only called functions that
remain in use elsewhere.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-23 15:36:51 -07:00
Joshua Hoblitt e9b59476a6 ci: extend the multisite replication read budget
The bucket sync checkpoint confirms the bucket's data sync markers are
caught up at the RADOS level, but the reading RGW can briefly serve 404
for the freshly synced object until it refreshes its view. The post-
checkpoint read was given only a 60s budget, which is occasionally too
short; give it 180s so a transient RGW serving lag does not fail the
test.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-22 19:22:24 -07:00
Joshua Hoblitt 325d314dd7 ci: converge wedged multisite secondary before asserting sync
After the sync-readiness gate was added, rgw-multisite-testing still
fails ~25% of master-push runs, now in the new "wait for multisite
sync to be established" step. Diagnosis from the failure diagnostics:
the secondary zone never initializes metadata sync. For the whole
600s its sync status reports "metadata sync failed to read sync
status: (2) No such file or directory".

The master pushes each new RGW configuration period to the secondary
RGW, but until the realm system user is resolvable locally on the
secondary that push is received unauthenticated and rejected with
HTTP 403. The master then re-pushes the same period unchanged every
30s, so a secondary that starts out unable to resolve the user never
recovers. In passing runs the same push later authenticates as
realm-a-system-user and returns 200, at which point the zone converges;
in failing runs it is 403 for the entire window (zero successful
pushes). The generated keys are clean and identical on both clusters,
so this is a fresh-zone bootstrap race in the gateway, not a key
problem.

Rather than wait out a doomed timeout, nudge the secondary to converge
when its metadata sync has not initialized within a couple of minutes:
pull the latest period directly (outbound auth works, so this does not
depend on the rejected inbound push) and restart the secondary RGW so
it re-runs metadata sync init against the now-ready master. This is
what ceph's own multisite QA does (restart zone gateways after the
period is committed). Healthy setups pass the first check in seconds,
so the nudge only runs when sync is genuinely stuck.

Also guard wait_for_sync_status against matching empty output from a
failed exec, and pass the realm/zonegroup/zone to the diagnostics
sync-status dump so the primary cluster no longer reports the unrelated
default zone.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-22 12:45:04 -07:00
Travis Nielsen dc50bce4e6 Merge pull request #17559 from OdedViner/network_policy_upstream
security: add example NetworkPolicy CRs for Ceph operand pods
2026-06-17 12:54:24 -06:00
Travis Nielsen 393cb7a25a Merge pull request #17711 from jhoblitt/ci-rgw-multisite-deterministic
ci: make rgw-multisite-testing wait for multisite readiness
2026-06-16 15:15:33 -06:00
Travis Nielsen 72dde3db1d Merge pull request #17712 from jhoblitt/ci-multus-setup-flakes
ci: fix multus integration test setup flakes
2026-06-16 15:14:45 -06:00
Travis Nielsen 923302d7e7 Merge pull request #17714 from jhoblitt/ci-upgrade-suite-flakes
ci: reduce CephUpgradeSuite flakes
2026-06-16 15:14:03 -06:00
Oded Viner 9dc70051a8 security: add example NetworkPolicy CRs for Ceph operand pods
Add a single example YAML with NetworkPolicy definitions
for all Ceph operand pods (mon, osd, mgr, mds, exporter,
osd-prepare, crashcollector, tools).
Users can apply these policies to restrict egress/ingress
traffic for Rook-Ceph daemon pods.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-06-16 17:44:34 +03:00
Joshua Hoblitt cff8b7f311 ci: reduce CephUpgradeSuite flakes
The CephUpgradeSuite workflow failed ~31% of PR runs vs the ~19%
broken-PR baseline of the other integration suites. Classifying the
failing step of every failed run from the last ~400 PR runs and the
logs of every failure on non-dependabot branches shows the excess
comes from a handful of too-tight waits, unretried network fetches,
and environment races rather than from the upgrade logic itself:

- The biggest class (21 of 44 genuine test-phase failure runs):
  "giving up waiting for deployment(s) with label
  app=rook-ceph-{rgw,osd},ceph-version!=..." during the
  Squid->Tentacle upgrade. The operator updates daemons sequentially
  (mons, mgr, osds with PG health gates between them, mds, rgw), so
  each later daemon's fixed 275s wait also absorbs the time spent on
  every daemon before it; rgw, last in the sequence, failed most
  often. The mon wait was already extended for slow image pulls in
  98bd03a42; extend the same retry count to all the daemon waits.

- ~7 runs: "snapshot controller is not ready" in the Helm upgrade
  path. WaitForSnapshotController(30) allows only 150s, and the logs
  show the deployment still converging (readyreplicas 1 < replicas 2)
  on the final poll. Raise to 90 retries.

- InstallOrUpgradeHelmRepoChart ran helm with no retry; an observed
  failure fetched the ceph-csi-drivers chart tarball from GitHub
  release assets and got a 504. Retry up to 5 times, as
  InstallLocalHelmChart already does.

- Raise the go test timeouts (2400s->3600s rook, 1800s->2700s helm).
  Failed runs ended in "panic: test timed out" during teardown, which
  aborts cleanup and log collection; the longer daemon waits above
  also need the headroom.

- The "setup cluster resources" composite action (shared by all
  suites) accounted for half of all failed jobs. The one steady class
  there (8 distinct runs across 7 different days): minikube exits with
  K8S_FAIL_CONNECT (code 40) when the requested kubernetes version is
  missing from its bundled version list, because it then validates the
  version with an anonymous GitHub API request (GITHUB_TOKEN is not
  honored on that code path), and anonymous requests from shared
  runner IPs are regularly rate-limited. minikube 1.38.x predates
  v1.35.5, so only the v1.35.5 jobs hit this class. Pass --force to
  minikube start to skip that check: the kubernetes versions used in
  CI are pinned constants already validated by the PRs that bump
  them, and an invalid version would still fail fast at the kubeadm
  download. Also seen twice: dpkg failing on a corrupt cri-dockerd .deb
  because curl ran without --fail and saved an HTTP error page as the
  package; add --fail and retries.

- On runners without the /mnt resource disk, the fallback OSD disk is
  an iSCSI (LIO) device and use_local_disk_for_integration_test
  returned early on those runners, skipping the udev nowatch
  workaround for the device re-probe storms of rook issue 8975. A burn-in
  failure on this PR captured the consequence with full kernel
  forensics: ~100 udev change events on the OSD disk, and ceph-volume
  activate wedged in uninterruptible sleep on the block device lock
  (blkdev_llseek -> rwsem_down_write_slowpath) for 18+ minutes during
  the ceph version upgrade while the cluster's only OSD stayed down.
  Install the nowatch rule before the early return so it applies to
  every runner type, and before the disk is first written rather than
  after.

- One burn-in round failed before the upgrade even began: the
  pre-upgrade PVC create on the Helm path expired WaitUntilPVCIsBound
  (RetryLoop, 275s) while the CSI provisioner pods were still
  starting. Set RETRY_MAX=110 for this workflow to double the
  framework's base wait budget, instead of extending individual waits
  one flake at a time.

- minikube also exits with GUEST_START (code 80) when its internal 6
  minute node-ready wait expires on slow runners (5 runs in the
  dataset, two more observed while burning in this PR, one of them on
  another suite). Pass --wait-timeout=15m to minikube start.

Not addressed (small or episodic classes): OSDs never coming up on
initial deploy (3 runs, possibly the same udev/iSCSI wedge), the
filesystem not becoming active (2 runs), and several day-clustered
minikube incidents (exit codes 65/67/90).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 15:27:02 -07:00
Joshua Hoblitt 98f088b54f ci: fix toolbox exec race in validate_cluster.sh
validate_cluster.sh built its ceph exec command around a toolbox pod
name captured once at script start, without waiting for the toolbox to
be ready. When the toolbox container was still being created, or the
pod had been replaced, every check failed with 'unable to upgrade
connection: container not found' for the entire wait window and the
mon quorum check timed out without ever observing the cluster. This
intermittently failed the 'wait for ceph cluster N to be ready' steps
of canary jobs (observed on master pushes of multi-cluster-mirroring
and the canary job between 2026-03-10 and 2026-04-24, always with the
container-not-found signature filling the whole window).

Wait for the toolbox deployment rollout up front, and exec through
deploy/rook-ceph-tools so each call resolves a currently-ready pod,
matching how the other CI scripts invoke the toolbox.

Also call wait_for_daemon directly instead of 'return $(...)'. The
command substitution ran wait_for_daemon in a subshell and captured
its output, so the 'current status' diagnostics printed on timeout
were never displayed; worse, the captured text became the argument of
'return', which fails with 'numeric argument required' and aborts the
script with a misleading exit code 2. The osd variant's captured
'Return value' debug echo had the same problem and is removed.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 14:19:52 -07:00
Joshua Hoblitt 4da654bae2 ci: fix multus integration test setup flakes
Multus CI failures on master pushes from 2026-04-08 to 2026-06-10 were
dominated by a race in setup-multus.sh: 'kubectl wait' was invoked on a
pod label selector immediately after 'kubectl create' of a daemonset,
and it exits with "error: no matching resources found" when the
daemonset controller has not created any pods yet. This caused 9 of the
10 genuine multus flakes across the standalone multus workflow and the
canary multus-public-and-cluster job (the two share setup-multus.sh).

- setup-multus.sh: wait with 'kubectl rollout status' on the daemonsets
  instead of 'kubectl wait' on pod label selectors. The daemonset
  object exists as soon as 'kubectl create' returns, and rollout status
  correctly waits for all desired pods to be created and become ready.
  The old selector wait also silently under-waited: pods created after
  its initial LIST were never waited on at all. The stricter wait
  revealed that full multi-node convergence can exceed 2 minutes on
  busy runners, so the wait timeout is raised to 5 minutes (the waits
  return as soon as the rollouts are ready, so this costs nothing on
  healthy runs).

- Wait for the host-net-config daemonset rollout in both workflows
  before proceeding; it configures the host routing that the multus
  public network depends on and was previously not waited on at all.
  Pin its jonlabelle/network-tools image.

- test_multus_connections: retry the osd dump / fs dump network checks
  for up to 2 minutes. Daemons register their addresses in the mon maps
  asynchronously after the cluster reports ready; one canary failure
  (2026-04-29) ran the MDS check at fsmap epoch 1, before any MDS had
  registered.

- Dump cluster state (pods, daemonsets, events, multus logs) when the
  multus workflow job fails; the workflow previously had no failure
  diagnostics at all.

Also remove the unused NUMBER_OF_COMPUTE_NODES env var from the multus
workflow (kind-config.yaml creates 3 workers; nothing consumes it).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-10 21:12:46 -07:00
Joshua Hoblitt bf00821056 ci: make rgw-multisite-testing wait for multisite readiness
The rgw-multisite-testing canary job has been chronically flaky
(roughly a third of recent canary runs, including on master and
release branches). Job logs and collected artifacts show two recurring
mechanisms, neither of which is addressed by bumping timeouts:

1. Every recent failure of the "write an object" step died within
   ~16s: the very first "s3cmd mb" against the primary zone got
   ECONNREFUSED three times across the 10s retry window. The RGW
   process never crashed; RGWs pause their HTTP frontends to reload
   the realm whenever the RGW configuration period changes ("rgw
   realm reloader: Pausing frontends for realm update..."), which
   keeps happening shortly after the second zone joins the zonegroup,
   exactly when the test starts writing. While paused, the readiness
   probe also times out, the endpoint is removed from the Service, and
   connections to the ClusterIP are refused.

2. The "kubectl wait --for=condition=available deployment/..."
   barrier added previously is a no-op for these deployments: with one
   replica and maxUnavailable=1, the availability threshold is zero
   ready pods, so the condition is always true.

Changes:

- Wait on RGW pod readiness instead of deployment availability.
- Add a "wait for multisite sync to be established" step that polls
  "radosgw-admin sync status" on both clusters until metadata and
  data sync report caught up, mirroring the checkpoints ceph's own
  multisite QA performs before asserting replication.
- Probe both RGW endpoints over HTTP until several consecutive
  requests succeed before writing, to ride out realm-reload pauses.
- Replace the fixed 3-attempt s3cmd retries with time-budgeted
  retries.
- Replace the s3cmd read polling loop and its SIGUSR1 timer (which was
  never cancelled on success, silently shrinking the second
  direction's time budget) with "radosgw-admin bucket sync
  checkpoint", the marker-based barrier used by ceph QA, followed by
  a single get and diff.
- Dump multisite diagnostics (pod state, RGW logs, sync status,
  period) when the job fails.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-10 15:42:51 -07:00
subhamkrai 34b0a87c9e csi: csi-addons should be disabled by default
having csi-addons enabled by default is causing random
pod restart on non-openshift cluster. Let's disable
it by default.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-06-05 13:42:17 -06:00
Travis Nielsen 1a37f6088c ci: canary tests only start one csi controller
Two csi controllers cannot start on minikube, so the CI
was failing in the canary and nvmeof tests when they validated
that the controllers were running. For the tests that validate
the csi controllers, ensure only a single controller is configured.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2026-06-04 11:07:40 -06:00
Santosh 2cc51e8eb0 ci: retry s3cmd bucket creation in rgw-mutisite test
The rgw-multisite job is failing frequently. Observed
"connectionRefusedERror: connection Refused" on the s3cmd mb call. It
might be happening because CephObjectSTore CR is ready but RGW gateway
is not fully accepting TCP connections.

This PR  wraps s3cmd mb and s3cmd put calls in a retry loop that
will attempt each command up to 3 times with a 5 second sleep between
attemps. This might help fix the flaky CI test.

Signed-off-by: Santosh <sapillai@redhat.com>
2026-05-25 20:16:16 +05:30
subhamkrai 4eefad42e8 csi: move csi management to admin
Going forward, admin will manage the csi operator
CR's and rook will only manage Ceph Connection cr
and client Profile cr.

The old csi driver is completely removed from Rook
and can no longer be used starting in Rook v1.20.

The upgrade guide will contain the needed transition steps
for managing the csi operator settings.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-04-29 14:24:26 -06:00
subhamkrai b50f1a3d60 ci: fix some minor canary test issues
The canary test was waiting for replicapool
instead of replicapool2, and a more reliable
wait for the toolbox pod start is added.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-04-29 14:01:42 -06:00
Oded Viner 86c920fa79 object: add support for aws sdk v2 in s3agent createbucket
initialize both aws sdk v1 and v2 clients in s3agent.
update createbucket to use sdk v2 while keeping all
other methods on sdk v1.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-03-31 17:41:51 +03:00
Oded Viner ebfbb49f15 nvmeof: add csi operator support for nvmeof canary test
update the nvmeof minikube canary test to use the
ceph-csi operator instead of manually deploying
the provisioner and node-plugin.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-03-24 12:18:45 +01:00
Travis Nielsen 24053d04d3 Merge pull request #17017 from parth-gr/test-mirror-ec
ci: add ec pool mirroring in our ci
2026-03-03 10:24:26 -07:00
parth-gr 50f5eec61d ci: add ec pool mirroring in our ci
we do support ec pool mirroring too
Add a test for that

Signed-off-by: parth-gr <partharora1010@gmail.com>
2026-03-03 11:35:05 +05:30
Oded Viner 2ebb5d968a nvmeof: add nvmeof minikube canary test without csi operator
adds a new nvmeof minikube canary job for ceph v20.
the test deploys rook with csi operator disabled for nvmeof flow.
it validates pvc and pod io, gateway restart, and data persistence.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-02-18 13:22:50 +02:00
Satoru Takeuchi 0137274b58 test: create iscsi disk if extra disk does not exist
We found a new github runner that doesn't have an extra
disk mounted on /mnt and it causes massive amounts
of CI failures. We can overcome this circumstance
by creating an iSCSI disk as an extra disk.

ref. https://github.com/rook/rook/issues/16978#issue-3862612835

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2026-02-02 16:29:44 +00:00
Blaine Gardner 526acaf868 ci: add initial kube-api-linter support to CI
Kubernetes has begun releasing a kube-api-linter application for linting
Kubernetes APIs to match against common best-practices. (a.k.a., KAL).
Of note: KAL is still new and has no semver releases.

I believe this linter is helpful in encouraging Rook devs to think
deeply about how end-users consume the API. It ensures that new APIs
account for set/unset status, zero values, upper/lower bounds, and
possibly others in the future.

This commit adds initial support to Rook, without making the results a
requirement. Maintainers still need to evaluate rule configuration and
determine how tightly Rook wants to hold to the results.

As of this commit, I observe around 650  "issues" with the current Rook APIs.
This doesn't mean that Rook is in danger of breaking. Issues with API
types are often just indications that upper/lower bounds are missing, or
that there is ambiguity between set-empty and unset-empty values.

This commit uses `--new` in CI to ensure only new API issues are
flagged.

Also of note, KAL is a golangci-lint plugin. It is possible to locally
build KAL (from source) into golangci-lint to run both at the same time.
However, this is time-consuming, and it's harder to coordinate tool
binary caching this way when it's related to 2 different versions.
Treating KAL as a standalone binary allows for simpler make scripting
and maintenance.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2026-01-21 11:34:22 -07:00
Praveen M 8ebf07bbfe csi: update csi-addons to v0.14.0
Signed-off-by: Praveen M <m.praveen@ibm.com>
2026-01-16 11:50:18 +05:30
subhamkrai 6c4d8205b6 ci: use minikube github action instead script
using github action for minikube will avoid
unwanted errors and probably will be more stable
than manual install. But both script and action
both uses almost same time for installation so
there we don't have preference.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-01-13 15:24:01 +05:30
subhamkrai a62a1392f8 ci: run daily nightly job in with gh arm runner
let's run daily nightly canary job with github action
arm runner as self-hosted runners is shutting down which
was provided by upstream user.

Signed-off-by: subhamkrai <srai@redhat.com>
2025-12-08 21:20:37 +05:30
Erik Sundell e7a704ccf1 helm: remove legacy PodSecurityPolicy resource
The helm charts allowed rendering a PodSecurityPolicy resource via the
configuration `pspEnable`. This option is removed and all references to
psp, PodSecurityPolicy, and Pod Security Policy have been cleaned up.

The PSP resource was only rendered if k8s version was lower than 1.25
when it was still supported. It has been deprecated since k8s 1.21.

Signed-off-by: Erik Sundell <erik@sundellopensource.se>
2025-09-30 18:25:35 +02:00
Michael Adam 605e820fc8 ci: move minikube to a better location
In the deb package's location of /usr/bin/minilube, a wrong version seems to be
reported but from /usr/local/bin it reports correctly.

Signed-off-by: Michael Adam <obnox@samba.org>
2025-09-16 18:24:18 +02:00
Michael AdamandTravis Nielsen 40427d5c0d ci: update latest k8s version to 1.34
this change updates the k8s version to 1.34 and also updates
the  cri-ctl version for minikube.

Signed-off-by: Michael Adam <obnox@samba.org>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2025-09-16 17:57:28 +02:00
subham rai 90804fd61a Merge pull request #16449 from parth-gr/fix-vault
ci: fix the ci canary vault tests
2025-09-10 14:32:17 +05:30
parth-gr b4fa906202 ci: fix the ci canary valut test
use the official way for helm installation
and not use the cdn installation

Signed-off-by: parth-gr <partharora1010@gmail.com>
2025-09-10 12:16:11 +05:30
Elias Carter 3ed0d928b8 test: fix create-dev-cluster.sh hanging since v1.18
Currently create-dev-cluster.sh will hang as the rook operator gets stuck like so:
```
2025-09-08 20:29:58.093909 E | ceph-cluster-controller: failed to reconcile CephCluster "rook-ceph/my-cluster". failed to reconcile cluster "my-cluster": failed to configure local ceph cluster: failed to create cluster: failed to start ceph monitors: failed to initialize ceph cluster info: failed to save mons: failed to create/update cephConnection: failed to get ceph connection CR: no matches for kind "CephConnection" in version "csi.ceph.io/v1"
```

Reconciling the mons is blocked until the new CephConnection CRD is available, and that CRD is new as of v1.18.

The simple fix is to simply install those CRDs in create-dev-cluster.sh. After doing this, the dev cluster bootstraps fine.

Signed-off-by: Elias Carter <elias@dropbox.com>
2025-09-08 13:36:22 -07:00
subhamkraiandTravis Nielsen c57a47f774 ci: run csi-operator only in canary and upgrade suite
this commit add check to only run the csi-operator in
all the canary tests and upgrade suite only, other suite
like smoke and object will still test csi-driver.

Also, adding changes to make CI happy.

Signed-off-by: subhamkrai <srai@redhat.com>
Co-Authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2025-08-19 11:23:14 +05:30