Commit Graph
817 Commits
Author SHA1 Message Date
Travis Nielsen c73f8ab1c2 ci: daily tests run stable and devel ceph versions
The daily CI smoke, object, and other suites now run
all of the known Ceph versions, including stable
Squid, Tentacle, and Umbrella, as well as devel
versions of the same releases.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2026-07-27 11:42:46 -06:00
Chiman Jain 2034be1f37 build: migrate gopkg.in/yaml to go.yaml.in/yaml
Signed-off-by: Chiman Jain <chimanjain15@gmail.com>
2026-07-22 20:05:39 +05:30
Joshua Hoblitt 452fb31d95 Merge pull request #17917 from jhoblitt/chart-object-multisite
helm: add multisite CR support to rook-ceph-cluster chart
2026-07-13 13:10:05 -07:00
subhamkrai 458a8a34f8 ci: add umbrella-devel smoke suite and canary coverage
Add nightly smoke tests and installer support for the umbrella-devel Ceph image.
Also, refactored the test to use reusable code, and since we run the tests in
k8s, removed the k8s matrix from the input.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-07-13 16:28:32 +05:30
Joshua Hoblitt 8a72d80127 test: exercise multisite chart values in CephHelmSuite
Configure the CephHelmSuite object store as an RGW multisite zone so the
chart's CephObjectRealm, CephObjectZoneGroup, and CephObjectZone templates
are exercised on a Helm install. The operator must reconcile the realm,
zone group, zone, and store to Ready, which the existing install check
already gates on via the RGW pod count.

The realm/zone group/zone and shared-pool layout mirror the fixture in
tests/integration/object/util/sharedstore: a zone backed by explicit
CephBlockPools referenced through sharedPools.poolPlacements. The rgw pools
are appended to the block pools already configured for CSI, and are torn
down along with the multisite CRs when the default storage CRs are removed.

Gate the multisite configuration behind a new UseMultisiteObjectStore
setting enabled only for CephHelmSuite. The other Helm-based suites keep the
plain object store; the upgrade suite in particular installs an older base
chart that lacks the multisite templates and cannot resolve a zone-based
store.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-09 09:59:18 -07:00
Joshua Hoblitt 6c2a3a5f2b Merge pull request #17881 from jhoblitt/convert-bucket-quota-to-sharedstore
test: convert the object bucket quota suite to the shared store
2026-07-07 15:35:29 -07:00
subhamkrai a3e02e607d csi: update csi-operator version to v1.0.4
Updating csi-operator to latest v1.0.4 and
updating the required doc changes as well.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-07-07 07:56:55 -06:00
Joshua Hoblitt dad0e21122 test: retire the unused object bucket framework helpers
Converting the bucket quota tests to typed clients left the framework
helpers and package globals they were the last consumers of unused. Remove
UpdateObc, CheckOBMaxObject, GetAccessKey, and GetSecretKey from the bucket
client, GetEndPointUrl from the object client, and the orphaned object-test
globals.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-06 10:35:37 -07:00
Joshua Hoblitt d59abb1956 ci: convert integration and canary CI from minikube to kind
The integration and canary suites ran on a single-node minikube
`driver: none` cluster, where the kubelet ran directly on the GitHub
runner, so the host docker daemon doubled as the cluster runtime and
host block devices and host paths were directly visible to pods. Replace
that with a kind cluster.

Every suite creates its cluster through the shared
integration-test-setup-cluster-resources composite action, so the
conversion is centralized there and converts the smoke, object, helm,
keystone, multi-cluster, upgrade, on-release, nightly, encryption-KMS and
all canary jobs at once.

- Replace the setup-minikube step with helm/kind-action, selecting the
  kubernetes version via the kindest/node image tag and creating a
  single-node cluster from a new kind config (kind pinned to v0.32.0 for
  reproducibility).
- Drop the cri-dockerd install; kind nodes use their built-in containerd.
- Add a kind config that bind-mounts the host /dev, /var/lib/rook and
  /run/udev into the node so the existing host-based disk-prep helpers
  (use_local_disk*, create_partitions_for_osds, blockDevicePV.sh,
  localPathPV.sh, ...) keep working unchanged: devices and partitions
  created on the host appear in the node and in the OSD pods that
  hostPath-mount the node /dev, and ceph-volume can read the host udev
  database it needs to inventory disks.
- Prepare the kind node for the host-level operations rook runs against the
  underlying host: remount /sys read-write so CSI's kernel RBD mapping
  (`rbd map --device-type krbd`, which writes /sys/bus/rbd) works, and install
  lvm2 and cryptsetup, which rook runs in the node's mount namespace to
  provision LVM- and encryption-backed OSDs. kindest/node images provide none
  of this; the minikube driver:none runner host did.
- Route the Service and pod CIDRs from the runner to the kind node so
  host-side tests (the `go test` process runs on the runner) can reach
  in-cluster ClusterIPs, e.g. an S3 request to the RGW service. With minikube
  driver:none the runner already shared the cluster network.
- Load locally built images into the cluster. Under minikube `driver: none`
  the built image was already in the cluster runtime; under kind it must be
  imported, so build_rook and create_helm_tag now import their images into
  each node's containerd through a new load_image_into_cluster helper (via the
  node's ctr, which avoids the kind/kindest-node containerd-config version
  skew that breaks `kind load docker-image`).
- Point Vault's kubernetes-auth at the in-cluster API endpoint
  (kubernetes.default.svc) instead of the kubeconfig server URL: kind exposes
  that as https://127.0.0.1:<port>, unreachable from the in-cluster Vault pod,
  so OSD encryption-key retrieval via k8s-auth failed.
- Replace the remaining direct minikube references in the canary workflow: a
  `minikube kubectl` call and the external-cluster topology values.
- Adapt host-name assumptions that only held under driver:none: resolve the
  disk-cleanup job by the k8s node name rather than the runner hostname, and
  let kind-action ignore post-job cluster-teardown failures (the runner is
  ephemeral; nvme/multus devices can wedge `docker rm` of the node).
- Update stale comments that described the CI environment as minikube.
- Move the multus integration test's kind config under tests/config too, so
  both kind cluster configs live in one place.

create-dev-cluster.sh and other local-dev tooling are intentionally left on
minikube.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-06 08:42:39 -07:00
Joshua Hoblitt c4ac54cea0 Merge pull request #17853 from jhoblitt/convert-bucket-notification-test
test: convert the bucket notification suite to the shared store
2026-07-01 09:22:44 -07:00
Joshua Hoblitt c0eb360041 docs: fix typos and grammar in code comments
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:34:11 -07:00
Joshua Hoblitt 031e449070 test: retire the unused bucket notification framework helpers
Converting the bucket notification test to typed clients left the
framework notification plumbing it was the last consumer of unused.
Remove the NotificationOperation client, TopicClient.CreateHTTPServer,
BucketClient.CreateObcNotification and CheckBucketNotificationSetonRGW
(and the now-dead UpdateObcNotification helpers), and the
GetBucketNotification and GetOBCNotification manifest builders.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:30:22 -07:00
Joshua Hoblitt e281600bb5 Merge pull request #17855 from jhoblitt/narrow-bare-nolint
core: narrow bare //nolint directives to specific linters
2026-06-30 08:14:46 -07:00
Joshua Hoblitt e5c6d75a73 core: narrow bare //nolint directives to specific linters
The //nolint directives at these sites omit the linter name, so each
suppresses every linter on its line rather than the one check it needs.
That hides any unrelated errcheck/gosec/govet finding later introduced
on the same line. Name the specific linter for each:

- staticcheck for the two operator sites: SA4004 (the intentional
  single-iteration loop in the OSD PVC host lookup) and SA1019 (the
  deliberate read of the deprecated S3.Enabled field in the RGW
  API-enable builder).
- errcheck for the rbd-mirror deferred token-file cleanup and the
  test-framework logging helpers (WriteString / writeHeader).

No behavior change.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 16:32:59 -07:00
Joshua Hoblitt 3a4dc64f5b test: convert the COSI driver suite to the object test toolkit
Rework the COSI driver integration test onto the shared-store object test
toolkit: an ordered t.Run script with the (t, k8sh, store) signature, typed
clients, and watch-based waits, replacing kubectl-string manifests and
fixed-interval polling.

Drive the COSI bucket resources (BucketClass, BucketClaim, Bucket) with the
upstream sigs.k8s.io/container-object-storage-interface typed client, exposed
as k8sh.COSIClientset alongside the existing OBC client. Create the
CephCOSIDriver and its privileged user through typed clients, wait for the
driver Deployment with wait4, and verify the provisioned bucket through the
shared store's rgw admin client.

Install the COSI CRDs and central controller from the consolidated upstream
repo pinned to v0.2.2 (the former -api and -controller repos are retired) via
a kubectl -k fixture that is removed with t.Cleanup. The driver cannot trust
a TLS RGW endpoint, so the suite skips itself in the TLS pass rather than
being special-cased in the dispatcher.

Retire the now-unused COSIOperation client and the GetCOSIDriver,
GetBucketClass, and GetBucketClaim manifest helpers.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 15:36:13 -07:00
Joshua Hoblitt 49612461a4 docs: fix function comments to match their declaration names
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-26 12:52:39 -07:00
daixihegu 063dcad054 core: use slices.Contains to simplify code
Signed-off-by: daixihegu <daixihegu@163.com>
2026-06-25 10:17:53 +08:00
Joshua Hoblitt d0885370fa Merge pull request #17732 from jhoblitt/ci-smoke-suite-flakes
ci: harden CephSmokeSuite against CSI cold-start and teardown flakes
2026-06-24 08:57:25 -07:00
Joshua Hoblitt dec475914e Merge pull request #17744 from jhoblitt/keystone-webhook-eventually
test: add Eventually polling helper for test setup waits
2026-06-23 10:39:46 -07:00
dependabot[bot] 516b46bf34 build(deps): bump the github-dependencies group with 4 updates
Bumps the github-dependencies group with 4 updates: [github.com/aws/aws-sdk-go-v2/service/s3](https://github.com/aws/aws-sdk-go-v2), [github.com/ceph/go-ceph](https://github.com/ceph/go-ceph), [github.com/prometheus-operator/prometheus-operator/pkg/apis/monitoring](https://github.com/prometheus-operator/prometheus-operator) and [github.com/prometheus-operator/prometheus-operator/pkg/client](https://github.com/prometheus-operator/prometheus-operator).

Updates `github.com/aws/aws-sdk-go-v2/service/s3` from 1.103.3 to 1.104.0
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/service/s3/v1.103.3...service/s3/v1.104.0)

Updates `github.com/ceph/go-ceph` from 0.39.0 to 0.40.0
- [Release notes](https://github.com/ceph/go-ceph/releases)
- [Changelog](https://github.com/ceph/go-ceph/blob/master/docs/release-process.md)
- [Commits](https://github.com/ceph/go-ceph/compare/v0.39.0...v0.40.0)

Updates `github.com/prometheus-operator/prometheus-operator/pkg/apis/monitoring` from 0.91.0 to 0.92.0
- [Release notes](https://github.com/prometheus-operator/prometheus-operator/releases)
- [Changelog](https://github.com/prometheus-operator/prometheus-operator/blob/main/CHANGELOG.md)
- [Commits](https://github.com/prometheus-operator/prometheus-operator/compare/v0.91.0...v0.92.0)

Updates `github.com/prometheus-operator/prometheus-operator/pkg/client` from 0.91.0 to 0.92.0
- [Release notes](https://github.com/prometheus-operator/prometheus-operator/releases)
- [Changelog](https://github.com/prometheus-operator/prometheus-operator/blob/main/CHANGELOG.md)
- [Commits](https://github.com/prometheus-operator/prometheus-operator/compare/v0.91.0...v0.92.0)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/service/s3
  dependency-version: 1.104.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-dependencies
- dependency-name: github.com/ceph/go-ceph
  dependency-version: 0.40.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-dependencies
- dependency-name: github.com/prometheus-operator/prometheus-operator/pkg/apis/monitoring
  dependency-version: 0.92.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-dependencies
- dependency-name: github.com/prometheus-operator/prometheus-operator/pkg/client
  dependency-version: 0.92.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-dependencies
...

Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2026-06-23 20:32:19 +05:30
Joshua Hoblitt c85643692e ci: retry snapshot controller and crd manifest fetches
Installing the external-snapshotter CRDs and controller runs several
kubectl commands against raw.githubusercontent.com manifest URLs with
no retry, and getManifestFromURL did not retry either, nor did it check
the HTTP status code, so an error page could be fed to kubectl as a
manifest. Transient fetch failures showed up in TestCephSmokeSuite as
NFS test failures during snapshot CRD installation.

Retry the URL-based kubectl invocations and the manifest download, fail
on non-200 responses, and bump the snapshot controller wait in the file
test from 75s to 450s to match the block test since the
snapshot-controller image pull can take longer in CI.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-22 16:27:35 -07:00
Joshua Hoblitt 53ab1c5136 ci: wait for csi driver deployments after cluster install
The csi-operator creates the driver deployments asynchronously, so the
ctrlplugin pods may still be pulling images when the ceph daemons are
already up and the integration tests start. If a PVC is created before
the provisioner is ready, the first provisioning attempt fails (the
mounted csi config may also not be visible to the driver yet) and the
external-provisioner sidecar enters exponential backoff, delaying the
volume creation by several minutes and timing out the test wait.

Observed in TestCephSmokeSuite runs as 'timed out waiting for image
count to reach 1' with the rbd ctrlplugin pod cold-starting mid-test
and the first CreateVolume failing with 'failed to fetch monitor list
using clusterID'. In one run the second provisioning attempt landed two
seconds after the test wait expired.

Wait for the ctrlplugin deployments of the enabled drivers to be ready
before declaring the cluster installed.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-22 16:27:35 -07:00
Joshua Hoblitt 4fb04e638d test: add Eventually polling helper for test setup waits
The keystone auth suite installs cert-manager and trust-manager with
helm --wait and then immediately applies Issuer, Certificate, and
Bundle resources. helm --wait returns when the webhook deployments
report ready, but the webhook service endpoints may not be programmed
on the apiserver's node yet, so the apply is rejected with 'failed
calling webhook ... connect: connection refused' and the suite fails
before any test runs. This is currently the dominant failure mode of
the keystone suite, seen on master and on PRs that cannot have caused
it (dependabot github-actions bumps, mergify backports).

Add utils.Eventually, a timeout-based polling primitive intended as
the shared replacement for hand-rolled retry loops in test code:

- cond is func(ctx) error rather than func() bool, so each failed
  attempt logs its reason via t.Logf and the final timeout error wraps
  the last attempt's error.
- cond runs on the calling goroutine and receives a context carrying
  the overall deadline, so cooperative operations stop at the deadline
  instead of overrunning it. It is never run on another goroutine: a
  hung attempt would keep executing concurrently with its retry, and
  require's FailNow is unsafe off the test goroutine.
- utils.AttemptTimeout decorates a cond with a per-attempt deadline
  for operations that can hang but would succeed if canceled and
  retried. The bound is cooperative; conds that cannot honor a context
  must be bounded at the operation level instead.

Expose it as a K8sHelper.ApplyWithRetry method that retries kubectl
apply until the webhooks accept connections, and use it for the
keystone setup applies instead of failing the suite on the first
attempt.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-22 16:17:46 -07:00
Travis Nielsen dc50bce4e6 Merge pull request #17559 from OdedViner/network_policy_upstream
security: add example NetworkPolicy CRs for Ceph operand pods
2026-06-17 12:54:24 -06:00
Travis Nielsen 923302d7e7 Merge pull request #17714 from jhoblitt/ci-upgrade-suite-flakes
ci: reduce CephUpgradeSuite flakes
2026-06-16 15:14:03 -06:00
Oded Viner 9dc70051a8 security: add example NetworkPolicy CRs for Ceph operand pods
Add a single example YAML with NetworkPolicy definitions
for all Ceph operand pods (mon, osd, mgr, mds, exporter,
osd-prepare, crashcollector, tools).
Users can apply these policies to restrict egress/ingress
traffic for Rook-Ceph daemon pods.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-06-16 17:44:34 +03:00
Joshua Hoblitt faa2fb684b ci: remove unused EnableCsiOperator test setting
The CSI operator is installed unconditionally by the integration test
installer, so the EnableCsiOperator setting no longer selects between
CSI management modes and is consumed nowhere in the test framework.
Remove the setting and the three suites' usages of it.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-12 11:32:05 -07:00
Joshua Hoblitt c4f8831a4d tests: remove unused k8s helper functions
The package-level VersionAtLeast function had no callers anywhere in the
repo (the K8sHelper.VersionAtLeast method remains in use), and
IsKubectlErrorNotFound's only caller was provisioners.go, which is
deleted in this PR. Verified with deadcode and repo-wide grep.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 15:42:39 -07:00
Joshua Hoblitt cff8b7f311 ci: reduce CephUpgradeSuite flakes
The CephUpgradeSuite workflow failed ~31% of PR runs vs the ~19%
broken-PR baseline of the other integration suites. Classifying the
failing step of every failed run from the last ~400 PR runs and the
logs of every failure on non-dependabot branches shows the excess
comes from a handful of too-tight waits, unretried network fetches,
and environment races rather than from the upgrade logic itself:

- The biggest class (21 of 44 genuine test-phase failure runs):
  "giving up waiting for deployment(s) with label
  app=rook-ceph-{rgw,osd},ceph-version!=..." during the
  Squid->Tentacle upgrade. The operator updates daemons sequentially
  (mons, mgr, osds with PG health gates between them, mds, rgw), so
  each later daemon's fixed 275s wait also absorbs the time spent on
  every daemon before it; rgw, last in the sequence, failed most
  often. The mon wait was already extended for slow image pulls in
  98bd03a42; extend the same retry count to all the daemon waits.

- ~7 runs: "snapshot controller is not ready" in the Helm upgrade
  path. WaitForSnapshotController(30) allows only 150s, and the logs
  show the deployment still converging (readyreplicas 1 < replicas 2)
  on the final poll. Raise to 90 retries.

- InstallOrUpgradeHelmRepoChart ran helm with no retry; an observed
  failure fetched the ceph-csi-drivers chart tarball from GitHub
  release assets and got a 504. Retry up to 5 times, as
  InstallLocalHelmChart already does.

- Raise the go test timeouts (2400s->3600s rook, 1800s->2700s helm).
  Failed runs ended in "panic: test timed out" during teardown, which
  aborts cleanup and log collection; the longer daemon waits above
  also need the headroom.

- The "setup cluster resources" composite action (shared by all
  suites) accounted for half of all failed jobs. The one steady class
  there (8 distinct runs across 7 different days): minikube exits with
  K8S_FAIL_CONNECT (code 40) when the requested kubernetes version is
  missing from its bundled version list, because it then validates the
  version with an anonymous GitHub API request (GITHUB_TOKEN is not
  honored on that code path), and anonymous requests from shared
  runner IPs are regularly rate-limited. minikube 1.38.x predates
  v1.35.5, so only the v1.35.5 jobs hit this class. Pass --force to
  minikube start to skip that check: the kubernetes versions used in
  CI are pinned constants already validated by the PRs that bump
  them, and an invalid version would still fail fast at the kubeadm
  download. Also seen twice: dpkg failing on a corrupt cri-dockerd .deb
  because curl ran without --fail and saved an HTTP error page as the
  package; add --fail and retries.

- On runners without the /mnt resource disk, the fallback OSD disk is
  an iSCSI (LIO) device and use_local_disk_for_integration_test
  returned early on those runners, skipping the udev nowatch
  workaround for the device re-probe storms of rook issue 8975. A burn-in
  failure on this PR captured the consequence with full kernel
  forensics: ~100 udev change events on the OSD disk, and ceph-volume
  activate wedged in uninterruptible sleep on the block device lock
  (blkdev_llseek -> rwsem_down_write_slowpath) for 18+ minutes during
  the ceph version upgrade while the cluster's only OSD stayed down.
  Install the nowatch rule before the early return so it applies to
  every runner type, and before the disk is first written rather than
  after.

- One burn-in round failed before the upgrade even began: the
  pre-upgrade PVC create on the Helm path expired WaitUntilPVCIsBound
  (RetryLoop, 275s) while the CSI provisioner pods were still
  starting. Set RETRY_MAX=110 for this workflow to double the
  framework's base wait budget, instead of extending individual waits
  one flake at a time.

- minikube also exits with GUEST_START (code 80) when its internal 6
  minute node-ready wait expires on slow runners (5 runs in the
  dataset, two more observed while burning in this PR, one of them on
  another suite). Pass --wait-timeout=15m to minikube start.

Not addressed (small or episodic classes): OSDs never coming up on
initial deploy (3 runs, possibly the same udev/iSCSI wedge), the
filesystem not becoming active (2 runs), and several day-clustered
minikube incidents (exit codes 65/67/90).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 15:27:02 -07:00
Joshua Hoblitt dbfdd818db tests: remove unused test framework helpers
Delete three files from tests/framework that are entirely dead code; no
symbol in any of them is referenced anywhere in the repo, including the
build-tagged integration suites (verified with deadcode and repo-wide
grep):

- clients/read_write.go: abandoned NFS test client (ReadWriteOperation
  and its methods)
- installer/provisioners.go: hostpath PV helpers orphaned when the
  cassandra and nfs installers were removed
- installer/device.go: IsAdditionalDeviceAvailableOnCluster, unused
  since the upgrade test rework

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 14:56:41 -07:00
Oded Viner 254965fff1 object: remove AWS SDK v1 dependency
Remove the AWS SDK v1 (github.com/aws/aws-sdk-go)
dependency entirely. All S3 operations now use AWS SDK v2
exclusively.

- Remove the v1 Client field from S3Agent struct and
  rename ClientV2 to Client
- Remove v1 session/client initialization from NewS3Agent
- Update all call sites referencing ClientV2
- Convert integration tests to use v2 API calling
  conventions (context parameter) and smithy error handling
- Remove aws-sdk-go v1.55.8 from go.mod

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-05-27 18:35:54 +03:00
subhamkrai 34c5ad7049 build: update csi-operator to v1.0.1
Updating csi-operator to latest v1.0.1 and
updating the required doc changes as well.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-05-27 18:05:56 +05:30
Travis Nielsen 888e725c40 Merge pull request #17555 from subhamkrai/update-csi-operator-1.0
csi: update csi-operator to v1.0.0
2026-05-21 12:32:46 -06:00
subhamkraiandTravis Nielsen 06452991bf csi: update csi-operator to v1.0.0
this commit update the csi-operator to latest v1.0.0.
we need to manually patch the csi drivers with the new serviceAccount
name that are being creating based on the latest ceph-csi-operator
release v1.0.0 where serviceAccount are being separate from csi-operator
chart.

co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2026-05-21 21:52:32 +05:30
Joshua Hoblitt 9eec9c730e core: rm unused consts
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-05-14 09:34:52 -07:00
Oded Viner f5df592106 object: migrate notification package to aws sdk v2
migrate the notification provisioner and s3ext packages from
aws sdk v1 to v2. the custom DeleteBucketNotification call is
rewritten to manually build and sign the HTTP request using the
v2 v4 signer, since this ceph-specific API has no sdk equivalent.

also fix t.Skipped() -> t.Skip() in notification integration test.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-05-13 19:06:52 +03:00
subhamkrai 4eefad42e8 csi: move csi management to admin
Going forward, admin will manage the csi operator
CR's and rook will only manage Ceph Connection cr
and client Profile cr.

The old csi driver is completely removed from Rook
and can no longer be used starting in Rook v1.20.

The upgrade guide will contain the needed transition steps
for managing the csi operator settings.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-04-29 14:24:26 -06:00
Travis Nielsen 1f205c4be4 ci: upgrade tests from v1.19 to master
The upgrade integration tests for the v1.20 release need to
validate the upgrade starting from v1.19 to the latest.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2026-04-29 09:19:40 -06:00
Oded Viner 7f3e2ace27 security: grant scc to rook-ceph-nvmeof service account
the rook-ceph-nvmeof service account is not listed in the
securitycontextconstraints users, causing cephnvmeofgateway
pods to fail on openshift. add the service account to the
scc in the helm charts, go scc builder, and test namespace
substitution.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-04-27 16:28:15 +03:00
Praveen M 6b55092abf csi: update CSI sidecars to latest versions available
Updated the following csi sidecars to their latest available versions:
- csi-attacher: v4.11.0
- csi-snapshotter: v8.5.0
- csi-resizer: v2.1.0
- csi-provisioner: v6.1.1
- csi-node-driver-registrar: v2.16.0

Signed-off-by: Praveen M <m.praveen@ibm.com>
2026-02-26 10:25:49 +05:30
Praveen M 31c7592d03 csi: update external sidecar images
Below sidecars are updated with latest available versions

csi-attacher: v4.10.0
csi-snapshotter: v8.4.0
csi-resizer: v2.0.0
csi-provisioner: v6.0.0
csi-node-driver-registrar: v2.15.0

Signed-off-by: Praveen M <m.praveen@ibm.com>
2026-01-14 14:47:11 +05:30
Travis Nielsen 94e9e2ba91 ci: test rook upgrade from 1.18
The upgrade test is expected to check from the previous minor
release to the current master or latest release. With the
release of 1.19 coming soon, the upgrade test will now
check from v1.18.x.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-12-15 11:44:38 -07:00
subhamkrai 304cdc8c83 core: remove support for reef v18 in Rook v1.19
For Rook v1.19,removing support for Ceph Reef. It is at end of life.
Users on Reef can continue to use Rook v1.18.x or older.
The latest two Ceph versions Squid and Tentacle are supported.

Signed-off-by: subhamkrai <srai@redhat.com>
2025-12-04 13:11:35 +05:30
Travis Nielsen 83047a17d1 operator: fix concurrency issues in the cluster controller
Use a sync map for the clusterMap in the cluster controller
and fix other issues to allow for concurrent cluster
reconcile.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-11-17 11:46:33 -07:00
subhamkraiandTravis Nielsen c57a47f774 ci: run csi-operator only in canary and upgrade suite
this commit add check to only run the csi-operator in
all the canary tests and upgrade suite only, other suite
like smoke and object will still test csi-driver.

Also, adding changes to make CI happy.

Signed-off-by: subhamkrai <srai@redhat.com>
Co-Authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2025-08-19 11:23:14 +05:30
Travis Nielsen 86d30082ab tests: upgrade from 1.17 to master
For the integration tests, now we start upgrading from v1.17.x
to the master branch, to test that the upgrades are stable
to what will become the 1.18 release.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-08-06 15:02:50 -06:00
subhamkrai 8bd4b38399 test: add pathType with ingress dashboard host
Recentally, helm related ci started failing due to
```
testutil: failed to install chart. result=. walk.go:75: found symbolic link in path: /home/runner/work/rook/rook/deploy/charts/rook-ceph-cluster/charts/library resolves to /home/runner/work/rook/rook/deploy/charts/library. Contents of linked file included and used
I0714 09:44:18.963620   18291 warnings.go:110] "Warning: path /ceph-dashboard(/|$)(.*) cannot be used with pathType Prefix"
Error: INSTALLATION FAILED: 1 error occurred:
	* admission webhook "validate.nginx.ingress.kubernetes.io" denied the request: ingress contains invalid paths: path /ceph-dashboard(/|$)(.*) cannot be used with pathType Prefix, err=failed to run helm command on args [install --create-namespace rook-ceph-cluster /home/runner/work/rook/rook/deploy/charts/rook-ceph-cluster --namespace helm-ns -f values-test.yaml] : . walk.go:75: found symbolic link in path: /home/runner/work/rook/rook/deploy/charts/rook-ceph-cluster/charts/library resolves to /home/runner/work/rook/rook/deploy/charts/library. Contents of linked file included and used
I0714 09:44:18.963620   18291 warnings.go:110] "Warning: path /ceph-dashboard(/|$)(.*) cannot be used with pathType Prefix"
Error: INSTALLATION FAILED: 1 error occurred:
```
as default `pathType` is `Prefix`, we need to change it to
`pathType:ImplementationSpecific`.

Signed-off-by: subhamkrai <srai@redhat.com>
2025-07-15 12:50:20 +05:30
Michael AdamandTravis Nielsen fc10f61444 ci: stop using helm.sh up
use setup-helm instead.

Signed-off-by: Michael Adam <obnox@samba.org>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2025-06-12 19:47:24 +02:00
Travis Nielsen d252466f6f ci: test tentacle devel in daily builds
With the tentacle release approaching, let's start
running the Rook tests against the tentacle devel
images to catch if any issues.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-05-30 13:52:56 -06:00
Carlos Barria 030ada45b0 core: fix golangci-lint check QF1003 QF1002
Refactored multiple conditional if-else blocks into switch statements to improve readability and maintainability. Also removed outdated staticcheck rule comments (QF1002, QF1003) from .golangci.yaml.

Signed-off-by: Carlos Barria <cbarria@yahoo.com>
2025-05-27 13:49:49 -04:00