208 Commits
Author SHA1 Message Date
Chiman Jain 2034be1f37 build: migrate gopkg.in/yaml to go.yaml.in/yaml
Signed-off-by: Chiman Jain <chimanjain15@gmail.com>
2026-07-22 20:05:39 +05:30
Joshua Hoblitt c0eb360041 docs: fix typos and grammar in code comments
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:34:11 -07:00
Joshua Hoblitt e281600bb5 Merge pull request #17855 from jhoblitt/narrow-bare-nolint
core: narrow bare //nolint directives to specific linters
2026-06-30 08:14:46 -07:00
Joshua Hoblitt e5c6d75a73 core: narrow bare //nolint directives to specific linters
The //nolint directives at these sites omit the linter name, so each
suppresses every linter on its line rather than the one check it needs.
That hides any unrelated errcheck/gosec/govet finding later introduced
on the same line. Name the specific linter for each:

- staticcheck for the two operator sites: SA4004 (the intentional
  single-iteration loop in the OSD PVC host lookup) and SA1019 (the
  deliberate read of the deprecated S3.Enabled field in the RGW
  API-enable builder).
- errcheck for the rbd-mirror deferred token-file cleanup and the
  test-framework logging helpers (WriteString / writeHeader).

No behavior change.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 16:32:59 -07:00
Joshua Hoblitt 3a4dc64f5b test: convert the COSI driver suite to the object test toolkit
Rework the COSI driver integration test onto the shared-store object test
toolkit: an ordered t.Run script with the (t, k8sh, store) signature, typed
clients, and watch-based waits, replacing kubectl-string manifests and
fixed-interval polling.

Drive the COSI bucket resources (BucketClass, BucketClaim, Bucket) with the
upstream sigs.k8s.io/container-object-storage-interface typed client, exposed
as k8sh.COSIClientset alongside the existing OBC client. Create the
CephCOSIDriver and its privileged user through typed clients, wait for the
driver Deployment with wait4, and verify the provisioned bucket through the
shared store's rgw admin client.

Install the COSI CRDs and central controller from the consolidated upstream
repo pinned to v0.2.2 (the former -api and -controller repos are retired) via
a kubectl -k fixture that is removed with t.Cleanup. The driver cannot trust
a TLS RGW endpoint, so the suite skips itself in the TLS pass rather than
being special-cased in the dispatcher.

Retire the now-unused COSIOperation client and the GetCOSIDriver,
GetBucketClass, and GetBucketClaim manifest helpers.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 15:36:13 -07:00
Joshua Hoblitt 49612461a4 docs: fix function comments to match their declaration names
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-26 12:52:39 -07:00
Joshua Hoblitt d0885370fa Merge pull request #17732 from jhoblitt/ci-smoke-suite-flakes
ci: harden CephSmokeSuite against CSI cold-start and teardown flakes
2026-06-24 08:57:25 -07:00
Joshua Hoblitt c85643692e ci: retry snapshot controller and crd manifest fetches
Installing the external-snapshotter CRDs and controller runs several
kubectl commands against raw.githubusercontent.com manifest URLs with
no retry, and getManifestFromURL did not retry either, nor did it check
the HTTP status code, so an error page could be fed to kubectl as a
manifest. Transient fetch failures showed up in TestCephSmokeSuite as
NFS test failures during snapshot CRD installation.

Retry the URL-based kubectl invocations and the manifest download, fail
on non-200 responses, and bump the snapshot controller wait in the file
test from 75s to 450s to match the block test since the
snapshot-controller image pull can take longer in CI.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-22 16:27:35 -07:00
Joshua Hoblitt 53ab1c5136 ci: wait for csi driver deployments after cluster install
The csi-operator creates the driver deployments asynchronously, so the
ctrlplugin pods may still be pulling images when the ceph daemons are
already up and the integration tests start. If a PVC is created before
the provisioner is ready, the first provisioning attempt fails (the
mounted csi config may also not be visible to the driver yet) and the
external-provisioner sidecar enters exponential backoff, delaying the
volume creation by several minutes and timing out the test wait.

Observed in TestCephSmokeSuite runs as 'timed out waiting for image
count to reach 1' with the rbd ctrlplugin pod cold-starting mid-test
and the first CreateVolume failing with 'failed to fetch monitor list
using clusterID'. In one run the second provisioning attempt landed two
seconds after the test wait expired.

Wait for the ctrlplugin deployments of the enabled drivers to be ready
before declaring the cluster installed.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-22 16:27:35 -07:00
Joshua Hoblitt 4fb04e638d test: add Eventually polling helper for test setup waits
The keystone auth suite installs cert-manager and trust-manager with
helm --wait and then immediately applies Issuer, Certificate, and
Bundle resources. helm --wait returns when the webhook deployments
report ready, but the webhook service endpoints may not be programmed
on the apiserver's node yet, so the apply is rejected with 'failed
calling webhook ... connect: connection refused' and the suite fails
before any test runs. This is currently the dominant failure mode of
the keystone suite, seen on master and on PRs that cannot have caused
it (dependabot github-actions bumps, mergify backports).

Add utils.Eventually, a timeout-based polling primitive intended as
the shared replacement for hand-rolled retry loops in test code:

- cond is func(ctx) error rather than func() bool, so each failed
  attempt logs its reason via t.Logf and the final timeout error wraps
  the last attempt's error.
- cond runs on the calling goroutine and receives a context carrying
  the overall deadline, so cooperative operations stop at the deadline
  instead of overrunning it. It is never run on another goroutine: a
  hung attempt would keep executing concurrently with its retry, and
  require's FailNow is unsafe off the test goroutine.
- utils.AttemptTimeout decorates a cond with a per-attempt deadline
  for operations that can hang but would succeed if canceled and
  retried. The bound is cooperative; conds that cannot honor a context
  must be bounded at the operation level instead.

Expose it as a K8sHelper.ApplyWithRetry method that retries kubectl
apply until the webhooks accept connections, and use it for the
keystone setup applies instead of failing the suite on the first
attempt.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-22 16:17:46 -07:00
Travis Nielsen 923302d7e7 Merge pull request #17714 from jhoblitt/ci-upgrade-suite-flakes
ci: reduce CephUpgradeSuite flakes
2026-06-16 15:14:03 -06:00
Joshua Hoblitt c4f8831a4d tests: remove unused k8s helper functions
The package-level VersionAtLeast function had no callers anywhere in the
repo (the K8sHelper.VersionAtLeast method remains in use), and
IsKubectlErrorNotFound's only caller was provisioners.go, which is
deleted in this PR. Verified with deadcode and repo-wide grep.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 15:42:39 -07:00
Joshua Hoblitt cff8b7f311 ci: reduce CephUpgradeSuite flakes
The CephUpgradeSuite workflow failed ~31% of PR runs vs the ~19%
broken-PR baseline of the other integration suites. Classifying the
failing step of every failed run from the last ~400 PR runs and the
logs of every failure on non-dependabot branches shows the excess
comes from a handful of too-tight waits, unretried network fetches,
and environment races rather than from the upgrade logic itself:

- The biggest class (21 of 44 genuine test-phase failure runs):
  "giving up waiting for deployment(s) with label
  app=rook-ceph-{rgw,osd},ceph-version!=..." during the
  Squid->Tentacle upgrade. The operator updates daemons sequentially
  (mons, mgr, osds with PG health gates between them, mds, rgw), so
  each later daemon's fixed 275s wait also absorbs the time spent on
  every daemon before it; rgw, last in the sequence, failed most
  often. The mon wait was already extended for slow image pulls in
  98bd03a42; extend the same retry count to all the daemon waits.

- ~7 runs: "snapshot controller is not ready" in the Helm upgrade
  path. WaitForSnapshotController(30) allows only 150s, and the logs
  show the deployment still converging (readyreplicas 1 < replicas 2)
  on the final poll. Raise to 90 retries.

- InstallOrUpgradeHelmRepoChart ran helm with no retry; an observed
  failure fetched the ceph-csi-drivers chart tarball from GitHub
  release assets and got a 504. Retry up to 5 times, as
  InstallLocalHelmChart already does.

- Raise the go test timeouts (2400s->3600s rook, 1800s->2700s helm).
  Failed runs ended in "panic: test timed out" during teardown, which
  aborts cleanup and log collection; the longer daemon waits above
  also need the headroom.

- The "setup cluster resources" composite action (shared by all
  suites) accounted for half of all failed jobs. The one steady class
  there (8 distinct runs across 7 different days): minikube exits with
  K8S_FAIL_CONNECT (code 40) when the requested kubernetes version is
  missing from its bundled version list, because it then validates the
  version with an anonymous GitHub API request (GITHUB_TOKEN is not
  honored on that code path), and anonymous requests from shared
  runner IPs are regularly rate-limited. minikube 1.38.x predates
  v1.35.5, so only the v1.35.5 jobs hit this class. Pass --force to
  minikube start to skip that check: the kubernetes versions used in
  CI are pinned constants already validated by the PRs that bump
  them, and an invalid version would still fail fast at the kubeadm
  download. Also seen twice: dpkg failing on a corrupt cri-dockerd .deb
  because curl ran without --fail and saved an HTTP error page as the
  package; add --fail and retries.

- On runners without the /mnt resource disk, the fallback OSD disk is
  an iSCSI (LIO) device and use_local_disk_for_integration_test
  returned early on those runners, skipping the udev nowatch
  workaround for the device re-probe storms of rook issue 8975. A burn-in
  failure on this PR captured the consequence with full kernel
  forensics: ~100 udev change events on the OSD disk, and ceph-volume
  activate wedged in uninterruptible sleep on the block device lock
  (blkdev_llseek -> rwsem_down_write_slowpath) for 18+ minutes during
  the ceph version upgrade while the cluster's only OSD stayed down.
  Install the nowatch rule before the early return so it applies to
  every runner type, and before the disk is first written rather than
  after.

- One burn-in round failed before the upgrade even began: the
  pre-upgrade PVC create on the Helm path expired WaitUntilPVCIsBound
  (RetryLoop, 275s) while the CSI provisioner pods were still
  starting. Set RETRY_MAX=110 for this workflow to double the
  framework's base wait budget, instead of extending individual waits
  one flake at a time.

- minikube also exits with GUEST_START (code 80) when its internal 6
  minute node-ready wait expires on slow runners (5 runs in the
  dataset, two more observed while burning in this PR, one of them on
  another suite). Pass --wait-timeout=15m to minikube start.

Not addressed (small or episodic classes): OSDs never coming up on
initial deploy (3 runs, possibly the same udev/iSCSI wedge), the
filesystem not becoming active (2 runs), and several day-clustered
minikube incidents (exit codes 65/67/90).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-11 15:27:02 -07:00
subhamkraiandTravis Nielsen 06452991bf csi: update csi-operator to v1.0.0
this commit update the csi-operator to latest v1.0.0.
we need to manually patch the csi drivers with the new serviceAccount
name that are being creating based on the latest ceph-csi-operator
release v1.0.0 where serviceAccount are being separate from csi-operator
chart.

co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2026-05-21 21:52:32 +05:30
subhamkrai 4eefad42e8 csi: move csi management to admin
Going forward, admin will manage the csi operator
CR's and rook will only manage Ceph Connection cr
and client Profile cr.

The old csi driver is completely removed from Rook
and can no longer be used starting in Rook v1.20.

The upgrade guide will contain the needed transition steps
for managing the csi operator settings.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-04-29 14:24:26 -06:00
Praveen M 6b55092abf csi: update CSI sidecars to latest versions available
Updated the following csi sidecars to their latest available versions:
- csi-attacher: v4.11.0
- csi-snapshotter: v8.5.0
- csi-resizer: v2.1.0
- csi-provisioner: v6.1.1
- csi-node-driver-registrar: v2.16.0

Signed-off-by: Praveen M <m.praveen@ibm.com>
2026-02-26 10:25:49 +05:30
Praveen M 31c7592d03 csi: update external sidecar images
Below sidecars are updated with latest available versions

csi-attacher: v4.10.0
csi-snapshotter: v8.4.0
csi-resizer: v2.0.0
csi-provisioner: v6.0.0
csi-node-driver-registrar: v2.15.0

Signed-off-by: Praveen M <m.praveen@ibm.com>
2026-01-14 14:47:11 +05:30
Carlos Barria 8afc4ee548 core: fix golangci-lint check QF1004
Signed-off-by: Carlos Barria <cbarria@yahoo.com>
2025-05-22 11:21:45 -04:00
Niels de Vos a168c40008 csi: update Kubernetes CSI sidecar images to current versions
The Kubernetes CSI sidecars have had several releases that were not
included in deployments by Rook yet, update them to the versions that
are available today:

- csi-attacher:v4.8.1
- csi-provisioner:v5.2.0
- csi-resizer:v1.13.2
- csi-snapshotter:v8.2.1

This change is important, because Ceph-CSI will implement the new
Controller.GetSnapshot CSI procedure. A bug in csi-lib-utils causes a
panic when a ControllerCapability is provided, but not (yet) known to
the CSI sidecars. The updated sidecars consume a version of
csi-lib-utils with a fix for that panic.

See-also: kubernetes-csi/csi-lib-utils#188
Signed-off-by: Niels de Vos <ndevos@ibm.com>
2025-05-20 14:24:02 +02:00
Carlos Barria 61494d1baf core: fix golangci-lint check ST1005
Signed-off-by: Carlos Barria <cbarria@yahoo.com>
2025-05-19 17:52:22 -04:00
Blaine Gardner 98bd03a420 ci: extend timeout for mons in upgrade test
When upgrading from one Ceph version to another, the new image to
upgrade to can take a long time to pull in some cases before the upgrade
can even begin. For example, some ceph-ci images regularly take 6+
minutes to pull, which exceeds the timeout waiting for mons to be ready.

Extend the timeout for mons to account for these cases.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2025-04-29 13:53:43 -06:00
Joshua Hoblitt 3cb343f62a core: run gofumpt on all files
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2025-03-26 10:41:48 -07:00
Travis Nielsen c7e5a4c236 tests: enable csi operator for go test suites
Create the CSI operator in the go integration test suites
to test the new CSI operator scenarios. For the upgrade
test suite, enable the CSI operator after the upgrade
to verify the working cluster after it is enabled.

Signed-off-by: subhamkrai <srai@redhat.com>
2025-03-19 21:02:05 +05:30
Steven Kreitzer 463d9fb420 csi: csi-snapshotter flag typo; upgrade csi-snapshotter
Signed-off-by: Steven Kreitzer <skre@skre.me>
2024-12-18 09:09:29 -06:00
df511fb58f ci: update golangci-lint to the latest version (v1.62)
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.

Additionally, it  silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
 and string format errors found by golangci-lint, while at it.

Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
2024-12-14 14:47:30 +01:00
Travis Nielsen b52ba6baca test: wait for mon daemons rather than mon canaries
The mon canaries may be created even when the mon daemons
are not created thereafter during the integration tests.
Therefore, the integration tests need to also query a label
specific to the mon daemon so the canaries are not a distraction
to the test.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-10-25 09:17:02 -06:00
Travis Nielsen 7ed77ddd13 core: remove obsolete creation of v1beta1 pruner cron jobs
The v1beta1 cron jobs have been obsolete since K8s 1.21,
and Rook has not supported that version of K8s
for many moons, so we can remove the obsolete code
for the handling of v1beta1 cron jobs for crash pruning.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-10-14 17:13:06 -06:00
Praveen M a1ddf4535d csi: update csi sidecars' image version
Below csi sidecars are updated with latest available versions

csi-resizer: v1.11.1
csi-provisioner: v5.0.1
csi-attacher: v4.6.1
csi-snapshotter: v8.0.1
csi-node-driver-registrar: v2.11.1

Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-08-20 22:08:19 +05:30
ee8bcad49d rgw: add support for keystone auth + swift/s3
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>

* extend the API object specs for swift and keystone integration

* adapt rgw to the new go-ceph version

  - The parameter lists of the API call have changes, as parameters
    ignored by the RGW Admin Ops API are no longer serialized, therefore
    the mock has to be adapted.

  - There is now validation for the user keys that are passed to the
    User get API, therefore things failed when we had empty keys in our
    User proxy object.

* expand the reconcile loop for the swift and keystone integration

* fix minor mistakes in design document

* add env var to pass extra args to minikube

  Minikube decides CPU cores and memory automatically based on the
  available resources on the machine which may be insufficient to
  run rook. This commit adds an environment variable to add arbitrary
  arguments to the minikube command, so both can be specified if
  desired.

* integration tests for swift and keystone

  The new integration of swift or s3 and keystone support by rook
  does not have any integration tests yet.

  This commit introduces integration tests for swift and keystone. The
  tests are done against a minimal keystone setup (keystone container
  image from Yaook-project (https://yaook.cloud), sqlite as database
  backend, cert-manager and trust-manager for test certificate setup).

  To prevent hardcoded credentials, passwords are generated
  by the tests. The integration tests use the openstack client
  (keystone- and swift-functionality) (https://docs.openstack.org/
  python-openstackclient/ latest/). This was a concious design decision
  to use client tooling as close as possible to the end user instead of
  using other go-libraries (such as gophercloud).

* add documentation on swift and keystone

  Currently there is no documentation on the use of Swift to access
  an object store as well as the use of OpenStack keystone for
  authentication.

  This commit adds documentation on the use of Swift and OpenStack
  keystone, as well as CRD-related documentation and an example setup.

* add integration tests for S3 via keystone

  This commit introduces integration tests for s3 and keystone. The
  tests are run against the same minimal keystone setup that the tests
  for swift and keystone use.

  The integration tests use the aws s3 client to use client tooling as
  close as possible to the end user instead of using other go-libraries.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
2024-08-08 14:26:21 +02:00
Travis Nielsen 043b446675 tests: retry helm upgrade during upgrade tests
The helm upgrade tests have been failing frequently, but not
always, on the oldest version of K8s that is tested in the CI
for the past few months. Add a retry to attempt to get
the CI passing more consistently.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-07-24 08:14:21 -06:00
Praveen M faf7837621 csi: update csi sidecars' image version
Below sidecars are updated with latest available versions

csi-node-driver-registrar: v2.10.1
csi-resizer: v1.10.1
csi-provisioner: v4.0.1
csi-attacher: v4.5.1
csi-snapshotter: v7.0.2

Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-04-25 18:58:00 +05:30
Madhu Rajanna 0d5bd70194 csi: install vgs CRD in tests
update the snapshot controller to 7.0.1
and install new Volumegroup CRD's

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-03-01 09:15:17 +01:00
Praveen M a80396df1b csi: update cmdline args as used by ceph-csi
This commit adds cmdline args to enable
1. RecoverVolumeExpansionFailure
2. PreventVolumeModeConversion
3. HonorPVReclaimPolicy

Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-01-12 15:24:07 +05:30
parth-gr 40295a989c ci: fix objectsuite flakiness
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-08 15:33:03 +05:30
Eng Zer Jun 77bff6a07c core: remove redundant len check
From the Go specification [1]:

  "1. For a nil slice, the number of iterations is 0."
  "3. If the map is nil, the number of iterations is 0."

`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.

[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len

Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
2023-10-05 20:16:46 +08:00
subhamkrai 1bd4ab9d81 ci: skip mgr pod restart count upgrade 1.22.x suite
for now, let's skip the mgr pod restart count
for upgrade suite 1.22.x to get the CI green
and so that we don't skip any other error in name
of mgr restart count.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:03:08 +05:30
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
parth-gr 9319eaa86d ci: watch for pods that restart unexpectedly
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives

Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
2022-12-14 21:50:15 +05:30
Travis Nielsen b693a5cca5 core: refactor crash collector for more node daemons
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md

Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-09 16:57:05 -07:00
Blaine Gardner fc56108865 object: allow status endpoints to be null
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.

This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-08 13:36:15 -07:00
Rakshith R be10dac98c ci: enable and fixes for nfs ci
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
  ci is more resource intensive.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-06 11:53:46 +00:00
Liang Zheng f9c360e609 osd: optimize device probe
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 10:41:51 +08:00
Alexander Trost 49440e8b20 test: fix test pod log collector
When a test has a slash in it's name the log collection can fail due to
the "directory" not existing, this makes sure the slashes are replaced
by underscores.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-08-12 23:44:13 +02:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Rakshith R 034c4756d3 ci: make sure PVC is deleted before proceeding for file test
Signed-off-by: Rakshith R <rar@redhat.com>
2022-07-04 11:36:37 +05:30
Blaine Gardner a63844f8bf file: block deletion on more dependents
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.

One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.

Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.

A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.

Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-06-14 13:02:35 -06:00
Travis Nielsen 2db3910052 test: remove dead test helper code
The K8sHelper has a number of methods that are no longer
in use that we can remove.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-25 15:28:29 -06:00
Travis Nielsen fb86955f01 core: examples set default priority class names
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-19 16:09:34 -06:00
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30