Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The //nolint directives at these sites omit the linter name, so each
suppresses every linter on its line rather than the one check it needs.
That hides any unrelated errcheck/gosec/govet finding later introduced
on the same line. Name the specific linter for each:
- staticcheck for the two operator sites: SA4004 (the intentional
single-iteration loop in the OSD PVC host lookup) and SA1019 (the
deliberate read of the deprecated S3.Enabled field in the RGW
API-enable builder).
- errcheck for the rbd-mirror deferred token-file cleanup and the
test-framework logging helpers (WriteString / writeHeader).
No behavior change.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Rework the COSI driver integration test onto the shared-store object test
toolkit: an ordered t.Run script with the (t, k8sh, store) signature, typed
clients, and watch-based waits, replacing kubectl-string manifests and
fixed-interval polling.
Drive the COSI bucket resources (BucketClass, BucketClaim, Bucket) with the
upstream sigs.k8s.io/container-object-storage-interface typed client, exposed
as k8sh.COSIClientset alongside the existing OBC client. Create the
CephCOSIDriver and its privileged user through typed clients, wait for the
driver Deployment with wait4, and verify the provisioned bucket through the
shared store's rgw admin client.
Install the COSI CRDs and central controller from the consolidated upstream
repo pinned to v0.2.2 (the former -api and -controller repos are retired) via
a kubectl -k fixture that is removed with t.Cleanup. The driver cannot trust
a TLS RGW endpoint, so the suite skips itself in the TLS pass rather than
being special-cased in the dispatcher.
Retire the now-unused COSIOperation client and the GetCOSIDriver,
GetBucketClass, and GetBucketClaim manifest helpers.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Installing the external-snapshotter CRDs and controller runs several
kubectl commands against raw.githubusercontent.com manifest URLs with
no retry, and getManifestFromURL did not retry either, nor did it check
the HTTP status code, so an error page could be fed to kubectl as a
manifest. Transient fetch failures showed up in TestCephSmokeSuite as
NFS test failures during snapshot CRD installation.
Retry the URL-based kubectl invocations and the manifest download, fail
on non-200 responses, and bump the snapshot controller wait in the file
test from 75s to 450s to match the block test since the
snapshot-controller image pull can take longer in CI.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The csi-operator creates the driver deployments asynchronously, so the
ctrlplugin pods may still be pulling images when the ceph daemons are
already up and the integration tests start. If a PVC is created before
the provisioner is ready, the first provisioning attempt fails (the
mounted csi config may also not be visible to the driver yet) and the
external-provisioner sidecar enters exponential backoff, delaying the
volume creation by several minutes and timing out the test wait.
Observed in TestCephSmokeSuite runs as 'timed out waiting for image
count to reach 1' with the rbd ctrlplugin pod cold-starting mid-test
and the first CreateVolume failing with 'failed to fetch monitor list
using clusterID'. In one run the second provisioning attempt landed two
seconds after the test wait expired.
Wait for the ctrlplugin deployments of the enabled drivers to be ready
before declaring the cluster installed.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The keystone auth suite installs cert-manager and trust-manager with
helm --wait and then immediately applies Issuer, Certificate, and
Bundle resources. helm --wait returns when the webhook deployments
report ready, but the webhook service endpoints may not be programmed
on the apiserver's node yet, so the apply is rejected with 'failed
calling webhook ... connect: connection refused' and the suite fails
before any test runs. This is currently the dominant failure mode of
the keystone suite, seen on master and on PRs that cannot have caused
it (dependabot github-actions bumps, mergify backports).
Add utils.Eventually, a timeout-based polling primitive intended as
the shared replacement for hand-rolled retry loops in test code:
- cond is func(ctx) error rather than func() bool, so each failed
attempt logs its reason via t.Logf and the final timeout error wraps
the last attempt's error.
- cond runs on the calling goroutine and receives a context carrying
the overall deadline, so cooperative operations stop at the deadline
instead of overrunning it. It is never run on another goroutine: a
hung attempt would keep executing concurrently with its retry, and
require's FailNow is unsafe off the test goroutine.
- utils.AttemptTimeout decorates a cond with a per-attempt deadline
for operations that can hang but would succeed if canceled and
retried. The bound is cooperative; conds that cannot honor a context
must be bounded at the operation level instead.
Expose it as a K8sHelper.ApplyWithRetry method that retries kubectl
apply until the webhooks accept connections, and use it for the
keystone setup applies instead of failing the suite on the first
attempt.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The package-level VersionAtLeast function had no callers anywhere in the
repo (the K8sHelper.VersionAtLeast method remains in use), and
IsKubectlErrorNotFound's only caller was provisioners.go, which is
deleted in this PR. Verified with deadcode and repo-wide grep.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The CephUpgradeSuite workflow failed ~31% of PR runs vs the ~19%
broken-PR baseline of the other integration suites. Classifying the
failing step of every failed run from the last ~400 PR runs and the
logs of every failure on non-dependabot branches shows the excess
comes from a handful of too-tight waits, unretried network fetches,
and environment races rather than from the upgrade logic itself:
- The biggest class (21 of 44 genuine test-phase failure runs):
"giving up waiting for deployment(s) with label
app=rook-ceph-{rgw,osd},ceph-version!=..." during the
Squid->Tentacle upgrade. The operator updates daemons sequentially
(mons, mgr, osds with PG health gates between them, mds, rgw), so
each later daemon's fixed 275s wait also absorbs the time spent on
every daemon before it; rgw, last in the sequence, failed most
often. The mon wait was already extended for slow image pulls in
98bd03a42; extend the same retry count to all the daemon waits.
- ~7 runs: "snapshot controller is not ready" in the Helm upgrade
path. WaitForSnapshotController(30) allows only 150s, and the logs
show the deployment still converging (readyreplicas 1 < replicas 2)
on the final poll. Raise to 90 retries.
- InstallOrUpgradeHelmRepoChart ran helm with no retry; an observed
failure fetched the ceph-csi-drivers chart tarball from GitHub
release assets and got a 504. Retry up to 5 times, as
InstallLocalHelmChart already does.
- Raise the go test timeouts (2400s->3600s rook, 1800s->2700s helm).
Failed runs ended in "panic: test timed out" during teardown, which
aborts cleanup and log collection; the longer daemon waits above
also need the headroom.
- The "setup cluster resources" composite action (shared by all
suites) accounted for half of all failed jobs. The one steady class
there (8 distinct runs across 7 different days): minikube exits with
K8S_FAIL_CONNECT (code 40) when the requested kubernetes version is
missing from its bundled version list, because it then validates the
version with an anonymous GitHub API request (GITHUB_TOKEN is not
honored on that code path), and anonymous requests from shared
runner IPs are regularly rate-limited. minikube 1.38.x predates
v1.35.5, so only the v1.35.5 jobs hit this class. Pass --force to
minikube start to skip that check: the kubernetes versions used in
CI are pinned constants already validated by the PRs that bump
them, and an invalid version would still fail fast at the kubeadm
download. Also seen twice: dpkg failing on a corrupt cri-dockerd .deb
because curl ran without --fail and saved an HTTP error page as the
package; add --fail and retries.
- On runners without the /mnt resource disk, the fallback OSD disk is
an iSCSI (LIO) device and use_local_disk_for_integration_test
returned early on those runners, skipping the udev nowatch
workaround for the device re-probe storms of rook issue 8975. A burn-in
failure on this PR captured the consequence with full kernel
forensics: ~100 udev change events on the OSD disk, and ceph-volume
activate wedged in uninterruptible sleep on the block device lock
(blkdev_llseek -> rwsem_down_write_slowpath) for 18+ minutes during
the ceph version upgrade while the cluster's only OSD stayed down.
Install the nowatch rule before the early return so it applies to
every runner type, and before the disk is first written rather than
after.
- One burn-in round failed before the upgrade even began: the
pre-upgrade PVC create on the Helm path expired WaitUntilPVCIsBound
(RetryLoop, 275s) while the CSI provisioner pods were still
starting. Set RETRY_MAX=110 for this workflow to double the
framework's base wait budget, instead of extending individual waits
one flake at a time.
- minikube also exits with GUEST_START (code 80) when its internal 6
minute node-ready wait expires on slow runners (5 runs in the
dataset, two more observed while burning in this PR, one of them on
another suite). Pass --wait-timeout=15m to minikube start.
Not addressed (small or episodic classes): OSDs never coming up on
initial deploy (3 runs, possibly the same udev/iSCSI wedge), the
filesystem not becoming active (2 runs), and several day-clustered
minikube incidents (exit codes 65/67/90).
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
this commit update the csi-operator to latest v1.0.0.
we need to manually patch the csi drivers with the new serviceAccount
name that are being creating based on the latest ceph-csi-operator
release v1.0.0 where serviceAccount are being separate from csi-operator
chart.
co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
Going forward, admin will manage the csi operator
CR's and rook will only manage Ceph Connection cr
and client Profile cr.
The old csi driver is completely removed from Rook
and can no longer be used starting in Rook v1.20.
The upgrade guide will contain the needed transition steps
for managing the csi operator settings.
Signed-off-by: subhamkrai <srai@redhat.com>
Updated the following csi sidecars to their latest available versions:
- csi-attacher: v4.11.0
- csi-snapshotter: v8.5.0
- csi-resizer: v2.1.0
- csi-provisioner: v6.1.1
- csi-node-driver-registrar: v2.16.0
Signed-off-by: Praveen M <m.praveen@ibm.com>
The Kubernetes CSI sidecars have had several releases that were not
included in deployments by Rook yet, update them to the versions that
are available today:
- csi-attacher:v4.8.1
- csi-provisioner:v5.2.0
- csi-resizer:v1.13.2
- csi-snapshotter:v8.2.1
This change is important, because Ceph-CSI will implement the new
Controller.GetSnapshot CSI procedure. A bug in csi-lib-utils causes a
panic when a ControllerCapability is provided, but not (yet) known to
the CSI sidecars. The updated sidecars consume a version of
csi-lib-utils with a fix for that panic.
See-also: kubernetes-csi/csi-lib-utils#188
Signed-off-by: Niels de Vos <ndevos@ibm.com>
When upgrading from one Ceph version to another, the new image to
upgrade to can take a long time to pull in some cases before the upgrade
can even begin. For example, some ceph-ci images regularly take 6+
minutes to pull, which exceeds the timeout waiting for mons to be ready.
Extend the timeout for mons to account for these cases.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Create the CSI operator in the go integration test suites
to test the new CSI operator scenarios. For the upgrade
test suite, enable the CSI operator after the upgrade
to verify the working cluster after it is enabled.
Signed-off-by: subhamkrai <srai@redhat.com>
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.
Additionally, it silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
and string format errors found by golangci-lint, while at it.
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
The mon canaries may be created even when the mon daemons
are not created thereafter during the integration tests.
Therefore, the integration tests need to also query a label
specific to the mon daemon so the canaries are not a distraction
to the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The v1beta1 cron jobs have been obsolete since K8s 1.21,
and Rook has not supported that version of K8s
for many moons, so we can remove the obsolete code
for the handling of v1beta1 cron jobs for crash pruning.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>
* extend the API object specs for swift and keystone integration
* adapt rgw to the new go-ceph version
- The parameter lists of the API call have changes, as parameters
ignored by the RGW Admin Ops API are no longer serialized, therefore
the mock has to be adapted.
- There is now validation for the user keys that are passed to the
User get API, therefore things failed when we had empty keys in our
User proxy object.
* expand the reconcile loop for the swift and keystone integration
* fix minor mistakes in design document
* add env var to pass extra args to minikube
Minikube decides CPU cores and memory automatically based on the
available resources on the machine which may be insufficient to
run rook. This commit adds an environment variable to add arbitrary
arguments to the minikube command, so both can be specified if
desired.
* integration tests for swift and keystone
The new integration of swift or s3 and keystone support by rook
does not have any integration tests yet.
This commit introduces integration tests for swift and keystone. The
tests are done against a minimal keystone setup (keystone container
image from Yaook-project (https://yaook.cloud), sqlite as database
backend, cert-manager and trust-manager for test certificate setup).
To prevent hardcoded credentials, passwords are generated
by the tests. The integration tests use the openstack client
(keystone- and swift-functionality) (https://docs.openstack.org/
python-openstackclient/ latest/). This was a concious design decision
to use client tooling as close as possible to the end user instead of
using other go-libraries (such as gophercloud).
* add documentation on swift and keystone
Currently there is no documentation on the use of Swift to access
an object store as well as the use of OpenStack keystone for
authentication.
This commit adds documentation on the use of Swift and OpenStack
keystone, as well as CRD-related documentation and an example setup.
* add integration tests for S3 via keystone
This commit introduces integration tests for s3 and keystone. The
tests are run against the same minimal keystone setup that the tests
for swift and keystone use.
The integration tests use the aws s3 client to use client tooling as
close as possible to the end user instead of using other go-libraries.
Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
The helm upgrade tests have been failing frequently, but not
always, on the oldest version of K8s that is tested in the CI
for the past few months. Add a retry to attempt to get
the CI passing more consistently.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition
Signed-off-by: parth-gr <paarora@redhat.com>
From the Go specification [1]:
"1. For a nil slice, the number of iterations is 0."
"3. If the map is nil, the number of iterations is 0."
`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.
[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len
Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
for now, let's skip the mgr pod restart count
for upgrade suite 1.22.x to get the CI green
and so that we don't skip any other error in name
of mgr restart count.
Signed-off-by: subhamkrai <srai@redhat.com>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives
Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md
Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.
This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
ci is more resource intensive.
Signed-off-by: Rakshith R <rar@redhat.com>
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
When a test has a slash in it's name the log collection can fail due to
the "directory" not existing, this makes sure the slashes are replaced
by underscores.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.
One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.
Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.
A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.
Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>