The daily CI smoke, object, and other suites now run
all of the known Ceph versions, including stable
Squid, Tentacle, and Umbrella, as well as devel
versions of the same releases.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The canary that checks every RGW zone.json *_pool field is covered by
Rook's zonePoolNSSuffix map was inline in runObjectE2ETest, running
against the legacy per-pass store before the shared store existed.
Move it into tests/integration/object/zonepools as a standalone
shared-store consumer. It now validates the shared store's zone, which
carries the real shared-pool placements the canary is meant to guard,
and runs first among the shared-store packages so it still sees a fresh
zone.
Add a Sharedstore.Installer() accessor so the packaged canary can run
radosgw-admin inside the cluster while keeping the uniform
(t, k8sh, store) package entry signature.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Changes:
- Use a dedicated .nvmeof pool (via CephBlockPool CR named
builtin-nvmeof) for the NVMe-oF gateway internal state.
- Use a separate nvmeof pool for the StorageClass data.
- Remove the pool field from the CephNVMeOFGateway CRD.
The gateway now always uses the .nvmeof pool, hardcoded
in the operator.
- Add .nvmeof to the allowed CephBlockPool name overrides.
- Create a production example nvmeof.yaml (instances: 2,
replicas: 3) and a CI-only nvmeof-test.yaml (instances: 1,
replicas: 1).
- Update documentation and CI test script accordingly.
Signed-off-by: Oded Viner <oviner@redhat.com>
Add tests/integration/object/bucket/lifecycle, converting the bucket
lifecycle slice of the legacy testObjectStoreOperations onto the shared
store. An OBC is created with a bucketLifecycle additionalConfig, the
rules are verified on the rgw bucket over the S3 API, then updated and
removed with the bucket polled until it reflects each change. The
lifecycleCmpOpts comparison options move out of the suite dispatcher into
the package, and the Ceph v19.2.3 removal-regression skip now discovers
the CephCluster by listing the store's namespace. It runs against the
shared store in both the TLS and non-TLS passes.
Drop the converted lifecycle block from the legacy test.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Add nightly smoke tests and installer support for the umbrella-devel Ceph image.
Also, refactored the test to use reusable code, and since we run the tests in
k8s, removed the k8s matrix from the input.
Signed-off-by: subhamkrai <srai@redhat.com>
Generalize the object suite's sharedstore.Create to accept the target
namespace, store name, and RGW instance count instead of the hardcoded
object-ns / sharedstore / single instance, then use it to stand up the
SmokeSuite object store.
TestObjectStorage_SmokeTest previously created its store with
runObjectE2ETestLite, which applies a CephObjectStore manifest and polls
for health on fixed intervals. It now builds the store through the
typed-client, watch-based sharedstore fixture and tears it down with
Destroy. lite-store is the only object store in the smoke namespace, so
the fixture safely owns the cluster-global .rgw.root realm pool that
Destroy deletes.
The object suite's own call is updated to pass its namespace,
"sharedstore", and one instance, preserving its existing behavior.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Configure the CephHelmSuite object store as an RGW multisite zone so the
chart's CephObjectRealm, CephObjectZoneGroup, and CephObjectZone templates
are exercised on a Helm install. The operator must reconcile the realm,
zone group, zone, and store to Ready, which the existing install check
already gates on via the RGW pod count.
The realm/zone group/zone and shared-pool layout mirror the fixture in
tests/integration/object/util/sharedstore: a zone backed by explicit
CephBlockPools referenced through sharedPools.poolPlacements. The rgw pools
are appended to the block pools already configured for CSI, and are torn
down along with the multisite CRs when the default storage CRs are removed.
Gate the multisite configuration behind a new UseMultisiteObjectStore
setting enabled only for CephHelmSuite. The other Helm-based suites keep the
plain object store; the upgrade suite in particular installs an older base
chart that lacks the multisite templates and cannot resolve a zone-based
store.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
This adds a make target lint.markdown-links
for checking forbroken links in the markdown
documentation sources.
Thid is intended to be used in the CI as well as
locally by developers
The link checking is implemented using the markdown-link-check tool
but it is self-contained in that it does not require the
tool to be installed on the system.
The only prerequisite for using this target is to have
a working docker or podman available.
Assisted-by GitHub copilot
Assisted-by: Google Antigravity/Gemini
Assisted-by: IBM Bob
Signed-off-by: Michael Adam <obnox@samba.org>
Add tests/integration/object/bucket/policy, converting the bucket policy
slice of the legacy testObjectStoreOperations onto the shared store: an OBC
created with a bucketPolicy additionalConfig, the policy verified verbatim
on the rgw bucket over the S3 API, then updated and removed with the bucket
polled until it reflects each change. The policy JSON is generated from the
bucket name, which the legacy test hardcoded in the Resource ARN. It runs
against the shared store in both the TLS and non-TLS passes.
Drop the converted policy block from the legacy test.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Converting the bucket quota tests to typed clients left the framework
helpers and package globals they were the last consumers of unused. Remove
UpdateObc, CheckOBMaxObject, GetAccessKey, and GetSecretKey from the bucket
client, GetEndPointUrl from the object client, and the orphaned object-test
globals.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Add tests/integration/object/bucket/quota, covering both quota behaviors
from the legacy testObjectStoreOperations on OBCs of its own: user quota
via the OBC maxObjects additionalConfig (enforced, then raised and
re-enforced, synced on the ObjectBucket reflecting the new limit) and
bucket quota via bucketMaxObjects switched to bucketMaxSize. It runs
against the shared store in both the TLS and non-TLS passes.
Drop the converted S3-access and bucket-quota blocks from the legacy test;
the main OBC now exists only as a dependents fixture for the store
deletion checks.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Add obc.Update, a get/mutate/update helper for live ObjectBucketClaims,
generalized from the notification suite's label updater, which becomes a
thin wrapper over it. The bucket quota conversion needs the same shape for
additionalConfig updates.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The integration and canary suites ran on a single-node minikube
`driver: none` cluster, where the kubelet ran directly on the GitHub
runner, so the host docker daemon doubled as the cluster runtime and
host block devices and host paths were directly visible to pods. Replace
that with a kind cluster.
Every suite creates its cluster through the shared
integration-test-setup-cluster-resources composite action, so the
conversion is centralized there and converts the smoke, object, helm,
keystone, multi-cluster, upgrade, on-release, nightly, encryption-KMS and
all canary jobs at once.
- Replace the setup-minikube step with helm/kind-action, selecting the
kubernetes version via the kindest/node image tag and creating a
single-node cluster from a new kind config (kind pinned to v0.32.0 for
reproducibility).
- Drop the cri-dockerd install; kind nodes use their built-in containerd.
- Add a kind config that bind-mounts the host /dev, /var/lib/rook and
/run/udev into the node so the existing host-based disk-prep helpers
(use_local_disk*, create_partitions_for_osds, blockDevicePV.sh,
localPathPV.sh, ...) keep working unchanged: devices and partitions
created on the host appear in the node and in the OSD pods that
hostPath-mount the node /dev, and ceph-volume can read the host udev
database it needs to inventory disks.
- Prepare the kind node for the host-level operations rook runs against the
underlying host: remount /sys read-write so CSI's kernel RBD mapping
(`rbd map --device-type krbd`, which writes /sys/bus/rbd) works, and install
lvm2 and cryptsetup, which rook runs in the node's mount namespace to
provision LVM- and encryption-backed OSDs. kindest/node images provide none
of this; the minikube driver:none runner host did.
- Route the Service and pod CIDRs from the runner to the kind node so
host-side tests (the `go test` process runs on the runner) can reach
in-cluster ClusterIPs, e.g. an S3 request to the RGW service. With minikube
driver:none the runner already shared the cluster network.
- Load locally built images into the cluster. Under minikube `driver: none`
the built image was already in the cluster runtime; under kind it must be
imported, so build_rook and create_helm_tag now import their images into
each node's containerd through a new load_image_into_cluster helper (via the
node's ctr, which avoids the kind/kindest-node containerd-config version
skew that breaks `kind load docker-image`).
- Point Vault's kubernetes-auth at the in-cluster API endpoint
(kubernetes.default.svc) instead of the kubeconfig server URL: kind exposes
that as https://127.0.0.1:<port>, unreachable from the in-cluster Vault pod,
so OSD encryption-key retrieval via k8s-auth failed.
- Replace the remaining direct minikube references in the canary workflow: a
`minikube kubectl` call and the external-cluster topology values.
- Adapt host-name assumptions that only held under driver:none: resolve the
disk-cleanup job by the k8s node name rather than the runner hostname, and
let kind-action ignore post-job cluster-teardown failures (the runner is
ephemeral; nvme/multus devices can wedge `docker rm` of the node).
- Update stale comments that described the CI environment as minikube.
- Move the multus integration test's kind config under tests/config too, so
both kind cluster configs live in one place.
create-dev-cluster.sh and other local-dev tooling are intentionally left on
minikube.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The README still described the util layout from before the package
consolidation, with ready, admin, sns, s3, and tls as standalone packages
and a code sample calling ready.ObjectStoreUser. The readiness predicates
live in wait4 and the client builders in client; update the package list,
the code sample, and the playbook references to match.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Add tests/integration/object/bucket/rw, the first slice carved from the
legacy testObjectStoreOperations. It provisions an OBC on the shared store,
writes, reads back, and deletes an object over S3, and checks the OBC stays
Bound. It runs against the shared store in both the TLS and non-TLS passes.
Drop the converted put/get and OBC-revert subtests from the legacy test;
the user quota subtest now seeds both of the objects it needs. The
remaining quota, policy, lifecycle, and dependents slices convert in later
PRs.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Introduce util/obc as the home for ObjectBucketClaim test helpers: the
provisioner StorageClass constructor (moved out of util/fixture, where it
was a misplaced pure constructor), the create/bound and delete/absent
lifecycle waiters, and a per-OBC S3 client. The lifecycle waiters and S3
client are promoted from the notification suite, which built them inline as
their only consumer.
Retarget the StorageClass callers (topic/kafka, bucket/owner) and rework
the notification suite onto the shared helpers.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Add client.NewS3Agent, which builds an rgw S3 agent for an object store
from a set of credentials and the store's endpoint. The object tests build
S3 clients from several credential sources (the store's dashboard admin
user, an OBC's provisioned secret); this factors out the endpoint and agent
construction they share.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
find_extra_block_dev piped find_extra_block_devs into head -1. Since
find_extra_block_devs prints every eligible disk and ends with echo
"$devs", head closes the pipe after the first line and the producer's
echo races into SIGPIPE (echo: write error: Broken pipe), failing the
command substitution under pipefail/errexit.
The race only fires reliably now that the runners present two data disks
(sdb and sdc), so find_extra_block_devs emits multiple lines; with a
single disk the one-line output almost never triggered it.
Read the producer to completion with mapfile via process substitution
and return the first element, mirroring find_second_block_dev, so there
is no downstream head to close the pipe early.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Converting the bucket notification test to typed clients left the
framework notification plumbing it was the last consumer of unused.
Remove the NotificationOperation client, TopicClient.CreateHTTPServer,
BucketClient.CreateObcNotification and CheckBucketNotificationSetonRGW
(and the now-dead UpdateObcNotification helpers), and the
GetBucketNotification and GetOBCNotification manifest builders.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Rework the bucket notification integration test onto the shared
CephObjectStore fixture and the wait4 toolkit, following the object
suite conversion playbook. The old test stood up its own object store
and drove everything through tests/framework/clients and kubectl with
fixed-interval polling.
The new tests/integration/object/notification package is an ordered
t.Run script on typed clients and watch-based waits. It deploys an HTTP
sink, drives a CephBucketTopic (HTTP endpoint), CephBucketNotification,
and a notification-labelled OBC, and verifies both end-to-end delivery
(matched in the sink's logs) and the notification configured on the rgw
bucket (read over the S3 API), covering the label add/remove and
notification-before-topic orderings.
Running against the shared store, it no longer creates a dedicated
object store and now exercises both the TLS and non-TLS suite passes.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The AWS SDK v1 to v2 migration in #17468 left the S3Agent wrapper methods
calling the SDK with context.TODO(), so S3 operations could not be cancelled
when the controller context is cancelled. Add a leading context.Context
parameter to CreateBucket, PutObjectInBucket, GetObjectInBucket,
DeleteObjectInBucket, PutBucketPolicy, and GetBucketPolicy, forward it to the
underlying client, and pass clusterInfo.Context from the bucket provisioner.
Resolves#17526
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
The //nolint directives at these sites omit the linter name, so each
suppresses every linter on its line rather than the one check it needs.
That hides any unrelated errcheck/gosec/govet finding later introduced
on the same line. Name the specific linter for each:
- staticcheck for the two operator sites: SA4004 (the intentional
single-iteration loop in the OSD PVC host lookup) and SA1019 (the
deliberate read of the deprecated S3.Enabled field in the RGW
API-enable builder).
- errcheck for the rbd-mirror deferred token-file cleanup and the
test-framework logging helpers (WriteString / writeHeader).
No behavior change.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Rework the COSI driver integration test onto the shared-store object test
toolkit: an ordered t.Run script with the (t, k8sh, store) signature, typed
clients, and watch-based waits, replacing kubectl-string manifests and
fixed-interval polling.
Drive the COSI bucket resources (BucketClass, BucketClaim, Bucket) with the
upstream sigs.k8s.io/container-object-storage-interface typed client, exposed
as k8sh.COSIClientset alongside the existing OBC client. Create the
CephCOSIDriver and its privileged user through typed clients, wait for the
driver Deployment with wait4, and verify the provisioned bucket through the
shared store's rgw admin client.
Install the COSI CRDs and central controller from the consolidated upstream
repo pinned to v0.2.2 (the former -api and -controller repos are retired) via
a kubectl -k fixture that is removed with t.Cleanup. The driver cannot trust
a TLS RGW endpoint, so the suite skips itself in the TLS pass rather than
being special-cased in the dispatcher.
Retire the now-unused COSIOperation client and the GetCOSIDriver,
GetBucketClass, and GetBucketClaim manifest helpers.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Add a conventions doc and conversion playbook so the remaining old-style
object tests (store lifecycle, notifications, COSI, bucket rw/quota/policy/
lifecycle) can be converted to this pattern mechanically.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The object store user (caps/keys/opmask), bucket-owner, and bucket topic
tests already shared a CephObjectStore but polled with count-based
utils.Retry behind a wide entry signature. Move them onto the wait4 toolkit
and a slim (t, k8sh, store) signature: each is an ordered t.Run script using
typed clients and watch-based waits, with per-package check helpers for
repeated verification. Extend the shared store fixture to provision the
realm/zonegroup/zone pools and serve TLS, build its rgw admin and SNS clients
through the consolidated util/client package (replacing the separate
admin/sns/s3 packages), and run every package in both the TLS and non-TLS
passes (now separate parallel jobs).
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Add the shared helpers the object test packages build on:
- wait4: watch-based Assert/Require waiters (Create/Delete/Condition/Absent)
that block on Kubernetes watch events rather than fixed-interval polling,
plus Eventually for non-k8s state (rgw admin/S3/SNS), a pod-log waiter, and
the readiness predicates used with them.
- secrets: verification of the Secret references CRDs publish in status.
- fixture: create-with-t.Cleanup helpers for pure-cleanup resources, and
constructors for ObjectBucketClaim test resources.
- client: rgw admin and SNS client builders, S3 credentials/endpoint, and the
store's TLS cert plus a verification-skipping HTTP client.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
find_extra_block_devs reads the runner's current extra disks with
`lsblk ... | egrep -v "($boot_dev|loop|nbd)"` to decide how many more to
provision. When the runner has no extra disk, that egrep matches nothing
and exits 1; under the script's `set -eo pipefail` the assignment aborts
the function before it ever calls create_extra_disk, so callers get zero
disks and fail with "expected >= N extra disks, found 0" (or a bare /dev/
path) on those runners.
Guard that first read with `|| true` (as the adjacent device-count line
already is) so an empty result means "no extra disks yet" and the
provisioning path runs. The second read, after create_extra_disk, is left
strict on purpose: if provisioning produced no disk, egrep's non-zero exit
should abort there rather than silently return nothing.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
build_rook() captures the make output into $o and inspects it for
transient network errors (connection reset, INTERNAL_ERROR, 503, 500)
to decide whether to retry the build. The capture was stdout-only,
but go and make write those errors to stderr, so $o never matched any
retry case. A transient module-proxy failure (e.g. proxy.golang.org
returning INTERNAL_ERROR while downloading modules) therefore fell
through to the "valid failure" branch and failed the job in the build
step, before the integration test even ran.
Redirect stderr into the captured output (2>&1) so the existing retry
logic actually triggers on network failures.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>