Commit Graph
1603 Commits
Author SHA1 Message Date
Travis Nielsen 2bfee41b0b Merge pull request #18063 from malayparida2000/fix/modcheck-error-message
ci: fix mod.check validation error message
2026-07-28 12:30:37 -06:00
Malay Kumar Parida 408dfdebd2 ci: fix mod.check validation error message
Advise running make mod.check and committing results instead of make clean.

Signed-off-by: Malay Kumar Parida <mparida@redhat.com>
2026-07-28 23:35:10 +05:30
Joshua Hoblitt 1207514422 Merge pull request #18033 from jhoblitt/object-zonepools
test: extract the zone.json pool canary into its own package
2026-07-27 16:14:58 -07:00
Travis Nielsen fe33e08780 ci: integration tests default to tentacle version
The integration tests now default to stable Ceph Tentacle
for the PR runs.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2026-07-27 11:43:40 -06:00
Travis Nielsen c73f8ab1c2 ci: daily tests run stable and devel ceph versions
The daily CI smoke, object, and other suites now run
all of the known Ceph versions, including stable
Squid, Tentacle, and Umbrella, as well as devel
versions of the same releases.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2026-07-27 11:42:46 -06:00
Joshua Hoblitt 6bf24ac507 test: extract the zone.json pool canary into its own package
The canary that checks every RGW zone.json *_pool field is covered by
Rook's zonePoolNSSuffix map was inline in runObjectE2ETest, running
against the legacy per-pass store before the shared store existed.

Move it into tests/integration/object/zonepools as a standalone
shared-store consumer. It now validates the shared store's zone, which
carries the real shared-pool placements the canary is meant to guard,
and runs first among the shared-store packages so it still sees a fresh
zone.

Add a Sharedstore.Installer() accessor so the packaged canary can run
radosgw-admin inside the cluster while keeping the uniform
(t, k8sh, store) package entry signature.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-22 16:04:42 -07:00
Joshua Hoblitt 012d69a076 Merge pull request #17894 from jhoblitt/convert-bucket-lifecycle-to-sharedstore
test: convert the object bucket lifecycle suite to the shared store
2026-07-22 12:55:11 -07:00
Chiman Jain 2034be1f37 build: migrate gopkg.in/yaml to go.yaml.in/yaml
Signed-off-by: Chiman Jain <chimanjain15@gmail.com>
2026-07-22 20:05:39 +05:30
Oded Viner ec502c3f80 nvmeof: use .nvmeof pool for gateway and remove pool field from CRD
Changes:
- Use a dedicated .nvmeof pool (via CephBlockPool CR named
  builtin-nvmeof) for the NVMe-oF gateway internal state.
- Use a separate nvmeof pool for the StorageClass data.
- Remove the pool field from the CephNVMeOFGateway CRD.
  The gateway now always uses the .nvmeof pool, hardcoded
  in the operator.
- Add .nvmeof to the allowed CephBlockPool name overrides.
- Create a production example nvmeof.yaml (instances: 2,
  replicas: 3) and a CI-only nvmeof-test.yaml (instances: 1,
  replicas: 1).
- Update documentation and CI test script accordingly.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-07-21 20:19:44 +03:00
Joshua Hoblitt 95a1c4d8ad test: convert the object bucket lifecycle suite to the shared store
Add tests/integration/object/bucket/lifecycle, converting the bucket
lifecycle slice of the legacy testObjectStoreOperations onto the shared
store. An OBC is created with a bucketLifecycle additionalConfig, the
rules are verified on the rgw bucket over the S3 API, then updated and
removed with the bucket polled until it reflects each change. The
lifecycleCmpOpts comparison options move out of the suite dispatcher into
the package, and the Ceph v19.2.3 removal-regression skip now discovers
the CephCluster by listing the store's namespace. It runs against the
shared store in both the TLS and non-TLS passes.

Drop the converted lifecycle block from the legacy test.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-20 13:46:22 -07:00
Joshua Hoblitt 452fb31d95 Merge pull request #17917 from jhoblitt/chart-object-multisite
helm: add multisite CR support to rook-ceph-cluster chart
2026-07-13 13:10:05 -07:00
Joshua Hoblitt 197b6ef843 Merge pull request #17857 from jhoblitt/convert-smoke-object-to-sharedstore
test: build the smoke object store with the shared store fixture
2026-07-13 13:09:31 -07:00
Joshua Hoblitt 64973c65d5 Merge pull request #17884 from jhoblitt/convert-bucket-policy-to-sharedstore
test: convert the object bucket policy suite to the shared store
2026-07-13 13:08:59 -07:00
subhamkrai 458a8a34f8 ci: add umbrella-devel smoke suite and canary coverage
Add nightly smoke tests and installer support for the umbrella-devel Ceph image.
Also, refactored the test to use reusable code, and since we run the tests in
k8s, removed the k8s matrix from the input.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-07-13 16:28:32 +05:30
Joshua Hoblitt 3c0be0d67a test: build the smoke object store with the shared store fixture
Generalize the object suite's sharedstore.Create to accept the target
namespace, store name, and RGW instance count instead of the hardcoded
object-ns / sharedstore / single instance, then use it to stand up the
SmokeSuite object store.

TestObjectStorage_SmokeTest previously created its store with
runObjectE2ETestLite, which applies a CephObjectStore manifest and polls
for health on fixed intervals. It now builds the store through the
typed-client, watch-based sharedstore fixture and tears it down with
Destroy. lite-store is the only object store in the smoke namespace, so
the fixture safely owns the cluster-global .rgw.root realm pool that
Destroy deletes.

The object suite's own call is updated to pass its namespace,
"sharedstore", and one instance, preserving its existing behavior.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-10 14:39:29 -07:00
Joshua Hoblitt 8a72d80127 test: exercise multisite chart values in CephHelmSuite
Configure the CephHelmSuite object store as an RGW multisite zone so the
chart's CephObjectRealm, CephObjectZoneGroup, and CephObjectZone templates
are exercised on a Helm install. The operator must reconcile the realm,
zone group, zone, and store to Ready, which the existing install check
already gates on via the RGW pod count.

The realm/zone group/zone and shared-pool layout mirror the fixture in
tests/integration/object/util/sharedstore: a zone backed by explicit
CephBlockPools referenced through sharedPools.poolPlacements. The rgw pools
are appended to the block pools already configured for CSI, and are torn
down along with the multisite CRs when the default storage CRs are removed.

Gate the multisite configuration behind a new UseMultisiteObjectStore
setting enabled only for CephHelmSuite. The other Helm-based suites keep the
plain object store; the upgrade suite in particular installs an older base
chart that lacks the multisite templates and cannot resolve a zone-based
store.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-09 09:59:18 -07:00
Michael Adam 031830d418 tests: add make lint.markdown-links
This adds a make target lint.markdown-links
for checking forbroken links in the markdown
documentation sources.

Thid is intended to be used in the CI as well as
locally by developers

The link checking is implemented using the markdown-link-check tool
but it is self-contained in that it does not require the
tool to be installed on the system.
The only prerequisite for using this target is to have
a working docker or podman available.

Assisted-by GitHub copilot
Assisted-by: Google Antigravity/Gemini
Assisted-by: IBM Bob
Signed-off-by: Michael Adam <obnox@samba.org>
2026-07-09 17:59:15 +02:00
Joshua Hoblitt 2cf5008bc8 test: convert the object bucket policy suite to the shared store
Add tests/integration/object/bucket/policy, converting the bucket policy
slice of the legacy testObjectStoreOperations onto the shared store: an OBC
created with a bucketPolicy additionalConfig, the policy verified verbatim
on the rgw bucket over the S3 API, then updated and removed with the bucket
polled until it reflects each change. The policy JSON is generated from the
bucket name, which the legacy test hardcoded in the Resource ARN. It runs
against the shared store in both the TLS and non-TLS passes.

Drop the converted policy block from the legacy test.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-07 15:41:48 -07:00
Joshua Hoblitt 6c2a3a5f2b Merge pull request #17881 from jhoblitt/convert-bucket-quota-to-sharedstore
test: convert the object bucket quota suite to the shared store
2026-07-07 15:35:29 -07:00
Travis Nielsen 5019b8e455 Merge pull request #17875 from subhamkrai/update-csi-operator-1.0.3
csi: update csi-operator version to v1.0.4
2026-07-07 10:10:45 -06:00
Joshua Hoblitt 0a7d64bc69 Merge pull request #17877 from jhoblitt/ci/fix-find-block-dev-broken-pipe
ci: fix broken-pipe race in find_extra_block_dev
2026-07-07 08:38:39 -07:00
subhamkrai a3e02e607d csi: update csi-operator version to v1.0.4
Updating csi-operator to latest v1.0.4 and
updating the required doc changes as well.

Signed-off-by: subhamkrai <srai@redhat.com>
2026-07-07 07:56:55 -06:00
Joshua Hoblitt cf54270bf5 Merge pull request #17822 from jhoblitt/ci-minikube-to-kind
ci: convert integration and canary CI from minikube to kind
2026-07-06 12:46:09 -07:00
Joshua Hoblitt dad0e21122 test: retire the unused object bucket framework helpers
Converting the bucket quota tests to typed clients left the framework
helpers and package globals they were the last consumers of unused. Remove
UpdateObc, CheckOBMaxObject, GetAccessKey, and GetSecretKey from the bucket
client, GetEndPointUrl from the object client, and the orphaned object-test
globals.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-06 10:35:37 -07:00
Joshua Hoblitt 84c10661da test: convert the object bucket quota suite to the shared store
Add tests/integration/object/bucket/quota, covering both quota behaviors
from the legacy testObjectStoreOperations on OBCs of its own: user quota
via the OBC maxObjects additionalConfig (enforced, then raised and
re-enforced, synced on the ObjectBucket reflecting the new limit) and
bucket quota via bucketMaxObjects switched to bucketMaxSize. It runs
against the shared store in both the TLS and non-TLS passes.

Drop the converted S3-access and bucket-quota blocks from the legacy test;
the main OBC now exists only as a dependents fixture for the store
deletion checks.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-06 10:35:37 -07:00
Joshua Hoblitt 8dfc5d408d test: add an OBC update helper to util/obc
Add obc.Update, a get/mutate/update helper for live ObjectBucketClaims,
generalized from the notification suite's label updater, which becomes a
thin wrapper over it. The bucket quota conversion needs the same shape for
additionalConfig updates.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-06 10:35:37 -07:00
Joshua Hoblitt d59abb1956 ci: convert integration and canary CI from minikube to kind
The integration and canary suites ran on a single-node minikube
`driver: none` cluster, where the kubelet ran directly on the GitHub
runner, so the host docker daemon doubled as the cluster runtime and
host block devices and host paths were directly visible to pods. Replace
that with a kind cluster.

Every suite creates its cluster through the shared
integration-test-setup-cluster-resources composite action, so the
conversion is centralized there and converts the smoke, object, helm,
keystone, multi-cluster, upgrade, on-release, nightly, encryption-KMS and
all canary jobs at once.

- Replace the setup-minikube step with helm/kind-action, selecting the
  kubernetes version via the kindest/node image tag and creating a
  single-node cluster from a new kind config (kind pinned to v0.32.0 for
  reproducibility).
- Drop the cri-dockerd install; kind nodes use their built-in containerd.
- Add a kind config that bind-mounts the host /dev, /var/lib/rook and
  /run/udev into the node so the existing host-based disk-prep helpers
  (use_local_disk*, create_partitions_for_osds, blockDevicePV.sh,
  localPathPV.sh, ...) keep working unchanged: devices and partitions
  created on the host appear in the node and in the OSD pods that
  hostPath-mount the node /dev, and ceph-volume can read the host udev
  database it needs to inventory disks.
- Prepare the kind node for the host-level operations rook runs against the
  underlying host: remount /sys read-write so CSI's kernel RBD mapping
  (`rbd map --device-type krbd`, which writes /sys/bus/rbd) works, and install
  lvm2 and cryptsetup, which rook runs in the node's mount namespace to
  provision LVM- and encryption-backed OSDs. kindest/node images provide none
  of this; the minikube driver:none runner host did.
- Route the Service and pod CIDRs from the runner to the kind node so
  host-side tests (the `go test` process runs on the runner) can reach
  in-cluster ClusterIPs, e.g. an S3 request to the RGW service. With minikube
  driver:none the runner already shared the cluster network.
- Load locally built images into the cluster. Under minikube `driver: none`
  the built image was already in the cluster runtime; under kind it must be
  imported, so build_rook and create_helm_tag now import their images into
  each node's containerd through a new load_image_into_cluster helper (via the
  node's ctr, which avoids the kind/kindest-node containerd-config version
  skew that breaks `kind load docker-image`).
- Point Vault's kubernetes-auth at the in-cluster API endpoint
  (kubernetes.default.svc) instead of the kubeconfig server URL: kind exposes
  that as https://127.0.0.1:<port>, unreachable from the in-cluster Vault pod,
  so OSD encryption-key retrieval via k8s-auth failed.
- Replace the remaining direct minikube references in the canary workflow: a
  `minikube kubectl` call and the external-cluster topology values.
- Adapt host-name assumptions that only held under driver:none: resolve the
  disk-cleanup job by the k8s node name rather than the runner hostname, and
  let kind-action ignore post-job cluster-teardown failures (the runner is
  ephemeral; nvme/multus devices can wedge `docker rm` of the node).
- Update stale comments that described the CI environment as minikube.
- Move the multus integration test's kind config under tests/config too, so
  both kind cluster configs live in one place.

create-dev-cluster.sh and other local-dev tooling are intentionally left on
minikube.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-06 08:42:39 -07:00
Joshua Hoblitt 064e37ae8d test: reconcile the object README with the merged util packages
The README still described the util layout from before the package
consolidation, with ready, admin, sns, s3, and tls as standalone packages
and a code sample calling ready.ObjectStoreUser. The readiness predicates
live in wait4 and the client builders in client; update the package list,
the code sample, and the playbook references to match.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-02 08:41:59 -07:00
Joshua Hoblitt b4d6335368 test: convert the object bucket read/write suite to the shared store
Add tests/integration/object/bucket/rw, the first slice carved from the
legacy testObjectStoreOperations. It provisions an OBC on the shared store,
writes, reads back, and deletes an object over S3, and checks the OBC stays
Bound. It runs against the shared store in both the TLS and non-TLS passes.

Drop the converted put/get and OBC-revert subtests from the legacy test;
the user quota subtest now seeds both of the objects it needs. The
remaining quota, policy, lifecycle, and dependents slices convert in later
PRs.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-02 08:41:59 -07:00
Joshua Hoblitt 2f82880de7 test: consolidate the OBC test helpers into util/obc
Introduce util/obc as the home for ObjectBucketClaim test helpers: the
provisioner StorageClass constructor (moved out of util/fixture, where it
was a misplaced pure constructor), the create/bound and delete/absent
lifecycle waiters, and a per-OBC S3 client. The lifecycle waiters and S3
client are promoted from the notification suite, which built them inline as
their only consumer.

Retarget the StorageClass callers (topic/kafka, bucket/owner) and rework
the notification suite onto the shared helpers.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 16:26:46 -07:00
Joshua Hoblitt b7725553d0 test: add a generic S3 agent builder to the object test client
Add client.NewS3Agent, which builds an rgw S3 agent for an object store
from a set of credentials and the store's endpoint. The object tests build
S3 clients from several credential sources (the store's dashboard admin
user, an OBC's provisioned secret); this factors out the endpoint and agent
construction they share.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 14:42:39 -07:00
Joshua Hoblitt c4ac54cea0 Merge pull request #17853 from jhoblitt/convert-bucket-notification-test
test: convert the bucket notification suite to the shared store
2026-07-01 09:22:44 -07:00
Joshua Hoblitt 7079c9bf62 ci: fix broken-pipe race in find_extra_block_dev
find_extra_block_dev piped find_extra_block_devs into head -1. Since
find_extra_block_devs prints every eligible disk and ends with echo
"$devs", head closes the pipe after the first line and the producer's
echo races into SIGPIPE (echo: write error: Broken pipe), failing the
command substitution under pipefail/errexit.

The race only fires reliably now that the runners present two data disks
(sdb and sdc), so find_extra_block_devs emits multiple lines; with a
single disk the one-line output almost never triggered it.

Read the producer to completion with mapfile via process substitution
and return the first element, mirroring find_second_block_dev, so there
is no downstream head to close the pipe early.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 09:17:20 -07:00
Joshua Hoblitt c0eb360041 docs: fix typos and grammar in code comments
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:34:11 -07:00
Joshua Hoblitt 031e449070 test: retire the unused bucket notification framework helpers
Converting the bucket notification test to typed clients left the
framework notification plumbing it was the last consumer of unused.
Remove the NotificationOperation client, TopicClient.CreateHTTPServer,
BucketClient.CreateObcNotification and CheckBucketNotificationSetonRGW
(and the now-dead UpdateObcNotification helpers), and the
GetBucketNotification and GetOBCNotification manifest builders.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:30:22 -07:00
Joshua Hoblitt 780f0cceb0 test: convert the bucket notification suite to the shared store
Rework the bucket notification integration test onto the shared
CephObjectStore fixture and the wait4 toolkit, following the object
suite conversion playbook. The old test stood up its own object store
and drove everything through tests/framework/clients and kubectl with
fixed-interval polling.

The new tests/integration/object/notification package is an ordered
t.Run script on typed clients and watch-based waits. It deploys an HTTP
sink, drives a CephBucketTopic (HTTP endpoint), CephBucketNotification,
and a notification-labelled OBC, and verifies both end-to-end delivery
(matched in the sink's logs) and the notification configured on the rgw
bucket (read over the S3 API), covering the label add/remove and
notification-before-topic orderings.

Running against the shared store, it no longer creates a dedicated
object store and now exercises both the TLS and non-TLS suite passes.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:30:22 -07:00
Anas Khan 67bf4baab8 object: use passed context for S3 v2 API calls
The AWS SDK v1 to v2 migration in #17468 left the S3Agent wrapper methods
calling the SDK with context.TODO(), so S3 operations could not be cancelled
when the controller context is cancelled. Add a leading context.Context
parameter to CreateBucket, PutObjectInBucket, GetObjectInBucket,
DeleteObjectInBucket, PutBucketPolicy, and GetBucketPolicy, forward it to the
underlying client, and pass clusterInfo.Context from the bucket provisioner.

Resolves #17526

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
2026-07-01 13:15:40 +05:30
Joshua Hoblitt e281600bb5 Merge pull request #17855 from jhoblitt/narrow-bare-nolint
core: narrow bare //nolint directives to specific linters
2026-06-30 08:14:46 -07:00
Joshua Hoblitt 5065ea7973 Merge pull request #17829 from jhoblitt/ci-provision-disks-no-extra
ci: provision extra disks when the runner starts with none
2026-06-30 08:13:13 -07:00
Joshua Hoblitt e5c6d75a73 core: narrow bare //nolint directives to specific linters
The //nolint directives at these sites omit the linter name, so each
suppresses every linter on its line rather than the one check it needs.
That hides any unrelated errcheck/gosec/govet finding later introduced
on the same line. Name the specific linter for each:

- staticcheck for the two operator sites: SA4004 (the intentional
  single-iteration loop in the OSD PVC host lookup) and SA1019 (the
  deliberate read of the deprecated S3.Enabled field in the RGW
  API-enable builder).
- errcheck for the rbd-mirror deferred token-file cleanup and the
  test-framework logging helpers (WriteString / writeHeader).

No behavior change.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 16:32:59 -07:00
Joshua Hoblitt 3a4dc64f5b test: convert the COSI driver suite to the object test toolkit
Rework the COSI driver integration test onto the shared-store object test
toolkit: an ordered t.Run script with the (t, k8sh, store) signature, typed
clients, and watch-based waits, replacing kubectl-string manifests and
fixed-interval polling.

Drive the COSI bucket resources (BucketClass, BucketClaim, Bucket) with the
upstream sigs.k8s.io/container-object-storage-interface typed client, exposed
as k8sh.COSIClientset alongside the existing OBC client. Create the
CephCOSIDriver and its privileged user through typed clients, wait for the
driver Deployment with wait4, and verify the provisioned bucket through the
shared store's rgw admin client.

Install the COSI CRDs and central controller from the consolidated upstream
repo pinned to v0.2.2 (the former -api and -controller repos are retired) via
a kubectl -k fixture that is removed with t.Cleanup. The driver cannot trust
a TLS RGW endpoint, so the suite skips itself in the TLS pass rather than
being special-cased in the dispatcher.

Retire the now-unused COSIOperation client and the GetCOSIDriver,
GetBucketClass, and GetBucketClaim manifest helpers.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 15:36:13 -07:00
Joshua Hoblitt 802c4e0e6e Merge pull request #17663 from jhoblitt/maint-integration-sharedstore-improvements
test: improve the shared-store object suite tests
2026-06-29 13:33:28 -07:00
Joshua Hoblitt 8ee5d92022 test: document the object suite test conventions
Add a conventions doc and conversion playbook so the remaining old-style
object tests (store lifecycle, notifications, COSI, bucket rw/quota/policy/
lifecycle) can be converted to this pattern mechanically.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 11:10:52 -07:00
Joshua Hoblitt eb404e6677 test: rework the shared-store user/bucket/topic suites
The object store user (caps/keys/opmask), bucket-owner, and bucket topic
tests already shared a CephObjectStore but polled with count-based
utils.Retry behind a wide entry signature. Move them onto the wait4 toolkit
and a slim (t, k8sh, store) signature: each is an ordered t.Run script using
typed clients and watch-based waits, with per-package check helpers for
repeated verification. Extend the shared store fixture to provision the
realm/zonegroup/zone pools and serve TLS, build its rgw admin and SNS clients
through the consolidated util/client package (replacing the separate
admin/sns/s3 packages), and run every package in both the TLS and non-TLS
passes (now separate parallel jobs).

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 11:10:52 -07:00
Joshua Hoblitt b59101d8e8 test: add the object suite test util toolkit
Add the shared helpers the object test packages build on:
- wait4: watch-based Assert/Require waiters (Create/Delete/Condition/Absent)
  that block on Kubernetes watch events rather than fixed-interval polling,
  plus Eventually for non-k8s state (rgw admin/S3/SNS), a pod-log waiter, and
  the readiness predicates used with them.
- secrets: verification of the Secret references CRDs publish in status.
- fixture: create-with-t.Cleanup helpers for pure-cleanup resources, and
  constructors for ObjectBucketClaim test resources.
- client: rgw admin and SNS client builders, S3 credentials/endpoint, and the
  store's TLS cert plus a verification-skipping HTTP client.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-29 11:10:52 -07:00
Joshua Hoblitt 029b13b000 ci: provision extra disks when the runner starts with none
find_extra_block_devs reads the runner's current extra disks with
`lsblk ... | egrep -v "($boot_dev|loop|nbd)"` to decide how many more to
provision. When the runner has no extra disk, that egrep matches nothing
and exits 1; under the script's `set -eo pipefail` the assignment aborts
the function before it ever calls create_extra_disk, so callers get zero
disks and fail with "expected >= N extra disks, found 0" (or a bare /dev/
path) on those runners.

Guard that first read with `|| true` (as the adjacent device-count line
already is) so an empty result means "no extra disks yet" and the
provisioning path runs. The second read, after create_extra_disk, is left
strict on purpose: if provisioning produced no disk, egrep's non-zero exit
should abort there rather than silently return nothing.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-26 15:00:37 -07:00
Joshua Hoblitt 49612461a4 docs: fix function comments to match their declaration names
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-26 12:52:39 -07:00
Joshua Hoblitt 04734d7150 ci: capture build_rook stderr so the network-retry can see it
build_rook() captures the make output into $o and inspects it for
transient network errors (connection reset, INTERNAL_ERROR, 503, 500)
to decide whether to retry the build. The capture was stdout-only,
but go and make write those errors to stderr, so $o never matched any
retry case. A transient module-proxy failure (e.g. proxy.golang.org
returning INTERNAL_ERROR while downloading modules) therefore fell
through to the "valid failure" branch and failed the job in the build
step, before the integration test even ran.

Redirect stderr into the captured output (2>&1) so the existing retry
logic actually triggers on network failures.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-06-25 12:36:33 -07:00
Travis Nielsen 01d49c2763 Merge pull request #17821 from daixihegu/master
core: use slices.Contains to simplify code
2026-06-25 10:40:35 -06:00
Joshua Hoblitt c1594a82ea Merge pull request #17801 from jhoblitt/ci-loop-to-iscsi
ci: replace loop devices with iscsi in canary osd tests
2026-06-24 22:25:16 -07:00