if the account cr has a force delete annotation, forcefully
remove the account, even if it contatins data in it, using the
purge data flag
Signed-off-by: parth-gr <partharora1010@gmail.com>
retrieveMultisiteZone is meant to gate the object-store reconcile on the
backing Ceph zone existing: it runs "radosgw-admin zone get" and, when
that fails, returns a non-nil error so the caller requeues. That gate is
dead code. The ENOENT check declares an inner err from exec.ExtractExitCode
that shadows the outer command error, and ExtractExitCode returns a nil
error for the ordinary exit failures radosgw-admin produces. Both the
ENOENT branch and the else branch then wrap that shadowed nil, and
errors.Wrapf(nil, ...) is nil, so the function returns
(waitForRequeueIfObjectStoreNotReady, nil). The caller only propagates the
requeue when the error is non-nil, so the requeue is dropped and reconcile
runs on.
The result is that a failed "zone get" no longer backs off. Reconcile
proceeds to stand up the object store anyway -- the RGW service, the
admin-ops endpoint, the deployment, and the pools radosgw scaffolds as it
comes up -- for a multisite store whose backing zone does not exist. This
is not the "normal multisite bootstrap" transient the original wording
suggested. getMultisiteResourceNames runs immediately before this and
already requeues until the CephObjectZone CR reports Ready, and the zone
controller marks it Ready only after it has created the Ceph zone, so a
healthy bootstrap never reaches this gate with a missing zone. That
CR-Ready gate, not this one, is what actually blocks bootstrap.
Where the dead gate does bite is the cases the CR status cannot cover:
- the Ceph zone deleted or renamed out of band while the CR still reads Ready
- a zone controller that reports Ready without leaving a usable zone behind
- any non-ENOENT "zone get" failure, e.g. a permission or connectivity
error, which the else branch swallows the same way
The check was correct when fc579f4520 introduced it in 2020 with
exec.ExitStatus, which returns (code, ok) and leaves err unshadowed.
bd58790c31 ("ceph: proxy ceph commands when multus is configured", 2021)
swapped it to exec.ExtractExitCode as an unrelated drive-by, inverting the
contract and killing the gate. The sibling realm, zonegroup, and zone
controllers were left untouched and still use exec.ExitStatus today.
Restore exec.ExitStatus so the outer error is no longer shadowed, both
branches wrap the real command error, and the caller requeues until the
zone exists -- matching the sibling controllers.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
In createOrUpdateCephUser, when the desired user carries no explicit
keys the reconciler falls back to the live RGW user's keys. If that live
user also has zero keys the code intends to fail, but it built the error
with errors.Wrapf(err, ...) at a point where err is already nil (the
prior SetUserQuota error was handled and returned just above).
errors.Wrapf(nil, ...) returns nil, so the failure was swallowed and
createOrUpdateCephUser returned success on a user the operator itself
flagged as broken.
The reconcile then continued to generateCephUserSecret, which indexes
userConfig.Keys[0] to populate the Kubernetes secret and panicked on the
empty key slice. The reconciler's deferred RecoverAndLogException caught
and logged that panic, so the reconcile was abandoned before it reached
the Ready status update; and because the recovered Reconcile returns a
zero Result with a nil error, the request was not requeued either. The
user was left neither marked Ready nor retried.
Construct the error with errors.Errorf so the intended failure is
surfaced instead of being swallowed. Add a regression test that returns
a keyless live user and asserts a non-nil "no keys set" error.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
this commit add support for updating rgw caps with
`user-info-without-key` and `accounts` to cephobjectstoreuser crd
Adding the check for min ceph version from which these caps
are available.
Co-Authoured-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: subhamkrai <srai@redhat.com>
currently there was a bug in the code where it didnt removed
the account from the ceph cluster during intial intialize
Signed-off-by: parth-gr <partharora1010@gmail.com>
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The AWS SDK v1 to v2 migration in #17468 left the S3Agent wrapper methods
calling the SDK with context.TODO(), so S3 operations could not be cancelled
when the controller context is cancelled. Add a leading context.Context
parameter to CreateBucket, PutObjectInBucket, GetObjectInBucket,
DeleteObjectInBucket, PutBucketPolicy, and GetBucketPolicy, forward it to the
underlying client, and pass clusterInfo.Context from the bucket provisioner.
Resolves#17526
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
The //nolint directives at these sites omit the linter name, so each
suppresses every linter on its line rather than the one check it needs.
That hides any unrelated errcheck/gosec/govet finding later introduced
on the same line. Name the specific linter for each:
- staticcheck for the two operator sites: SA4004 (the intentional
single-iteration loop in the OSD PVC host lookup) and SA1019 (the
deliberate read of the deprecated S3.Enabled field in the RGW
API-enable builder).
- errcheck for the rbd-mirror deferred token-file cleanup and the
test-framework logging helpers (WriteString / writeHeader).
No behavior change.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The realm system user's access and secret keys are generated by
GeneratePassword() and then wrapped in base64. GeneratePassword()
deliberately excludes '/' from the access key character set, but the
base64.StdEncoding wrap reintroduces it: its alphabet contains '/' and
'+', and the 14-character input always produces trailing '='. The
encoded string is the literal key used in S3 requests.
An access key containing '/' breaks AWS SigV4 credential scope parsing
("<access-key>/<date>/<region>/<service>/aws4_request" is split on
'/'), so "radosgw-admin realm pull" against the realm endpoint fails
permanently with "request failed: (22) Invalid argument" (HTTP 400),
and a CephObjectRealm pulling that realm can never reconcile.
This is the dominant cause of the "deploy second cluster rook"
failures in the rgw-multisite-testing canary job: every sampled failure
had a generated access key containing '/' and looped on EINVAL for the
whole 600s wait window, while runs with slash-free keys pulled the
realm successfully.
Encode both keys with base64.RawURLEncoding instead, whose alphabet
(A-Za-z0-9-_, unpadded) is safe in credential scopes, URLs, and shell
arguments. Only newly created realm secrets are affected; existing
secrets are not modified by the reconciler.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Remove the following unreferenced symbols:
- (*S3Agent).CreateBucketNoInfoLogging
- (*S3Agent).DeleteBucket (wrapper; callers use the SDK client directly)
- GetBucketsStats
- ObjectBuckets type and its Len/Less/Swap sort.Interface methods
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Remove (*BucketPolicy).EjectPrincipals and (*PolicyStatement).EjectPrincipals,
which had no callers outside policy.go. Also drop the duplicate PutBucketVersioning
entry in AllowedActions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
ModifyBucketPolicy merged the caller's statement into the policy fetched
from the bucket, matching by SID. Besides a missing match flag that
appended a duplicate when a SID matched, this preserved any pre-existing
statements on the bucket and reapplied them on every reconcile.
Overwrite the policy with the provided statements instead, so a managed
bucket always ends up with exactly the intended policy and cannot retain
unexpected statements.
Signed-off-by: Artem Muterko <artem@sopho.tech>
Wire a rookLogger adapter implementing smithy's Logger
interface so that aws.LogSigning output is emitted through
rook's capnslog when debug mode is enabled.
Signed-off-by: Oded Viner <oviner@redhat.com>
Remove the AWS SDK v1 (github.com/aws/aws-sdk-go)
dependency entirely. All S3 operations now use AWS SDK v2
exclusively.
- Remove the v1 Client field from S3Agent struct and
rename ClientV2 to Client
- Remove v1 session/client initialization from NewS3Agent
- Update all call sites referencing ClientV2
- Convert integration tests to use v2 API calling
conventions (context parameter) and smithy error handling
- Remove aws-sdk-go v1.55.8 from go.mod
Signed-off-by: Oded Viner <oviner@redhat.com>
adding new tls ssl_ciphersuites supporting tls 1.3 and the existing,
ssl_cipher supports tls 1.2 and below. Adding, the docs and unit-test
changs as well.
Signed-off-by: subhamkrai <srai@redhat.com>
migrate the sns-based topic provisioner from aws-sdk-go (v1)
to aws-sdk-go-v2. replace v1 session, credentials, and
request handlers with v2 aws.Config, static credentials
provider, and endpoint resolver. convert the custom ceph rgw
v2 signer hack from a v1 handler swap to a smithy finalize
middleware. update sns api calls to use context-first
signatures, map[string]string attributes, and typed
snstypes.NotFoundException error handling via errors.As.
update unit tests accordingly.
Signed-off-by: Oded Viner <oviner@redhat.com>
migrate the notification provisioner and s3ext packages from
aws sdk v1 to v2. the custom DeleteBucketNotification call is
rewritten to manually build and sign the HTTP request using the
v2 v4 signer, since this ceph-specific API has no sdk equivalent.
also fix t.Skipped() -> t.Skip() in notification integration test.
Signed-off-by: Oded Viner <oviner@redhat.com>
this commit add option in the ceph objecstore CR
to configure TLS profile and TLS ciphersuite for
rgw beast.
Signed-off-by: subhamkrai <srai@redhat.com>
migrate bucket provisioner from aws sdk v1 to v2.
replace awserr.Error type assertions with smithy.APIError
using errors.As for v2-style error handling.
switch setBucketPolicy and setBucketLifecycle to use
the v2 s3 client (ClientV2) directly with context.
replace v1 s3 types (BucketLifecycleConfiguration,
GetBucketLifecycleConfigurationOutput) with v2 equivalents
from s3types and s3v2 packages.
add cmpopts.IgnoreUnexported to cmp.Diff calls to handle
unexported noSmithyDocumentSerde fields in v2 sdk types.
add nil guard on lifecycle output before accessing rules
to prevent nil dereference when get returns an error.
depends on #17414
Signed-off-by: Oded Viner <oviner@redhat.com>
Support using SSE-S3 encryption with RGW using vault Agent auth.
RGW sends requests to the agent instead of directly to Vault, and the agent transparently injects the authentication token. This eliminates using and managing a static token
Signed-off-by: Santosh <sapillai@redhat.com>
This prevents overwriting the user-specified capabilities
and correctly copies the Capabilities into UserCapabilities
after the AdminOpsClient calls.
Signed-off-by: hjk068 <hello.hyunjin@gmail.com>
initialize both aws sdk v1 and v2 clients in s3agent.
update createbucket to use sdk v2 while keeping all
other methods on sdk v1.
Signed-off-by: Oded Viner <oviner@redhat.com>
Allow users to configure custom labels on CephObjectStore RGW services
by adding a Labels field to the RGWServiceSpec, following the existing
pattern used for Annotations. This enables use cases such as service
mesh integration and monitoring discovery that require specific labels
on Kubernetes services.
Fixes: #17235
Signed-off-by: majiayu000 <1835304752@qq.com>