Commit Graph
156 Commits
Author SHA1 Message Date
Anas Khan 0bf6f84ff5 object: return an error when the multisite zone is not found
retrieveMultisiteZone is meant to gate the object-store reconcile on the
backing Ceph zone existing: it runs "radosgw-admin zone get" and, when
that fails, returns a non-nil error so the caller requeues. That gate is
dead code. The ENOENT check declares an inner err from exec.ExtractExitCode
that shadows the outer command error, and ExtractExitCode returns a nil
error for the ordinary exit failures radosgw-admin produces. Both the
ENOENT branch and the else branch then wrap that shadowed nil, and
errors.Wrapf(nil, ...) is nil, so the function returns
(waitForRequeueIfObjectStoreNotReady, nil). The caller only propagates the
requeue when the error is non-nil, so the requeue is dropped and reconcile
runs on.

The result is that a failed "zone get" no longer backs off. Reconcile
proceeds to stand up the object store anyway -- the RGW service, the
admin-ops endpoint, the deployment, and the pools radosgw scaffolds as it
comes up -- for a multisite store whose backing zone does not exist. This
is not the "normal multisite bootstrap" transient the original wording
suggested. getMultisiteResourceNames runs immediately before this and
already requeues until the CephObjectZone CR reports Ready, and the zone
controller marks it Ready only after it has created the Ceph zone, so a
healthy bootstrap never reaches this gate with a missing zone. That
CR-Ready gate, not this one, is what actually blocks bootstrap.

Where the dead gate does bite is the cases the CR status cannot cover:

- the Ceph zone deleted or renamed out of band while the CR still reads Ready
- a zone controller that reports Ready without leaving a usable zone behind
- any non-ENOENT "zone get" failure, e.g. a permission or connectivity
  error, which the else branch swallows the same way

The check was correct when fc579f4520 introduced it in 2020 with
exec.ExitStatus, which returns (code, ok) and leaves err unshadowed.
bd58790c31 ("ceph: proxy ceph commands when multus is configured", 2021)
swapped it to exec.ExtractExitCode as an unrelated drive-by, inverting the
contract and killing the gate. The sibling realm, zonegroup, and zone
controllers were left untouched and still use exec.ExitStatus today.

Restore exec.ExitStatus so the outer error is no longer shadowed, both
branches wrap the real command error, and the caller requeues until the
zone exists -- matching the sibling controllers.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-21 12:40:25 -07:00
Joshua Hoblitt c0eb360041 docs: fix typos and grammar in code comments
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-01 08:34:11 -07:00
Artem Torubarov a7f176b366 rgw: skip non-raw commands on object deletion path when pools not exists
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-02-26 09:48:15 +01:00
Artem Torubarov b47d2c2f79 rgw: force objectStore controller to wait until zone and sharedPools are ready
Fixes race condition for rgw multisite deployments caused by ObjectStore
and Zone controllers calling period commit concurrently. See: https://github.com/rook/rook/issues/17013

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-02-24 11:43:57 +01:00
subhamkrai 08344f39f6 build: update deprecated api GetEventRecorderFor
with controller runtime latest version vO.23.0, golangci lint
is complaining about deprecated api deprecated.
This commit updates the api and adds the requried rbacs

Signed-off-by: subhamkrai <srai@redhat.com>
2026-02-09 21:48:02 +05:30
Travis Nielsen 023608e6fd core: enhance logging with namespaced names
For all of the controllers besides the cluster controller,
the logging now includes the namespaced name of the resource
that is being reconciled. This will help with log troubleshooting
to help analyze logs consistently for the resource being
reconciled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-12-04 12:00:42 -07:00
Travis Nielsen ac1fa3a4c3 core: shorten the logger names
The logger names are seen on each line of logging. As
long as they are unique, they really don't need to be
long and descriptive, so let's make them a bit more
concise, and independent from controller names.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-11-21 13:12:56 -07:00
SantoshPillai 67b44b47df rgw: make hpa update rgw resource
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2025-09-19 12:22:30 +05:30
cuiweixie 20b3b9deef operator: refactor to use reflect.TypeFor
Signed-off-by: cuiweixie <cuiweixie@gmail.com>
2025-08-28 23:02:19 +08:00
Oded Viner 182cd211ca object: mark realm as default to avoid zone creation
add defaultRealm field to cephobjectstore to allow rook to mark the
created realm as default. this prevents ceph from auto-creating a
"default" zonegroup and zone when the default is not explicitly set.
affects only non-multisite setups.

Signed-off-by: Oded Viner <oviner@redhat.com>
2025-08-07 14:48:21 +03:00
Oded Viner cf13deee6f core: log panics in controller reconcile functions
add RecoverAndLogException() helper to log panics with stack trace.
added defer call in all rook controller Reconcile() methods for
better error visibility in operator logs

Signed-off-by: Oded Viner <oviner@redhat.com>
2025-07-29 13:20:38 +03:00
parth-gr b8517eb54f rgw: fail reconcile on status update error
if the status update is failed, return an error
and do a new reconcile for the rgw controller

Signed-off-by: parth-gr <partharora1010@gmail.com>
2025-07-16 14:37:03 +05:30
Blaine Gardner 5f50e85a25 object: implement cephx key rotation for rgw keys
This is the first implementation of CephX key rotation in Rook.
Adds key rotation API from design PR 15915.

Also add helper methods for rotating keys, determining when keys
need to be rotated, and for generating CephX key statuses. These will be
reusable for other Rook reconciles beyond object/RGW.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2025-06-10 09:06:58 -06:00
Travis Nielsen 95a5920bef object: log all reconcile errors during object store creation
The object store controller was swallowing errors of the
type not found, assuming a certain category of errors
when ceph may still be initializing. But some of these
errors could be user error and they need to troubleshoot
the details of the error, so we need to log the specific
error instead of swallowing it.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-04-18 11:35:09 -06:00
Travis Nielsen f7fb1bc0f2 core: skip reconcile when adding the finalizers
When finalizers are added to the CRs, a follow-up reconcile
will be triggered due to the increased generation on the CR.
Therefore, abort the initial reconcile when adding the finalizer,
and allow the follow-up reconcile to complete the configuration.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-04-08 13:42:58 -06:00
Joshua Hoblitt 9f1ed201db core: typed watch handlers and predicates
All existing controller runtime watches are converted to use "typed"
handlers and predicates instead of operating on `client.Object`.  The
intent is to be bug for bug equivalent with the existing logic while
replacing run time type assertions and switch statements with compile
time type constraints and type casts. In several cases, functions using
assertions were split up such that each function only handles a single
Kind at a time. It is hoped that this will improve readability and
maintainability while facilitating future refactoring such as migrating
some watches to using IndexFields.

Of particular note is that the massive switch statement in
`WatchControllerPredicate()`  from
`pkg/operator/ceph/controller/predicate.go` has been replaced with
generics, reflection, and splitting the obc logic into its own predicate
function. There are still many helper functions operating on
`client.Object`. These were not updated unless required by the compiler
in order to limit the size of this change. The type safety of these
funcs should be improved as followup work.

 It is strongly suggested that going forward, handlers and predicates
 only handle a single Kind (generic or not) and that switches / type
 assertions are heavily discouraged or forbidden. This PR removed all
 but a single switch statement in a predicate, which should be addressed
 in future work.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2025-04-04 09:40:56 -07:00
Joshua Hoblitt 3cb343f62a core: run gofumpt on all files
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2025-03-26 10:41:48 -07:00
Artem Torubarov e2f7caf193 rgw: simplify map secret func
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2025-03-21 16:11:34 +01:00
Artem Torubarov dd0d500980 rgw: watch referenced secrets
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2025-03-17 14:09:05 +01:00
Artem Torubarov 00aa282074 rgw: use list in secret watch map func
Signed-off-by: Artem Torubarov <artem.torubarov@clyso.com>
2025-03-13 11:19:45 +01:00
Artem Torubarov bed0c23274 rgw: watch config secret changes
Implements #15460. If ObjectStore reading value from a secret
provided by user, the secret will be annotated with ObjectStore
name and ObjectStore will be reconciled on secret change.

Signed-off-by: Artem Torubarov <artem.torubarov@clyso.com>
2025-03-12 11:37:12 +01:00
Blaine Gardner fc08e87d44 Revert "object: create cosi user for each object store"
This reverts commit a941b3c33f.

Stop creating the 'cosi' user in the CephObjectStore reconcile. This
step often fails for some amount of time during initial object store
creation, causing frequent user concern. It has also been the source of
some reported failures that would otherwise be non-breaking for certain
users.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-11-21 16:03:32 -07:00
Artem Torubarov 59175f0b40 rgw: pool placement
Signed-off-by: Artem Torubarov <torubarov.a.a@gmail.com>
2024-09-06 16:02:53 +02:00
ee8bcad49d rgw: add support for keystone auth + swift/s3
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>

* extend the API object specs for swift and keystone integration

* adapt rgw to the new go-ceph version

  - The parameter lists of the API call have changes, as parameters
    ignored by the RGW Admin Ops API are no longer serialized, therefore
    the mock has to be adapted.

  - There is now validation for the user keys that are passed to the
    User get API, therefore things failed when we had empty keys in our
    User proxy object.

* expand the reconcile loop for the swift and keystone integration

* fix minor mistakes in design document

* add env var to pass extra args to minikube

  Minikube decides CPU cores and memory automatically based on the
  available resources on the machine which may be insufficient to
  run rook. This commit adds an environment variable to add arbitrary
  arguments to the minikube command, so both can be specified if
  desired.

* integration tests for swift and keystone

  The new integration of swift or s3 and keystone support by rook
  does not have any integration tests yet.

  This commit introduces integration tests for swift and keystone. The
  tests are done against a minimal keystone setup (keystone container
  image from Yaook-project (https://yaook.cloud), sqlite as database
  backend, cert-manager and trust-manager for test certificate setup).

  To prevent hardcoded credentials, passwords are generated
  by the tests. The integration tests use the openstack client
  (keystone- and swift-functionality) (https://docs.openstack.org/
  python-openstackclient/ latest/). This was a concious design decision
  to use client tooling as close as possible to the end user instead of
  using other go-libraries (such as gophercloud).

* add documentation on swift and keystone

  Currently there is no documentation on the use of Swift to access
  an object store as well as the use of OpenStack keystone for
  authentication.

  This commit adds documentation on the use of Swift and OpenStack
  keystone, as well as CRD-related documentation and an example setup.

* add integration tests for S3 via keystone

  This commit introduces integration tests for s3 and keystone. The
  tests are run against the same minimal keystone setup that the tests
  for swift and keystone use.

  The integration tests use the aws s3 client to use client tooling as
  close as possible to the end user instead of using other go-libraries.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
2024-08-08 14:26:21 +02:00
Blaine Gardner a2b0b6449c object: add hosting.advertiseEndpoint config
Add CephObjectStore spec.hosting.advertiseEndpoint configuration. This
provides a clear documented default for which endpoint Rook "advertises"
to dependent resources like CephObjectStores, OBCs, and COSI
Buckets/Accesses and allows users to override the default behavior if
desired.

The current default is to round-robin an endpoint from
spec.hosting.dnsNames, which has proven to be troublesome for some
users' object store configurations. This change provides much-needed
disambiguation for users.

This may be a breaking change for some existing spec.hosting.dnsNames
users. This is unexpected but is documented.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-07-22 14:43:51 -06:00
subhamkrai d429ed8be4 build: update controller runtime to v0.18.4
this commit update cntrl runtime to v0.18.4 and other related deps/

Signed-off-by: subhamkrai <srai@redhat.com>
2024-07-05 09:05:12 +05:30
Travis Nielsen fdacfd51c5 object: create an object store based on shared pools
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.

Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-03-11 11:20:57 -06:00
Jiffin Tony Thottan a941b3c33f object: create cosi user for each object store
Create each cosi user for each object store and secret which holds
credentials.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-09-19 13:48:58 +05:30
travisn 4ff18dadf4 object: remove obsolete bucket health checker removal
The bucket health checker was removed in 1.10. Now in 1.12
we no longer need this removal of the bucket health
checker since it will no longer exist to remove.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-28 16:07:51 -06:00
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
Jiffin Tony Thottan 1f45cfa581 object: use networkspec from clusterinfo spec while running radosgw-admin
The radosgw-admin command uses the network spec from ceph cluster spec
in object context but it is not filled properly in the object package.
But with PR 10898, network spec is available in clusterinfo which can
be used directly. Also removed cluserspec from object context.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-05-30 11:06:03 +05:30
sp98 e87338fbc4 core: skip OBC and Notification controllers
Skip running Object Bucket and Object bucket notification
controllers based on env variable.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-04-24 19:16:06 +05:30
Blaine Gardner a777b1d7d1 object: do not create service for external object stores
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.

This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.

The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-04 17:30:30 -06:00
Blaine Gardner a7c0c7ee93 object: remove health checker
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.

Full discussion: https://github.com/rook/rook/issues/11031

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-18 13:41:46 -06:00
parth-gr 26584fc6e5 core: update loadclusterInfo with multus check
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor

Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-22 14:46:51 +05:30
Jiffin Tony Thottan 8003e764f9 rgw: add custom endpoint list option for zone
User can define his desired endpoint list in Zone CR so that it will
overwrite the default service name for rgw.

Resolves #6432

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2022-08-17 10:03:12 +05:30
Joseph Lee 9639b60fc4 core: skip ceph upgrade check in external cluster
Closes: rook#10688
Signed-off-by: Joseph Lee <joseph@jc-lab.net>
2022-08-11 19:09:40 +09:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Alexander Trost 3005a6fc68 Merge pull request #10137 from koor-tech/fix_9099
rgw: fix dashboard admin creation for multiple object stores
2022-04-27 16:13:38 +00:00
Alexander Trost 5d03061d47 rgw: fix dashboard admin creation for multiple object stores
Check the radosgw-admin realm user list per object store instead of relying
on the ceph dashboard get-rgw-api- command.

Resolves #9099

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-27 14:57:17 +02:00
Sébastien Han 583791c45c core: move clusterInfo code to the controller package
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 11:05:02 +02:00
Sébastien Han 05506e7a68 core: reload go routine after CR is edited
Previously, the struct maintaining the list of cluster was still
initialized with a cluster item. Then the monitoring check will see that
the cluster is part of the struct already and thus won't run the
monitoring go routine again.
Now each time we cancel the context, we also remove the cluster item
from the map so that when the controller runs again, the monitoring
struct is re-populated and the go routine runs and statuses are updated.

Closes: https://github.com/rook/rook/issues/9911
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-03-31 08:49:42 +02:00
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
parth-gr 2dfd64a97c core: add observedGeneration to CR status
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller

Closes: https://github.com/rook/rook/issues/9673

Signed-off-by: parth-gr <paarora@redhat.com>
2022-03-16 19:58:35 +05:30
Blaine Gardner c92c6fbc36 core: rework usage of ReportReconcileResult
ReportReconcileResult should never be given an object that is nil. It is
impossible to force compile-time checking for this because the object's
type is an interface. We can't force this, but we can rework the
function definition to handle more cases where the object may not be
returned fully-complete, and we can rework callers to return structs
(not pointers-to-structs) so it is less likely a caller will pass nil
to the function.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-03-10 16:53:25 -07:00
Sébastien Han 7f32461b8f rgw: fix variable assignment
We need to explicitly assign the ceph version within each scope since
one function returns a pointer and one is not.
Previously, the `desiredCephVersion` was recreated in the scope of the
`else` and thus a new variable `desiredCephVersion` was initialized (see
the `:=`).
Now, each scope creates and assign to clusterInfo a `desiredCephVersion`
value.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-02-14 12:42:30 +01:00
Sébastien Han 856b14e60a object: do not check for upgrade on external mode
The object child controller should not compare versions if the object
store is external. The current version comparison is using the
cmdreporter to check the ceph version of the image in the CephCluster spec.
In external mode, this image is not set so the cmdreporter fails.
Also, this check is only valid for converged mode, so external mode
should be skipped.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-02-10 17:41:01 +01:00
Sébastien Han 5404ec13a2 core: dereference pointer before trying to compare with deepequal
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.

Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-27 17:41:42 +01:00
Blaine Gardner 9dc41f5796 Merge pull request #9427 from BlaineEXE/always-report-reconcile-events
operator: always report events
2021-12-20 08:42:08 -07:00
Blaine Gardner da61ac1ae8 operator: always report events
The original PR which added event reporting unnecessarily "optimized" to
prevent spamming the API controller[1].

It is sometimes important to get events as they happen and not hide new
events behind preexisting older events. For example, in integration
tests, we may often want to wait for a controller to finish processing
an update, and the best way to do that is to wait for the
"ReconcileSucceeded" event. In order for this to be useful, the events
must be reported each time.

If we begin having problems with events being reported too often, then
we should fix the underlying issue of reconciles happening too often
instead of relying on a time-based "optimization" that hides recent
event reports that may be useful.

[1]: https://github.com/rook/rook/pull/7222

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-15 15:24:17 -07:00