Commit Graph
74 Commits
Author SHA1 Message Date
Anas Khan 0bf6f84ff5 object: return an error when the multisite zone is not found
retrieveMultisiteZone is meant to gate the object-store reconcile on the
backing Ceph zone existing: it runs "radosgw-admin zone get" and, when
that fails, returns a non-nil error so the caller requeues. That gate is
dead code. The ENOENT check declares an inner err from exec.ExtractExitCode
that shadows the outer command error, and ExtractExitCode returns a nil
error for the ordinary exit failures radosgw-admin produces. Both the
ENOENT branch and the else branch then wrap that shadowed nil, and
errors.Wrapf(nil, ...) is nil, so the function returns
(waitForRequeueIfObjectStoreNotReady, nil). The caller only propagates the
requeue when the error is non-nil, so the requeue is dropped and reconcile
runs on.

The result is that a failed "zone get" no longer backs off. Reconcile
proceeds to stand up the object store anyway -- the RGW service, the
admin-ops endpoint, the deployment, and the pools radosgw scaffolds as it
comes up -- for a multisite store whose backing zone does not exist. This
is not the "normal multisite bootstrap" transient the original wording
suggested. getMultisiteResourceNames runs immediately before this and
already requeues until the CephObjectZone CR reports Ready, and the zone
controller marks it Ready only after it has created the Ceph zone, so a
healthy bootstrap never reaches this gate with a missing zone. That
CR-Ready gate, not this one, is what actually blocks bootstrap.

Where the dead gate does bite is the cases the CR status cannot cover:

- the Ceph zone deleted or renamed out of band while the CR still reads Ready
- a zone controller that reports Ready without leaving a usable zone behind
- any non-ENOENT "zone get" failure, e.g. a permission or connectivity
  error, which the else branch swallows the same way

The check was correct when fc579f4520 introduced it in 2020 with
exec.ExitStatus, which returns (code, ok) and leaves err unshadowed.
bd58790c31 ("ceph: proxy ceph commands when multus is configured", 2021)
swapped it to exec.ExtractExitCode as an unrelated drive-by, inverting the
contract and killing the gate. The sibling realm, zonegroup, and zone
controllers were left untouched and still use exec.ExitStatus today.

Restore exec.ExitStatus so the outer error is no longer shadowed, both
branches wrap the real command error, and the caller requeues until the
zone exists -- matching the sibling controllers.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-21 12:40:25 -07:00
Artem Torubarov a7f176b366 rgw: skip non-raw commands on object deletion path when pools not exists
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-02-26 09:48:15 +01:00
Artem Torubarov b47d2c2f79 rgw: force objectStore controller to wait until zone and sharedPools are ready
Fixes race condition for rgw multisite deployments caused by ObjectStore
and Zone controllers calling period commit concurrently. See: https://github.com/rook/rook/issues/17013

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-02-24 11:43:57 +01:00
subhamkrai 08344f39f6 build: update deprecated api GetEventRecorderFor
with controller runtime latest version vO.23.0, golangci lint
is complaining about deprecated api deprecated.
This commit updates the api and adds the requried rbacs

Signed-off-by: subhamkrai <srai@redhat.com>
2026-02-09 21:48:02 +05:30
subhamkrai 304cdc8c83 core: remove support for reef v18 in Rook v1.19
For Rook v1.19,removing support for Ceph Reef. It is at end of life.
Users on Reef can continue to use Rook v1.18.x or older.
The latest two Ceph versions Squid and Tentacle are supported.

Signed-off-by: subhamkrai <srai@redhat.com>
2025-12-04 13:11:35 +05:30
Blaine Gardner 029c345372 Merge pull request #15813 from BlaineEXE/auth-rotate
object: automate RGW cephx key rotation
2025-06-10 12:51:50 -06:00
Blaine Gardner 5f50e85a25 object: implement cephx key rotation for rgw keys
This is the first implementation of CephX key rotation in Rook.
Adds key rotation API from design PR 15915.

Also add helper methods for rotating keys, determining when keys
need to be rotated, and for generating CephX key statuses. These will be
reusable for other Rook reconciles beyond object/RGW.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2025-06-10 09:06:58 -06:00
Carlos Barria 030ada45b0 core: fix golangci-lint check QF1003 QF1002
Refactored multiple conditional if-else blocks into switch statements to improve readability and maintainability. Also removed outdated staticcheck rule comments (QF1002, QF1003) from .golangci.yaml.

Signed-off-by: Carlos Barria <cbarria@yahoo.com>
2025-05-27 13:49:49 -04:00
Travis Nielsen f7fb1bc0f2 core: skip reconcile when adding the finalizers
When finalizers are added to the CRs, a follow-up reconcile
will be triggered due to the increased generation on the CR.
Therefore, abort the initial reconcile when adding the finalizer,
and allow the follow-up reconcile to complete the configuration.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-04-08 13:42:58 -06:00
Artem Torubarov dd0d500980 rgw: watch referenced secrets
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2025-03-17 14:09:05 +01:00
Artem Torubarov 00aa282074 rgw: use list in secret watch map func
Signed-off-by: Artem Torubarov <artem.torubarov@clyso.com>
2025-03-13 11:19:45 +01:00
Artem Torubarov bed0c23274 rgw: watch config secret changes
Implements #15460. If ObjectStore reading value from a secret
provided by user, the secret will be annotated with ObjectStore
name and ObjectStore will be reconciled on secret change.

Signed-off-by: Artem Torubarov <artem.torubarov@clyso.com>
2025-03-12 11:37:12 +01:00
Blaine Gardner fc08e87d44 Revert "object: create cosi user for each object store"
This reverts commit a941b3c33f.

Stop creating the 'cosi' user in the CephObjectStore reconcile. This
step often fails for some amount of time during initial object store
creation, causing frequent user concern. It has also been the source of
some reported failures that would otherwise be non-breaking for certain
users.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-11-21 16:03:32 -07:00
Travis Nielsen b665d7a7b7 core: remove support for ceph quincy
Given that Ceph Quincy (v17) is past end of life,
remove Quincy from the supported Ceph versions,
examples, and documentation.

Supported versions now include only Reef and Squid.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-10-03 11:12:55 -06:00
sp98 a37c47641f core: disable mirroring in image mode
For image mode mirroring, if cephBlockPool.Pool.Spec.Mirroring.Enable
is set to false, then remove the peer cluster and disable mirroring on
all the pool if the user has disabled mirroring on all the pool images.

If mirroring is not disabled on all the pool images, then reconcile will
fail asking the users to manually disable mirroring on those images.

Signed-off-by: sp98 <sapillai@redhat.com>
2024-03-19 14:20:51 +05:30
travisn 806608cdbc pool: allow setting the application on a pool
Rook has been setting the application automatically on all
pools to rbd for CephBlockPools, rook-ceph-rgw for
CephObjectStores, mgr on the built-in .mgr pool,
and nfs on the built-in .nfs pool.

The legacy pool device_health_metrics is long gone
from Pacific which is no longer supported, so we can
remove special handling for that pool in the upgrade
guide and in the code.

The application setting is now available on the pool spec
although it is not expected to commonly need to override
the default applications set by Rook.

The application for CephFilesystem pools is now being
set to cephfs, where it was previously blank.

Signed-off-by: travisn <tnielsen@redhat.com>
2024-02-14 16:53:25 -07:00
travisn 6f8e42422d core: fix golang linter issues with variables in loops
Loop variables cannot be reliably uses since they will
change with each iteration. Update these loop variable
uses to be safe by indexing the slice rather than
using the loop variable directly.

Also suppress the linter issues for passwords used
in tests.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-12-05 14:14:58 -07:00
travisn 03d077aa6b core: remove support for ceph pacific
Pacific is end of life and no longer necessary to
support in Rook with v1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-14 17:07:03 -07:00
travisn 80244fa6ba object: change is_master from string to bool
In Reef the is_master changed from a string to a bool
so we must update the type for proper json
serialization.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-10-25 15:26:17 -06:00
Blaine Gardner 3c7499facf test: mark unit test secrets as not secret
There are 2 cases of randomly generated secrets copied into Rook's unit
test code that have been flagged by Gitleaks. Add a comment to both
cases to help the tool understand that these aren't real production
secrets -- just unit test stand-ins.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-10-11 16:42:29 -06:00
Jiffin Tony Thottan a941b3c33f object: create cosi user for each object store
Create each cosi user for each object store and secret which holds
credentials.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-09-19 13:48:58 +05:30
travisn 4ff18dadf4 object: remove obsolete bucket health checker removal
The bucket health checker was removed in 1.10. Now in 1.12
we no longer need this removal of the bucket health
checker since it will no longer exist to remove.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-28 16:07:51 -06:00
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
parth-gr 8e317ee074 ci: update golangci-lint version as it fails for some k8s version in 1.10
Closes: https://github.com/rook/rook/issues/11896

Signed-off-by: parth-gr <paarora@redhat.com>
2023-03-16 20:59:22 +05:30
Blaine Gardner a777b1d7d1 object: do not create service for external object stores
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.

This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.

The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-04 17:30:30 -06:00
Blaine Gardner a7c0c7ee93 object: remove health checker
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.

Full discussion: https://github.com/rook/rook/issues/11031

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-18 13:41:46 -06:00
Jiffin Tony Thottan 8003e764f9 rgw: add custom endpoint list option for zone
User can define his desired endpoint list in Zone CR so that it will
overwrite the default service name for rgw.

Resolves #6432

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2022-08-17 10:03:12 +05:30
Travis Nielsen dad97f3425 core: remove support for ceph octopus
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-07 15:03:26 -06:00
Alexander Trost 5d03061d47 rgw: fix dashboard admin creation for multiple object stores
Check the radosgw-admin realm user list per object store instead of relying
on the ceph dashboard get-rgw-api- command.

Resolves #9099

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-27 14:57:17 +02:00
Alexander Trost 8686296e17 core: remove double imported packages
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-25 13:51:45 +02:00
Pranshu Srivastava 5e414666cc rgw: add unit tests for the external ceph object store reconciler
Added few test cases to check for the following:
* creation of an external object store
* deletion of an external object store
* creation of an external object store with a missing secret
* creation of an external object store with ExternalRGWEndpoints set to
  nil.

Signed-off-by: Pranshu Srivastava <rexagod@gmail.com>
2022-03-15 18:08:54 +05:30
Sébastien Han 7f32461b8f rgw: fix variable assignment
We need to explicitly assign the ceph version within each scope since
one function returns a pointer and one is not.
Previously, the `desiredCephVersion` was recreated in the scope of the
`else` and thus a new variable `desiredCephVersion` was initialized (see
the `:=`).
Now, each scope creates and assign to clusterInfo a `desiredCephVersion`
value.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-02-14 12:42:30 +01:00
Travis Nielsen 221b3dfb72 object: update pool properties during reconcile
The reconcile was skipping updating most pool properties for
object stores. The implementation of pools between the file,
object, and pool controllers had some duplicate code, so
this change also factors out the common code for better
reuse in a single place. Anytime a pool is created or updated,
it will now consistently update all the pool properties
that are expected to be modifiable.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-02-04 13:41:12 -07:00
Sébastien Han 5404ec13a2 core: dereference pointer before trying to compare with deepequal
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.

Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-27 17:41:42 +01:00
Blaine Gardner da61ac1ae8 operator: always report events
The original PR which added event reporting unnecessarily "optimized" to
prevent spamming the API controller[1].

It is sometimes important to get events as they happen and not hide new
events behind preexisting older events. For example, in integration
tests, we may often want to wait for a controller to finish processing
an update, and the best way to do that is to wait for the
"ReconcileSucceeded" event. In order for this to be useful, the events
must be reported each time.

If we begin having problems with events being reported too often, then
we should fix the underlying issue of reconciles happening too often
instead of relying on a time-based "optimization" that hides recent
event reports that may be useful.

[1]: https://github.com/rook/rook/pull/7222

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-15 15:24:17 -07:00
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Blaine Gardner 956430826c rgw: add integration test for committing period
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.

Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.

To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
   of its operation that didn't cause failures in the past

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 15:24:59 -06:00
Blaine Gardner eadcd757b3 rgw: replace period update --commit with function
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.

Resolves #8879

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 14:45:39 -06:00
Blaine Gardner 5383ba2df2 ceph: retry object health check if creation fails
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:24:18 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Santosh Pillai 113376f1ab ceph: timeout radosgw-admin cli commands
When creating object store, `radosgw-admin realm get ..` command is stuck forever when required number of OSDs are not available. Because of this the uninstall of object store is also stuck. User has to manually remove the finalizer to delete the object store.This PR uses `ExecuteCommandWithTimeout` for running `radosgw-admin` command. Timeout during installation will be reconciled. Cleanup will be treated as best effort. Any errors during uninstalling of single site object store will only be logged.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-02-01 10:57:04 +05:30
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Ali Maredia 8964941d44 ceph: add unit tests for object store controller with multisite
Add unit tests for when the "Zone" is configured
in an object store so that the object store joins
the CephObjectZone and the zone's corresponding
multisite configuration.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-10-22 11:36:06 -04:00
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30