Commit Graph
135 Commits
Author SHA1 Message Date
Blaine Gardner fc08e87d44 Revert "object: create cosi user for each object store"
This reverts commit a941b3c33f.

Stop creating the 'cosi' user in the CephObjectStore reconcile. This
step often fails for some amount of time during initial object store
creation, causing frequent user concern. It has also been the source of
some reported failures that would otherwise be non-breaking for certain
users.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-11-21 16:03:32 -07:00
Artem Torubarov 59175f0b40 rgw: pool placement
Signed-off-by: Artem Torubarov <torubarov.a.a@gmail.com>
2024-09-06 16:02:53 +02:00
ee8bcad49d rgw: add support for keystone auth + swift/s3
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>

* extend the API object specs for swift and keystone integration

* adapt rgw to the new go-ceph version

  - The parameter lists of the API call have changes, as parameters
    ignored by the RGW Admin Ops API are no longer serialized, therefore
    the mock has to be adapted.

  - There is now validation for the user keys that are passed to the
    User get API, therefore things failed when we had empty keys in our
    User proxy object.

* expand the reconcile loop for the swift and keystone integration

* fix minor mistakes in design document

* add env var to pass extra args to minikube

  Minikube decides CPU cores and memory automatically based on the
  available resources on the machine which may be insufficient to
  run rook. This commit adds an environment variable to add arbitrary
  arguments to the minikube command, so both can be specified if
  desired.

* integration tests for swift and keystone

  The new integration of swift or s3 and keystone support by rook
  does not have any integration tests yet.

  This commit introduces integration tests for swift and keystone. The
  tests are done against a minimal keystone setup (keystone container
  image from Yaook-project (https://yaook.cloud), sqlite as database
  backend, cert-manager and trust-manager for test certificate setup).

  To prevent hardcoded credentials, passwords are generated
  by the tests. The integration tests use the openstack client
  (keystone- and swift-functionality) (https://docs.openstack.org/
  python-openstackclient/ latest/). This was a concious design decision
  to use client tooling as close as possible to the end user instead of
  using other go-libraries (such as gophercloud).

* add documentation on swift and keystone

  Currently there is no documentation on the use of Swift to access
  an object store as well as the use of OpenStack keystone for
  authentication.

  This commit adds documentation on the use of Swift and OpenStack
  keystone, as well as CRD-related documentation and an example setup.

* add integration tests for S3 via keystone

  This commit introduces integration tests for s3 and keystone. The
  tests are run against the same minimal keystone setup that the tests
  for swift and keystone use.

  The integration tests use the aws s3 client to use client tooling as
  close as possible to the end user instead of using other go-libraries.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
2024-08-08 14:26:21 +02:00
Blaine Gardner a2b0b6449c object: add hosting.advertiseEndpoint config
Add CephObjectStore spec.hosting.advertiseEndpoint configuration. This
provides a clear documented default for which endpoint Rook "advertises"
to dependent resources like CephObjectStores, OBCs, and COSI
Buckets/Accesses and allows users to override the default behavior if
desired.

The current default is to round-robin an endpoint from
spec.hosting.dnsNames, which has proven to be troublesome for some
users' object store configurations. This change provides much-needed
disambiguation for users.

This may be a breaking change for some existing spec.hosting.dnsNames
users. This is unexpected but is documented.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-07-22 14:43:51 -06:00
subhamkrai d429ed8be4 build: update controller runtime to v0.18.4
this commit update cntrl runtime to v0.18.4 and other related deps/

Signed-off-by: subhamkrai <srai@redhat.com>
2024-07-05 09:05:12 +05:30
Travis Nielsen fdacfd51c5 object: create an object store based on shared pools
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.

Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-03-11 11:20:57 -06:00
Jiffin Tony Thottan a941b3c33f object: create cosi user for each object store
Create each cosi user for each object store and secret which holds
credentials.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-09-19 13:48:58 +05:30
travisn 4ff18dadf4 object: remove obsolete bucket health checker removal
The bucket health checker was removed in 1.10. Now in 1.12
we no longer need this removal of the bucket health
checker since it will no longer exist to remove.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-28 16:07:51 -06:00
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
Jiffin Tony Thottan 1f45cfa581 object: use networkspec from clusterinfo spec while running radosgw-admin
The radosgw-admin command uses the network spec from ceph cluster spec
in object context but it is not filled properly in the object package.
But with PR 10898, network spec is available in clusterinfo which can
be used directly. Also removed cluserspec from object context.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-05-30 11:06:03 +05:30
sp98 e87338fbc4 core: skip OBC and Notification controllers
Skip running Object Bucket and Object bucket notification
controllers based on env variable.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-04-24 19:16:06 +05:30
Blaine Gardner a777b1d7d1 object: do not create service for external object stores
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.

This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.

The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-04 17:30:30 -06:00
Blaine Gardner a7c0c7ee93 object: remove health checker
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.

Full discussion: https://github.com/rook/rook/issues/11031

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-18 13:41:46 -06:00
parth-gr 26584fc6e5 core: update loadclusterInfo with multus check
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor

Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-22 14:46:51 +05:30
Jiffin Tony Thottan 8003e764f9 rgw: add custom endpoint list option for zone
User can define his desired endpoint list in Zone CR so that it will
overwrite the default service name for rgw.

Resolves #6432

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2022-08-17 10:03:12 +05:30
Joseph Lee 9639b60fc4 core: skip ceph upgrade check in external cluster
Closes: rook#10688
Signed-off-by: Joseph Lee <joseph@jc-lab.net>
2022-08-11 19:09:40 +09:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Alexander Trost 3005a6fc68 Merge pull request #10137 from koor-tech/fix_9099
rgw: fix dashboard admin creation for multiple object stores
2022-04-27 16:13:38 +00:00
Alexander Trost 5d03061d47 rgw: fix dashboard admin creation for multiple object stores
Check the radosgw-admin realm user list per object store instead of relying
on the ceph dashboard get-rgw-api- command.

Resolves #9099

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-27 14:57:17 +02:00
Sébastien Han 583791c45c core: move clusterInfo code to the controller package
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 11:05:02 +02:00
Sébastien Han 05506e7a68 core: reload go routine after CR is edited
Previously, the struct maintaining the list of cluster was still
initialized with a cluster item. Then the monitoring check will see that
the cluster is part of the struct already and thus won't run the
monitoring go routine again.
Now each time we cancel the context, we also remove the cluster item
from the map so that when the controller runs again, the monitoring
struct is re-populated and the go routine runs and statuses are updated.

Closes: https://github.com/rook/rook/issues/9911
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-03-31 08:49:42 +02:00
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
parth-gr 2dfd64a97c core: add observedGeneration to CR status
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller

Closes: https://github.com/rook/rook/issues/9673

Signed-off-by: parth-gr <paarora@redhat.com>
2022-03-16 19:58:35 +05:30
Blaine Gardner c92c6fbc36 core: rework usage of ReportReconcileResult
ReportReconcileResult should never be given an object that is nil. It is
impossible to force compile-time checking for this because the object's
type is an interface. We can't force this, but we can rework the
function definition to handle more cases where the object may not be
returned fully-complete, and we can rework callers to return structs
(not pointers-to-structs) so it is less likely a caller will pass nil
to the function.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-03-10 16:53:25 -07:00
Sébastien Han 7f32461b8f rgw: fix variable assignment
We need to explicitly assign the ceph version within each scope since
one function returns a pointer and one is not.
Previously, the `desiredCephVersion` was recreated in the scope of the
`else` and thus a new variable `desiredCephVersion` was initialized (see
the `:=`).
Now, each scope creates and assign to clusterInfo a `desiredCephVersion`
value.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-02-14 12:42:30 +01:00
Sébastien Han 856b14e60a object: do not check for upgrade on external mode
The object child controller should not compare versions if the object
store is external. The current version comparison is using the
cmdreporter to check the ceph version of the image in the CephCluster spec.
In external mode, this image is not set so the cmdreporter fails.
Also, this check is only valid for converged mode, so external mode
should be skipped.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-02-10 17:41:01 +01:00
Sébastien Han 5404ec13a2 core: dereference pointer before trying to compare with deepequal
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.

Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-27 17:41:42 +01:00
Blaine Gardner 9dc41f5796 Merge pull request #9427 from BlaineEXE/always-report-reconcile-events
operator: always report events
2021-12-20 08:42:08 -07:00
Blaine Gardner da61ac1ae8 operator: always report events
The original PR which added event reporting unnecessarily "optimized" to
prevent spamming the API controller[1].

It is sometimes important to get events as they happen and not hide new
events behind preexisting older events. For example, in integration
tests, we may often want to wait for a controller to finish processing
an update, and the best way to do that is to wait for the
"ReconcileSucceeded" event. In order for this to be useful, the events
must be reported each time.

If we begin having problems with events being reported too often, then
we should fix the underlying issue of reconciles happening too often
instead of relying on a time-based "optimization" that hides recent
event reports that may be useful.

[1]: https://github.com/rook/rook/pull/7222

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-15 15:24:17 -07:00
Blaine Gardner 0c76be8c38 pool: file: object: clean up health checkers for force deletion
When CephBlockPool, CephFilesystem, or CephObjectStore resources are
deleted after removing their finalizer, the code path to stop monitoring
was not stopping monitoring since a non-present resource does not have a
name and namespace attached. When the object is deleted, ensure the
internal representation used to stop monitoring has a name and namespace
to fix the issue.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-13 16:07:01 -07:00
Blaine Gardner 03ba7dec64 pool: file: object: clean up stop health checkers
Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.

Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-19 10:29:12 -07:00
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Blaine Gardner 2f850b6ae6 Merge pull request #8613 from subhamkrai/remove-nautilus
ceph: remove ceph nautilus, ceph octopus to default
2021-10-25 09:09:18 -06:00
Sébastien Han fc9c4b6fc9 Merge pull request #9010 from thotz/rgw-external-endpoint
ceph: update endpoint with IP for external RGW server
2021-10-25 15:49:08 +02:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
Jiffin Tony Thottan d4562f6b83 ceph: update endpoint with IP for external RGW server
For external RGW server use the IP mentioned in Gateway for admin Ops
operattions.

Fixes: #8916
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-10-21 19:13:10 +05:30
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
PixelJonas 8733bf6272 ceph: use normal quotation marks for comment
replaces all occurences of lduo/rduo quotation marks to make
sure that using the snippets in a k8s manifest will work

This fixes an issue with ArgoCD not being able
to apply `common.yaml` because of an encoding issue

Signed-off-by: PixelJonas <jonas@janz.digital>
2021-10-20 10:09:22 +02:00
Blaine Gardner eadcd757b3 rgw: replace period update --commit with function
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.

Resolves #8879

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 14:45:39 -06:00
Blaine Gardner 5383ba2df2 ceph: retry object health check if creation fails
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:24:18 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 2d55e69416 ceph: move scheme initialization to the same place
Let's initialize the schemes in a single place instead of doing it
when each controller initializes.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-06 11:33:10 +02:00
Sébastien Han daf53ea585 ceph: add unit test for object store user dependents
Now that we have the mock client for the admin ops api we can add unit
tests.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 10:01:53 +02:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Sébastien Han 677c93f3aa ceph: operate over object store pools concurrently
We now create and delete pools in parallel.
Before this patch, the deletion:

```
2021-06-08 13:48:26.557328 I | cephclient: no images/snapshosts present in pool "my-store.rgw.control"
2021-06-08 13:48:26.557383 I | cephclient: purging pool "my-store.rgw.control" (id=1)
2021-06-08 13:48:27.030021 I | ceph-object-controller: done disabling the dashboard api secret key
2021-06-08 13:48:28.800661 I | cephclient: purge completed for pool "my-store.rgw.control"
2021-06-08 13:48:29.085744 I | cephclient: no images/snapshosts present in pool "my-store.rgw.meta"
2021-06-08 13:48:29.085762 I | cephclient: purging pool "my-store.rgw.meta" (id=3)
2021-06-08 13:48:30.835487 I | cephclient: purge completed for pool "my-store.rgw.meta"
2021-06-08 13:48:31.128349 I | cephclient: no images/snapshosts present in pool "my-store.rgw.log"
2021-06-08 13:48:31.128368 I | cephclient: purging pool "my-store.rgw.log" (id=4)
2021-06-08 13:48:32.864376 I | cephclient: purge completed for pool "my-store.rgw.log"
2021-06-08 13:48:33.180923 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.index"
2021-06-08 13:48:33.181049 I | cephclient: purging pool "my-store.rgw.buckets.index" (id=5)
2021-06-08 13:48:34.892163 I | cephclient: purge completed for pool "my-store.rgw.buckets.index"
2021-06-08 13:48:35.179952 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.non-ec"
2021-06-08 13:48:35.179994 I | cephclient: purging pool "my-store.rgw.buckets.non-ec" (id=6)
2021-06-08 13:48:37.376067 I | cephclient: purge completed for pool "my-store.rgw.buckets.non-ec"
2021-06-08 13:48:37.665201 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.data"
2021-06-08 13:48:37.665218 I | cephclient: purging pool "my-store.rgw.buckets.data" (id=8)
2021-06-08 13:48:39.394717 I | cephclient: purge completed for pool "my-store.rgw.buckets.data"
2021-06-08 13:48:39.725528 I | cephclient: no images/snapshosts present in pool ".rgw.root"
2021-06-08 13:48:39.725567 I | cephclient: purging pool ".rgw.root" (id=7)
2021-06-08 13:48:41.424628 I | cephclient: purge completed for pool ".rgw.root"
```

It took 15sec to cleanup...
Now with this patch:

```
2021-06-08 14:15:25.621472 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.index"
2021-06-08 14:15:25.621570 I | cephclient: purging pool "my-store.rgw.buckets.index" (id=12)
2021-06-08 14:15:25.668708 I | cephclient: no images/snapshosts present in pool "my-store.rgw.log"
2021-06-08 14:15:25.668729 I | cephclient: purging pool "my-store.rgw.log" (id=11)
2021-06-08 14:15:25.693002 I | cephclient: no images/snapshosts present in pool "my-store.rgw.meta"
2021-06-08 14:15:25.693050 I | cephclient: purging pool "my-store.rgw.meta" (id=10)
2021-06-08 14:15:25.698830 I | cephclient: no images/snapshosts present in pool "my-store.rgw.control"
2021-06-08 14:15:25.698854 I | cephclient: purging pool "my-store.rgw.control" (id=9)
2021-06-08 14:15:25.701732 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.non-ec"
2021-06-08 14:15:25.701758 I | cephclient: purging pool "my-store.rgw.buckets.non-ec" (id=13)
2021-06-08 14:15:25.702816 I | cephclient: no images/snapshosts present in pool ".rgw.root"
2021-06-08 14:15:25.702836 I | cephclient: purging pool ".rgw.root" (id=14)
2021-06-08 14:15:25.717144 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.data"
2021-06-08 14:15:25.717212 I | cephclient: purging pool "my-store.rgw.buckets.data" (id=15)
2021-06-08 14:15:26.393246 I | ceph-object-controller: done disabling the dashboard api secret key
```

It tool around 1sec.

The creation before this patch:

```

2021-06-08 16:47:10.669484 I | ceph-spec: adding finalizer "cephobjectstore.ceph.rook.io" on "my-store"
2021-06-08 16:47:10.677253 E | ceph-object-controller: failed to set object store "rook-ceph/my-store" status to "Progressing". failed to update object "my-store" status: Operation cannot be fulfilled on cephobjectstores.ceph.rook.io "my-store": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:10.682231 I | op-mon: parsing mon endpoints: b=10.111.63.108:6789,c=10.108.123.222:6789,a=10.100.223.211:6789
2021-06-08 16:47:10.985002 I | ceph-object-controller: reconciling object store deployments
2021-06-08 16:47:11.002364 I | ceph-object-controller: ceph object store gateway service running at 10.99.31.5
2021-06-08 16:47:11.002384 I | ceph-object-controller: reconciling object store pools
2021-06-08 16:47:14.336086 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.control"
2021-06-08 16:47:16.115072 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:16.138224 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:16.391914 I | cephclient: creating replicated pool my-store.rgw.control succeeded
2021-06-08 16:47:16.391943 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.control"
2021-06-08 16:47:20.390562 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.meta"
2021-06-08 16:47:22.423269 I | cephclient: creating replicated pool my-store.rgw.meta succeeded
2021-06-08 16:47:22.423294 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.meta"
2021-06-08 16:47:26.502874 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.log"
2021-06-08 16:47:28.555252 I | cephclient: creating replicated pool my-store.rgw.log succeeded
2021-06-08 16:47:28.555308 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.log"
2021-06-08 16:47:32.619546 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.index"
2021-06-08 16:47:34.663020 I | cephclient: creating replicated pool my-store.rgw.buckets.index succeeded
2021-06-08 16:47:34.663050 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.index"
2021-06-08 16:47:38.712886 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:47:40.748430 I | cephclient: creating replicated pool my-store.rgw.buckets.non-ec succeeded
2021-06-08 16:47:40.748467 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:47:44.830659 I | cephclient: setting pool property "compression_mode" to "none" on pool ".rgw.root"
2021-06-08 16:47:46.875834 I | cephclient: creating replicated pool .rgw.root succeeded
2021-06-08 16:47:46.875863 I | cephclient: setting pool property "pg_num_min" to "8" on pool ".rgw.root"
2021-06-08 16:47:50.933717 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.data"
2021-06-08 16:47:52.969461 I | cephclient: creating replicated pool my-store.rgw.buckets.data succeeded
2021-06-08 16:47:52.969549 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-06-08 16:47:53.585737 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-06-08 16:47:53.585777 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-06-08 16:47:53.585787 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-06-08 16:47:53.585802 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-06-08 16:47:53.930416 I | ceph-object-controller: setting rgw config flags
2021-06-08 16:47:53.931363 I | op-config: setting "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:47:54.201200 I | op-config: successfully set "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:47:54.201225 I | op-config: setting "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:47:54.454589 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:47:54.454614 I | op-config: setting "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:47:54.710218 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:47:54.710237 I | op-config: setting "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:47:54.969049 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:47:54.969069 I | op-config: setting "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:47:55.255491 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:47:55.255631 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-06-08 16:47:55.279316 I | ceph-object-controller: enabling rgw dashboard
2021-06-08 16:47:55.361208 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:56.225904 I | ceph-object-controller: setting the dashboard api secret key
2021-06-08 16:47:56.225984 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
2021-06-08 16:47:56.589452 I | ceph-object-controller: starting rgw healthcheck
2021-06-08 16:47:56.650045 I | ceph-object-controller: done setting the dashboard api secret key
```

It took 46sec, and I've seen it taking almost a 1min sometimes.
Now with this patch:

```
2021-06-08 16:51:35.259558 I | ceph-spec: adding finalizer "cephobjectstore.ceph.rook.io" on "my-store"
2021-06-08 16:51:35.270524 E | ceph-object-controller: failed to set object store "rook-ceph/my-store" status to "Progressing". failed to update object "my-store" status: Operation cannot be fulfilled on cephobjectstores.ceph.rook.io "my-store": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:51:35.274387 I | op-mon: parsing mon endpoints: b=10.111.63.108:6789,c=10.108.123.222:6789,a=10.100.223.211:6789
2021-06-08 16:51:35.599023 I | ceph-object-controller: reconciling object store deployments
2021-06-08 16:51:35.607256 I | ceph-object-controller: ceph object store gateway service running at 10.104.254.110
2021-06-08 16:51:35.607315 I | ceph-object-controller: reconciling object store pools
2021-06-08 16:51:39.337735 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:51:40.343290 I | cephclient: setting pool property "compression_mode" to "none" on pool ".rgw.root"
2021-06-08 16:51:40.346335 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.index"
2021-06-08 16:51:40.362670 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.meta"
2021-06-08 16:51:40.363445 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.log"
2021-06-08 16:51:40.364321 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.control"
2021-06-08 16:51:41.412870 I | cephclient: creating replicated pool my-store.rgw.buckets.non-ec succeeded
2021-06-08 16:51:41.412904 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:51:42.460439 I | cephclient: creating replicated pool my-store.rgw.control succeeded
2021-06-08 16:51:42.460503 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.control"
2021-06-08 16:51:42.470666 I | cephclient: creating replicated pool my-store.rgw.meta succeeded
2021-06-08 16:51:42.470727 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.meta"
2021-06-08 16:51:42.473029 I | cephclient: creating replicated pool .rgw.root succeeded
2021-06-08 16:51:42.473064 I | cephclient: setting pool property "pg_num_min" to "8" on pool ".rgw.root"
2021-06-08 16:51:42.473895 I | cephclient: creating replicated pool my-store.rgw.buckets.index succeeded
2021-06-08 16:51:42.473923 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.index"
2021-06-08 16:51:42.478625 I | cephclient: creating replicated pool my-store.rgw.log succeeded
2021-06-08 16:51:42.478671 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.log"
2021-06-08 16:51:46.606905 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.data"
2021-06-08 16:51:48.627320 I | cephclient: creating replicated pool my-store.rgw.buckets.data succeeded
2021-06-08 16:51:48.627368 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-06-08 16:51:49.325086 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-06-08 16:51:49.325108 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-06-08 16:51:49.325121 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-06-08 16:51:49.325134 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-06-08 16:51:49.657898 I | ceph-object-controller: setting rgw config flags
2021-06-08 16:51:49.657920 I | op-config: setting "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:51:49.917898 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:51:49.917917 I | op-config: setting "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:51:50.184939 I | op-config: successfully set "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:51:50.184970 I | op-config: setting "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:51:50.463799 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:51:50.463821 I | op-config: setting "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:51:50.720953 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:51:50.720999 I | op-config: setting "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:51:50.974875 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:51:50.975024 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-06-08 16:51:51.019820 I | ceph-object-controller: enabling rgw dashboard
2021-06-08 16:51:51.980117 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
2021-06-08 16:51:51.980199 I | ceph-object-controller: setting the dashboard api secret key
2021-06-08 16:51:52.399997 I | ceph-object-controller: done setting the dashboard api secret key
2021-06-08 16:51:52.435936 I | ceph-object-controller: starting rgw healthcheck
```

It took 17sec.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 16:09:22 +02:00
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Sébastien Han e478bdadb1 ceph: stop the monitoring before cleanup
Saw this today in the logs:

```
2021-06-02 13:02:54.787838 I | ceph-block-pool-controller: deleting pool "testpool"
2021-06-02 13:02:56.178094 I | cephclient: no images/snapshosts present in pool "testpool"
2021-06-02 13:02:56.178125 I | cephclient: purging pool "testpool" (id=17)
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0x1cef49f]

goroutine 1607 [running]:
github.com/rook/rook/pkg/operator/ceph/pool.(*mirrorChecker).checkMirroringHealth(0xc001a98ba0, 0x8, 0xc000d34c60)
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/health.go:115 +0xdf
github.com/rook/rook/pkg/operator/ceph/pool.(*mirrorChecker).checkMirroring(0xc001a98ba0, 0xc0016fd500)
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/health.go:82 +0x147
created by github.com/rook/rook/pkg/operator/ceph/pool.(*ReconcileCephBlockPool).reconcile
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/controller.go:298 +0xee5
```

Essentially, it's intermittent but when deleting the pool the
healthcheck kicked in and fetched the mirroring status, which returned
empty. THe subsequent code tried to access content of a nil pointer,
  hence the error.
So now we stop monitoring first, then we proceed with the deletion.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-02 16:14:10 +02:00
Jiffin Tony Thottan aea21d9c84 ceph: service server cert support for rgw
Service serving certificates are intended to applications that require
encryption in openshift. These certificates are issued as TLS web server
certificates. Currently RGW supports TLS authentication with help of
certs passed as secrets, in this case we add following details as
`service.annotations` in the Objectstore Gateway Spec :
```
service:
  annotations:
    service.beta.openshift.io/serving-cert-secret-name: <name for
autogenerated secret>
```

More details about service serving cert can be found at :
https://docs.openshift.com/container-platform/4.6/security/certificates/service-serving-certificate.html

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-26 18:15:12 +05:30