Commit Graph
129 Commits
Author SHA1 Message Date
Jiffin Tony Thottan a941b3c33f object: create cosi user for each object store
Create each cosi user for each object store and secret which holds
credentials.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-09-19 13:48:58 +05:30
travisn 4ff18dadf4 object: remove obsolete bucket health checker removal
The bucket health checker was removed in 1.10. Now in 1.12
we no longer need this removal of the bucket health
checker since it will no longer exist to remove.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-28 16:07:51 -06:00
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
Jiffin Tony Thottan 1f45cfa581 object: use networkspec from clusterinfo spec while running radosgw-admin
The radosgw-admin command uses the network spec from ceph cluster spec
in object context but it is not filled properly in the object package.
But with PR 10898, network spec is available in clusterinfo which can
be used directly. Also removed cluserspec from object context.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-05-30 11:06:03 +05:30
sp98 e87338fbc4 core: skip OBC and Notification controllers
Skip running Object Bucket and Object bucket notification
controllers based on env variable.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-04-24 19:16:06 +05:30
Blaine Gardner a777b1d7d1 object: do not create service for external object stores
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.

This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.

The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-04 17:30:30 -06:00
Blaine Gardner a7c0c7ee93 object: remove health checker
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.

Full discussion: https://github.com/rook/rook/issues/11031

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-18 13:41:46 -06:00
parth-gr 26584fc6e5 core: update loadclusterInfo with multus check
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor

Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-22 14:46:51 +05:30
Jiffin Tony Thottan 8003e764f9 rgw: add custom endpoint list option for zone
User can define his desired endpoint list in Zone CR so that it will
overwrite the default service name for rgw.

Resolves #6432

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2022-08-17 10:03:12 +05:30
Joseph Lee 9639b60fc4 core: skip ceph upgrade check in external cluster
Closes: rook#10688
Signed-off-by: Joseph Lee <joseph@jc-lab.net>
2022-08-11 19:09:40 +09:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Alexander Trost 3005a6fc68 Merge pull request #10137 from koor-tech/fix_9099
rgw: fix dashboard admin creation for multiple object stores
2022-04-27 16:13:38 +00:00
Alexander Trost 5d03061d47 rgw: fix dashboard admin creation for multiple object stores
Check the radosgw-admin realm user list per object store instead of relying
on the ceph dashboard get-rgw-api- command.

Resolves #9099

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-27 14:57:17 +02:00
Sébastien Han 583791c45c core: move clusterInfo code to the controller package
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 11:05:02 +02:00
Sébastien Han 05506e7a68 core: reload go routine after CR is edited
Previously, the struct maintaining the list of cluster was still
initialized with a cluster item. Then the monitoring check will see that
the cluster is part of the struct already and thus won't run the
monitoring go routine again.
Now each time we cancel the context, we also remove the cluster item
from the map so that when the controller runs again, the monitoring
struct is re-populated and the go routine runs and statuses are updated.

Closes: https://github.com/rook/rook/issues/9911
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-03-31 08:49:42 +02:00
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
parth-gr 2dfd64a97c core: add observedGeneration to CR status
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller

Closes: https://github.com/rook/rook/issues/9673

Signed-off-by: parth-gr <paarora@redhat.com>
2022-03-16 19:58:35 +05:30
Blaine Gardner c92c6fbc36 core: rework usage of ReportReconcileResult
ReportReconcileResult should never be given an object that is nil. It is
impossible to force compile-time checking for this because the object's
type is an interface. We can't force this, but we can rework the
function definition to handle more cases where the object may not be
returned fully-complete, and we can rework callers to return structs
(not pointers-to-structs) so it is less likely a caller will pass nil
to the function.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-03-10 16:53:25 -07:00
Sébastien Han 7f32461b8f rgw: fix variable assignment
We need to explicitly assign the ceph version within each scope since
one function returns a pointer and one is not.
Previously, the `desiredCephVersion` was recreated in the scope of the
`else` and thus a new variable `desiredCephVersion` was initialized (see
the `:=`).
Now, each scope creates and assign to clusterInfo a `desiredCephVersion`
value.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-02-14 12:42:30 +01:00
Sébastien Han 856b14e60a object: do not check for upgrade on external mode
The object child controller should not compare versions if the object
store is external. The current version comparison is using the
cmdreporter to check the ceph version of the image in the CephCluster spec.
In external mode, this image is not set so the cmdreporter fails.
Also, this check is only valid for converged mode, so external mode
should be skipped.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-02-10 17:41:01 +01:00
Sébastien Han 5404ec13a2 core: dereference pointer before trying to compare with deepequal
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.

Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-27 17:41:42 +01:00
Blaine Gardner 9dc41f5796 Merge pull request #9427 from BlaineEXE/always-report-reconcile-events
operator: always report events
2021-12-20 08:42:08 -07:00
Blaine Gardner da61ac1ae8 operator: always report events
The original PR which added event reporting unnecessarily "optimized" to
prevent spamming the API controller[1].

It is sometimes important to get events as they happen and not hide new
events behind preexisting older events. For example, in integration
tests, we may often want to wait for a controller to finish processing
an update, and the best way to do that is to wait for the
"ReconcileSucceeded" event. In order for this to be useful, the events
must be reported each time.

If we begin having problems with events being reported too often, then
we should fix the underlying issue of reconciles happening too often
instead of relying on a time-based "optimization" that hides recent
event reports that may be useful.

[1]: https://github.com/rook/rook/pull/7222

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-15 15:24:17 -07:00
Blaine Gardner 0c76be8c38 pool: file: object: clean up health checkers for force deletion
When CephBlockPool, CephFilesystem, or CephObjectStore resources are
deleted after removing their finalizer, the code path to stop monitoring
was not stopping monitoring since a non-present resource does not have a
name and namespace attached. When the object is deleted, ensure the
internal representation used to stop monitoring has a name and namespace
to fix the issue.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-13 16:07:01 -07:00
Blaine Gardner 03ba7dec64 pool: file: object: clean up stop health checkers
Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.

Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-19 10:29:12 -07:00
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Blaine Gardner 2f850b6ae6 Merge pull request #8613 from subhamkrai/remove-nautilus
ceph: remove ceph nautilus, ceph octopus to default
2021-10-25 09:09:18 -06:00
Sébastien Han fc9c4b6fc9 Merge pull request #9010 from thotz/rgw-external-endpoint
ceph: update endpoint with IP for external RGW server
2021-10-25 15:49:08 +02:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
Jiffin Tony Thottan d4562f6b83 ceph: update endpoint with IP for external RGW server
For external RGW server use the IP mentioned in Gateway for admin Ops
operattions.

Fixes: #8916
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-10-21 19:13:10 +05:30
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
PixelJonas 8733bf6272 ceph: use normal quotation marks for comment
replaces all occurences of lduo/rduo quotation marks to make
sure that using the snippets in a k8s manifest will work

This fixes an issue with ArgoCD not being able
to apply `common.yaml` because of an encoding issue

Signed-off-by: PixelJonas <jonas@janz.digital>
2021-10-20 10:09:22 +02:00
Blaine Gardner eadcd757b3 rgw: replace period update --commit with function
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.

Resolves #8879

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 14:45:39 -06:00
Blaine Gardner 5383ba2df2 ceph: retry object health check if creation fails
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:24:18 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 2d55e69416 ceph: move scheme initialization to the same place
Let's initialize the schemes in a single place instead of doing it
when each controller initializes.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-06 11:33:10 +02:00
Sébastien Han daf53ea585 ceph: add unit test for object store user dependents
Now that we have the mock client for the admin ops api we can add unit
tests.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 10:01:53 +02:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Sébastien Han 677c93f3aa ceph: operate over object store pools concurrently
We now create and delete pools in parallel.
Before this patch, the deletion:

```
2021-06-08 13:48:26.557328 I | cephclient: no images/snapshosts present in pool "my-store.rgw.control"
2021-06-08 13:48:26.557383 I | cephclient: purging pool "my-store.rgw.control" (id=1)
2021-06-08 13:48:27.030021 I | ceph-object-controller: done disabling the dashboard api secret key
2021-06-08 13:48:28.800661 I | cephclient: purge completed for pool "my-store.rgw.control"
2021-06-08 13:48:29.085744 I | cephclient: no images/snapshosts present in pool "my-store.rgw.meta"
2021-06-08 13:48:29.085762 I | cephclient: purging pool "my-store.rgw.meta" (id=3)
2021-06-08 13:48:30.835487 I | cephclient: purge completed for pool "my-store.rgw.meta"
2021-06-08 13:48:31.128349 I | cephclient: no images/snapshosts present in pool "my-store.rgw.log"
2021-06-08 13:48:31.128368 I | cephclient: purging pool "my-store.rgw.log" (id=4)
2021-06-08 13:48:32.864376 I | cephclient: purge completed for pool "my-store.rgw.log"
2021-06-08 13:48:33.180923 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.index"
2021-06-08 13:48:33.181049 I | cephclient: purging pool "my-store.rgw.buckets.index" (id=5)
2021-06-08 13:48:34.892163 I | cephclient: purge completed for pool "my-store.rgw.buckets.index"
2021-06-08 13:48:35.179952 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.non-ec"
2021-06-08 13:48:35.179994 I | cephclient: purging pool "my-store.rgw.buckets.non-ec" (id=6)
2021-06-08 13:48:37.376067 I | cephclient: purge completed for pool "my-store.rgw.buckets.non-ec"
2021-06-08 13:48:37.665201 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.data"
2021-06-08 13:48:37.665218 I | cephclient: purging pool "my-store.rgw.buckets.data" (id=8)
2021-06-08 13:48:39.394717 I | cephclient: purge completed for pool "my-store.rgw.buckets.data"
2021-06-08 13:48:39.725528 I | cephclient: no images/snapshosts present in pool ".rgw.root"
2021-06-08 13:48:39.725567 I | cephclient: purging pool ".rgw.root" (id=7)
2021-06-08 13:48:41.424628 I | cephclient: purge completed for pool ".rgw.root"
```

It took 15sec to cleanup...
Now with this patch:

```
2021-06-08 14:15:25.621472 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.index"
2021-06-08 14:15:25.621570 I | cephclient: purging pool "my-store.rgw.buckets.index" (id=12)
2021-06-08 14:15:25.668708 I | cephclient: no images/snapshosts present in pool "my-store.rgw.log"
2021-06-08 14:15:25.668729 I | cephclient: purging pool "my-store.rgw.log" (id=11)
2021-06-08 14:15:25.693002 I | cephclient: no images/snapshosts present in pool "my-store.rgw.meta"
2021-06-08 14:15:25.693050 I | cephclient: purging pool "my-store.rgw.meta" (id=10)
2021-06-08 14:15:25.698830 I | cephclient: no images/snapshosts present in pool "my-store.rgw.control"
2021-06-08 14:15:25.698854 I | cephclient: purging pool "my-store.rgw.control" (id=9)
2021-06-08 14:15:25.701732 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.non-ec"
2021-06-08 14:15:25.701758 I | cephclient: purging pool "my-store.rgw.buckets.non-ec" (id=13)
2021-06-08 14:15:25.702816 I | cephclient: no images/snapshosts present in pool ".rgw.root"
2021-06-08 14:15:25.702836 I | cephclient: purging pool ".rgw.root" (id=14)
2021-06-08 14:15:25.717144 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.data"
2021-06-08 14:15:25.717212 I | cephclient: purging pool "my-store.rgw.buckets.data" (id=15)
2021-06-08 14:15:26.393246 I | ceph-object-controller: done disabling the dashboard api secret key
```

It tool around 1sec.

The creation before this patch:

```

2021-06-08 16:47:10.669484 I | ceph-spec: adding finalizer "cephobjectstore.ceph.rook.io" on "my-store"
2021-06-08 16:47:10.677253 E | ceph-object-controller: failed to set object store "rook-ceph/my-store" status to "Progressing". failed to update object "my-store" status: Operation cannot be fulfilled on cephobjectstores.ceph.rook.io "my-store": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:10.682231 I | op-mon: parsing mon endpoints: b=10.111.63.108:6789,c=10.108.123.222:6789,a=10.100.223.211:6789
2021-06-08 16:47:10.985002 I | ceph-object-controller: reconciling object store deployments
2021-06-08 16:47:11.002364 I | ceph-object-controller: ceph object store gateway service running at 10.99.31.5
2021-06-08 16:47:11.002384 I | ceph-object-controller: reconciling object store pools
2021-06-08 16:47:14.336086 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.control"
2021-06-08 16:47:16.115072 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:16.138224 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:16.391914 I | cephclient: creating replicated pool my-store.rgw.control succeeded
2021-06-08 16:47:16.391943 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.control"
2021-06-08 16:47:20.390562 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.meta"
2021-06-08 16:47:22.423269 I | cephclient: creating replicated pool my-store.rgw.meta succeeded
2021-06-08 16:47:22.423294 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.meta"
2021-06-08 16:47:26.502874 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.log"
2021-06-08 16:47:28.555252 I | cephclient: creating replicated pool my-store.rgw.log succeeded
2021-06-08 16:47:28.555308 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.log"
2021-06-08 16:47:32.619546 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.index"
2021-06-08 16:47:34.663020 I | cephclient: creating replicated pool my-store.rgw.buckets.index succeeded
2021-06-08 16:47:34.663050 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.index"
2021-06-08 16:47:38.712886 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:47:40.748430 I | cephclient: creating replicated pool my-store.rgw.buckets.non-ec succeeded
2021-06-08 16:47:40.748467 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:47:44.830659 I | cephclient: setting pool property "compression_mode" to "none" on pool ".rgw.root"
2021-06-08 16:47:46.875834 I | cephclient: creating replicated pool .rgw.root succeeded
2021-06-08 16:47:46.875863 I | cephclient: setting pool property "pg_num_min" to "8" on pool ".rgw.root"
2021-06-08 16:47:50.933717 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.data"
2021-06-08 16:47:52.969461 I | cephclient: creating replicated pool my-store.rgw.buckets.data succeeded
2021-06-08 16:47:52.969549 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-06-08 16:47:53.585737 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-06-08 16:47:53.585777 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-06-08 16:47:53.585787 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-06-08 16:47:53.585802 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-06-08 16:47:53.930416 I | ceph-object-controller: setting rgw config flags
2021-06-08 16:47:53.931363 I | op-config: setting "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:47:54.201200 I | op-config: successfully set "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:47:54.201225 I | op-config: setting "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:47:54.454589 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:47:54.454614 I | op-config: setting "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:47:54.710218 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:47:54.710237 I | op-config: setting "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:47:54.969049 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:47:54.969069 I | op-config: setting "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:47:55.255491 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:47:55.255631 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-06-08 16:47:55.279316 I | ceph-object-controller: enabling rgw dashboard
2021-06-08 16:47:55.361208 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:56.225904 I | ceph-object-controller: setting the dashboard api secret key
2021-06-08 16:47:56.225984 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
2021-06-08 16:47:56.589452 I | ceph-object-controller: starting rgw healthcheck
2021-06-08 16:47:56.650045 I | ceph-object-controller: done setting the dashboard api secret key
```

It took 46sec, and I've seen it taking almost a 1min sometimes.
Now with this patch:

```
2021-06-08 16:51:35.259558 I | ceph-spec: adding finalizer "cephobjectstore.ceph.rook.io" on "my-store"
2021-06-08 16:51:35.270524 E | ceph-object-controller: failed to set object store "rook-ceph/my-store" status to "Progressing". failed to update object "my-store" status: Operation cannot be fulfilled on cephobjectstores.ceph.rook.io "my-store": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:51:35.274387 I | op-mon: parsing mon endpoints: b=10.111.63.108:6789,c=10.108.123.222:6789,a=10.100.223.211:6789
2021-06-08 16:51:35.599023 I | ceph-object-controller: reconciling object store deployments
2021-06-08 16:51:35.607256 I | ceph-object-controller: ceph object store gateway service running at 10.104.254.110
2021-06-08 16:51:35.607315 I | ceph-object-controller: reconciling object store pools
2021-06-08 16:51:39.337735 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:51:40.343290 I | cephclient: setting pool property "compression_mode" to "none" on pool ".rgw.root"
2021-06-08 16:51:40.346335 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.index"
2021-06-08 16:51:40.362670 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.meta"
2021-06-08 16:51:40.363445 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.log"
2021-06-08 16:51:40.364321 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.control"
2021-06-08 16:51:41.412870 I | cephclient: creating replicated pool my-store.rgw.buckets.non-ec succeeded
2021-06-08 16:51:41.412904 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:51:42.460439 I | cephclient: creating replicated pool my-store.rgw.control succeeded
2021-06-08 16:51:42.460503 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.control"
2021-06-08 16:51:42.470666 I | cephclient: creating replicated pool my-store.rgw.meta succeeded
2021-06-08 16:51:42.470727 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.meta"
2021-06-08 16:51:42.473029 I | cephclient: creating replicated pool .rgw.root succeeded
2021-06-08 16:51:42.473064 I | cephclient: setting pool property "pg_num_min" to "8" on pool ".rgw.root"
2021-06-08 16:51:42.473895 I | cephclient: creating replicated pool my-store.rgw.buckets.index succeeded
2021-06-08 16:51:42.473923 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.index"
2021-06-08 16:51:42.478625 I | cephclient: creating replicated pool my-store.rgw.log succeeded
2021-06-08 16:51:42.478671 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.log"
2021-06-08 16:51:46.606905 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.data"
2021-06-08 16:51:48.627320 I | cephclient: creating replicated pool my-store.rgw.buckets.data succeeded
2021-06-08 16:51:48.627368 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-06-08 16:51:49.325086 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-06-08 16:51:49.325108 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-06-08 16:51:49.325121 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-06-08 16:51:49.325134 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-06-08 16:51:49.657898 I | ceph-object-controller: setting rgw config flags
2021-06-08 16:51:49.657920 I | op-config: setting "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:51:49.917898 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:51:49.917917 I | op-config: setting "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:51:50.184939 I | op-config: successfully set "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:51:50.184970 I | op-config: setting "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:51:50.463799 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:51:50.463821 I | op-config: setting "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:51:50.720953 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:51:50.720999 I | op-config: setting "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:51:50.974875 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:51:50.975024 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-06-08 16:51:51.019820 I | ceph-object-controller: enabling rgw dashboard
2021-06-08 16:51:51.980117 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
2021-06-08 16:51:51.980199 I | ceph-object-controller: setting the dashboard api secret key
2021-06-08 16:51:52.399997 I | ceph-object-controller: done setting the dashboard api secret key
2021-06-08 16:51:52.435936 I | ceph-object-controller: starting rgw healthcheck
```

It took 17sec.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 16:09:22 +02:00
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Sébastien Han e478bdadb1 ceph: stop the monitoring before cleanup
Saw this today in the logs:

```
2021-06-02 13:02:54.787838 I | ceph-block-pool-controller: deleting pool "testpool"
2021-06-02 13:02:56.178094 I | cephclient: no images/snapshosts present in pool "testpool"
2021-06-02 13:02:56.178125 I | cephclient: purging pool "testpool" (id=17)
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0x1cef49f]

goroutine 1607 [running]:
github.com/rook/rook/pkg/operator/ceph/pool.(*mirrorChecker).checkMirroringHealth(0xc001a98ba0, 0x8, 0xc000d34c60)
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/health.go:115 +0xdf
github.com/rook/rook/pkg/operator/ceph/pool.(*mirrorChecker).checkMirroring(0xc001a98ba0, 0xc0016fd500)
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/health.go:82 +0x147
created by github.com/rook/rook/pkg/operator/ceph/pool.(*ReconcileCephBlockPool).reconcile
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/controller.go:298 +0xee5
```

Essentially, it's intermittent but when deleting the pool the
healthcheck kicked in and fetched the mirroring status, which returned
empty. THe subsequent code tried to access content of a nil pointer,
  hence the error.
So now we stop monitoring first, then we proceed with the deletion.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-02 16:14:10 +02:00
Jiffin Tony Thottan aea21d9c84 ceph: service server cert support for rgw
Service serving certificates are intended to applications that require
encryption in openshift. These certificates are issued as TLS web server
certificates. Currently RGW supports TLS authentication with help of
certs passed as secrets, in this case we add following details as
`service.annotations` in the Objectstore Gateway Spec :
```
service:
  annotations:
    service.beta.openshift.io/serving-cert-secret-name: <name for
autogenerated secret>
```

More details about service serving cert can be found at :
https://docs.openshift.com/container-platform/4.6/security/certificates/service-serving-certificate.html

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-26 18:15:12 +05:30
Jiffin Tony Thottan 70f94a2d3c ceph: removing serviceIP from bucketChecker struct
serviceIP is not used by the health bucket checker, hence removing it.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-20 14:24:11 +05:30
Jiffin Tony Thottan bcbe707bcb ceph: fix healthcheck incase ssl enabled for rgw
In case ssl enabled for RGW, bucket healthcheck won't work since access is denied from the server.
In that case configure S3 agent using "InsecureSkipVerify: true" option.

Fixes: 7288
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-20 14:23:48 +05:30
Lars Lehtonen f1b341ba23 ceph: fix multiple imports
This fixes libraries that were being imported multiple times in
pkg/operator/ceph/object and its subpackages.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-12 02:57:13 -07:00
Sébastien Han 3ad85ad4f6 ceph: force pg_autoscaler mode for rgw pools
Let's make sure that Nautilus clients are using the pg_autoscaler when
setting up the cluster. We have seen cases where this was not enforced
after a re-installation.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-31 11:00:12 +02:00
Sébastien Han 09cb5d6b4a ceph: remove unsed type field
isExternal was not used anymore so let's remove it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-23 10:31:37 +01:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00