Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.
Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
For external RGW server use the IP mentioned in Gateway for admin Ops
operattions.
Fixes: #8916
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
replaces all occurences of lduo/rduo quotation marks to make
sure that using the snippets in a k8s manifest will work
This fixes an issue with ArgoCD not being able
to apply `common.yaml` because of an encoding issue
Signed-off-by: PixelJonas <jonas@janz.digital>
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.
Resolves#8879
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
We now create and delete pools in parallel.
Before this patch, the deletion:
```
2021-06-08 13:48:26.557328 I | cephclient: no images/snapshosts present in pool "my-store.rgw.control"
2021-06-08 13:48:26.557383 I | cephclient: purging pool "my-store.rgw.control" (id=1)
2021-06-08 13:48:27.030021 I | ceph-object-controller: done disabling the dashboard api secret key
2021-06-08 13:48:28.800661 I | cephclient: purge completed for pool "my-store.rgw.control"
2021-06-08 13:48:29.085744 I | cephclient: no images/snapshosts present in pool "my-store.rgw.meta"
2021-06-08 13:48:29.085762 I | cephclient: purging pool "my-store.rgw.meta" (id=3)
2021-06-08 13:48:30.835487 I | cephclient: purge completed for pool "my-store.rgw.meta"
2021-06-08 13:48:31.128349 I | cephclient: no images/snapshosts present in pool "my-store.rgw.log"
2021-06-08 13:48:31.128368 I | cephclient: purging pool "my-store.rgw.log" (id=4)
2021-06-08 13:48:32.864376 I | cephclient: purge completed for pool "my-store.rgw.log"
2021-06-08 13:48:33.180923 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.index"
2021-06-08 13:48:33.181049 I | cephclient: purging pool "my-store.rgw.buckets.index" (id=5)
2021-06-08 13:48:34.892163 I | cephclient: purge completed for pool "my-store.rgw.buckets.index"
2021-06-08 13:48:35.179952 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.non-ec"
2021-06-08 13:48:35.179994 I | cephclient: purging pool "my-store.rgw.buckets.non-ec" (id=6)
2021-06-08 13:48:37.376067 I | cephclient: purge completed for pool "my-store.rgw.buckets.non-ec"
2021-06-08 13:48:37.665201 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.data"
2021-06-08 13:48:37.665218 I | cephclient: purging pool "my-store.rgw.buckets.data" (id=8)
2021-06-08 13:48:39.394717 I | cephclient: purge completed for pool "my-store.rgw.buckets.data"
2021-06-08 13:48:39.725528 I | cephclient: no images/snapshosts present in pool ".rgw.root"
2021-06-08 13:48:39.725567 I | cephclient: purging pool ".rgw.root" (id=7)
2021-06-08 13:48:41.424628 I | cephclient: purge completed for pool ".rgw.root"
```
It took 15sec to cleanup...
Now with this patch:
```
2021-06-08 14:15:25.621472 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.index"
2021-06-08 14:15:25.621570 I | cephclient: purging pool "my-store.rgw.buckets.index" (id=12)
2021-06-08 14:15:25.668708 I | cephclient: no images/snapshosts present in pool "my-store.rgw.log"
2021-06-08 14:15:25.668729 I | cephclient: purging pool "my-store.rgw.log" (id=11)
2021-06-08 14:15:25.693002 I | cephclient: no images/snapshosts present in pool "my-store.rgw.meta"
2021-06-08 14:15:25.693050 I | cephclient: purging pool "my-store.rgw.meta" (id=10)
2021-06-08 14:15:25.698830 I | cephclient: no images/snapshosts present in pool "my-store.rgw.control"
2021-06-08 14:15:25.698854 I | cephclient: purging pool "my-store.rgw.control" (id=9)
2021-06-08 14:15:25.701732 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.non-ec"
2021-06-08 14:15:25.701758 I | cephclient: purging pool "my-store.rgw.buckets.non-ec" (id=13)
2021-06-08 14:15:25.702816 I | cephclient: no images/snapshosts present in pool ".rgw.root"
2021-06-08 14:15:25.702836 I | cephclient: purging pool ".rgw.root" (id=14)
2021-06-08 14:15:25.717144 I | cephclient: no images/snapshosts present in pool "my-store.rgw.buckets.data"
2021-06-08 14:15:25.717212 I | cephclient: purging pool "my-store.rgw.buckets.data" (id=15)
2021-06-08 14:15:26.393246 I | ceph-object-controller: done disabling the dashboard api secret key
```
It tool around 1sec.
The creation before this patch:
```
2021-06-08 16:47:10.669484 I | ceph-spec: adding finalizer "cephobjectstore.ceph.rook.io" on "my-store"
2021-06-08 16:47:10.677253 E | ceph-object-controller: failed to set object store "rook-ceph/my-store" status to "Progressing". failed to update object "my-store" status: Operation cannot be fulfilled on cephobjectstores.ceph.rook.io "my-store": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:10.682231 I | op-mon: parsing mon endpoints: b=10.111.63.108:6789,c=10.108.123.222:6789,a=10.100.223.211:6789
2021-06-08 16:47:10.985002 I | ceph-object-controller: reconciling object store deployments
2021-06-08 16:47:11.002364 I | ceph-object-controller: ceph object store gateway service running at 10.99.31.5
2021-06-08 16:47:11.002384 I | ceph-object-controller: reconciling object store pools
2021-06-08 16:47:14.336086 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.control"
2021-06-08 16:47:16.115072 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:16.138224 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:16.391914 I | cephclient: creating replicated pool my-store.rgw.control succeeded
2021-06-08 16:47:16.391943 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.control"
2021-06-08 16:47:20.390562 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.meta"
2021-06-08 16:47:22.423269 I | cephclient: creating replicated pool my-store.rgw.meta succeeded
2021-06-08 16:47:22.423294 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.meta"
2021-06-08 16:47:26.502874 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.log"
2021-06-08 16:47:28.555252 I | cephclient: creating replicated pool my-store.rgw.log succeeded
2021-06-08 16:47:28.555308 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.log"
2021-06-08 16:47:32.619546 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.index"
2021-06-08 16:47:34.663020 I | cephclient: creating replicated pool my-store.rgw.buckets.index succeeded
2021-06-08 16:47:34.663050 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.index"
2021-06-08 16:47:38.712886 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:47:40.748430 I | cephclient: creating replicated pool my-store.rgw.buckets.non-ec succeeded
2021-06-08 16:47:40.748467 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:47:44.830659 I | cephclient: setting pool property "compression_mode" to "none" on pool ".rgw.root"
2021-06-08 16:47:46.875834 I | cephclient: creating replicated pool .rgw.root succeeded
2021-06-08 16:47:46.875863 I | cephclient: setting pool property "pg_num_min" to "8" on pool ".rgw.root"
2021-06-08 16:47:50.933717 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.data"
2021-06-08 16:47:52.969461 I | cephclient: creating replicated pool my-store.rgw.buckets.data succeeded
2021-06-08 16:47:52.969549 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-06-08 16:47:53.585737 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-06-08 16:47:53.585777 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-06-08 16:47:53.585787 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-06-08 16:47:53.585802 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-06-08 16:47:53.930416 I | ceph-object-controller: setting rgw config flags
2021-06-08 16:47:53.931363 I | op-config: setting "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:47:54.201200 I | op-config: successfully set "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:47:54.201225 I | op-config: setting "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:47:54.454589 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:47:54.454614 I | op-config: setting "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:47:54.710218 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:47:54.710237 I | op-config: setting "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:47:54.969049 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:47:54.969069 I | op-config: setting "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:47:55.255491 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:47:55.255631 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-06-08 16:47:55.279316 I | ceph-object-controller: enabling rgw dashboard
2021-06-08 16:47:55.361208 E | ceph-crashcollector-controller: node reconcile failed on op "unchanged": Operation cannot be fulfilled on deployments.apps "rook-ceph-crashcollector-minikube": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:47:56.225904 I | ceph-object-controller: setting the dashboard api secret key
2021-06-08 16:47:56.225984 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
2021-06-08 16:47:56.589452 I | ceph-object-controller: starting rgw healthcheck
2021-06-08 16:47:56.650045 I | ceph-object-controller: done setting the dashboard api secret key
```
It took 46sec, and I've seen it taking almost a 1min sometimes.
Now with this patch:
```
2021-06-08 16:51:35.259558 I | ceph-spec: adding finalizer "cephobjectstore.ceph.rook.io" on "my-store"
2021-06-08 16:51:35.270524 E | ceph-object-controller: failed to set object store "rook-ceph/my-store" status to "Progressing". failed to update object "my-store" status: Operation cannot be fulfilled on cephobjectstores.ceph.rook.io "my-store": the object has been modified; please apply your changes to the latest version and try again
2021-06-08 16:51:35.274387 I | op-mon: parsing mon endpoints: b=10.111.63.108:6789,c=10.108.123.222:6789,a=10.100.223.211:6789
2021-06-08 16:51:35.599023 I | ceph-object-controller: reconciling object store deployments
2021-06-08 16:51:35.607256 I | ceph-object-controller: ceph object store gateway service running at 10.104.254.110
2021-06-08 16:51:35.607315 I | ceph-object-controller: reconciling object store pools
2021-06-08 16:51:39.337735 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:51:40.343290 I | cephclient: setting pool property "compression_mode" to "none" on pool ".rgw.root"
2021-06-08 16:51:40.346335 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.index"
2021-06-08 16:51:40.362670 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.meta"
2021-06-08 16:51:40.363445 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.log"
2021-06-08 16:51:40.364321 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.control"
2021-06-08 16:51:41.412870 I | cephclient: creating replicated pool my-store.rgw.buckets.non-ec succeeded
2021-06-08 16:51:41.412904 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.non-ec"
2021-06-08 16:51:42.460439 I | cephclient: creating replicated pool my-store.rgw.control succeeded
2021-06-08 16:51:42.460503 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.control"
2021-06-08 16:51:42.470666 I | cephclient: creating replicated pool my-store.rgw.meta succeeded
2021-06-08 16:51:42.470727 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.meta"
2021-06-08 16:51:42.473029 I | cephclient: creating replicated pool .rgw.root succeeded
2021-06-08 16:51:42.473064 I | cephclient: setting pool property "pg_num_min" to "8" on pool ".rgw.root"
2021-06-08 16:51:42.473895 I | cephclient: creating replicated pool my-store.rgw.buckets.index succeeded
2021-06-08 16:51:42.473923 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.buckets.index"
2021-06-08 16:51:42.478625 I | cephclient: creating replicated pool my-store.rgw.log succeeded
2021-06-08 16:51:42.478671 I | cephclient: setting pool property "pg_num_min" to "8" on pool "my-store.rgw.log"
2021-06-08 16:51:46.606905 I | cephclient: setting pool property "compression_mode" to "none" on pool "my-store.rgw.buckets.data"
2021-06-08 16:51:48.627320 I | cephclient: creating replicated pool my-store.rgw.buckets.data succeeded
2021-06-08 16:51:48.627368 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-06-08 16:51:49.325086 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-06-08 16:51:49.325108 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-06-08 16:51:49.325121 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-06-08 16:51:49.325134 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-06-08 16:51:49.657898 I | ceph-object-controller: setting rgw config flags
2021-06-08 16:51:49.657920 I | op-config: setting "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:51:49.917898 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_object_name_utc"="true" option to the mon configuration database
2021-06-08 16:51:49.917917 I | op-config: setting "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:51:50.184939 I | op-config: successfully set "client.rgw.my.store.a"="rgw_enable_usage_log"="true" option to the mon configuration database
2021-06-08 16:51:50.184970 I | op-config: setting "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:51:50.463799 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zone"="my-store" option to the mon configuration database
2021-06-08 16:51:50.463821 I | op-config: setting "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:51:50.720953 I | op-config: successfully set "client.rgw.my.store.a"="rgw_zonegroup"="my-store" option to the mon configuration database
2021-06-08 16:51:50.720999 I | op-config: setting "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:51:50.974875 I | op-config: successfully set "client.rgw.my.store.a"="rgw_log_nonexistent_bucket"="true" option to the mon configuration database
2021-06-08 16:51:50.975024 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-06-08 16:51:51.019820 I | ceph-object-controller: enabling rgw dashboard
2021-06-08 16:51:51.980117 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
2021-06-08 16:51:51.980199 I | ceph-object-controller: setting the dashboard api secret key
2021-06-08 16:51:52.399997 I | ceph-object-controller: done setting the dashboard api secret key
2021-06-08 16:51:52.435936 I | ceph-object-controller: starting rgw healthcheck
```
It took 17sec.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.
Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
Saw this today in the logs:
```
2021-06-02 13:02:54.787838 I | ceph-block-pool-controller: deleting pool "testpool"
2021-06-02 13:02:56.178094 I | cephclient: no images/snapshosts present in pool "testpool"
2021-06-02 13:02:56.178125 I | cephclient: purging pool "testpool" (id=17)
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0x1cef49f]
goroutine 1607 [running]:
github.com/rook/rook/pkg/operator/ceph/pool.(*mirrorChecker).checkMirroringHealth(0xc001a98ba0, 0x8, 0xc000d34c60)
/home/runner/work/rook/rook/pkg/operator/ceph/pool/health.go:115 +0xdf
github.com/rook/rook/pkg/operator/ceph/pool.(*mirrorChecker).checkMirroring(0xc001a98ba0, 0xc0016fd500)
/home/runner/work/rook/rook/pkg/operator/ceph/pool/health.go:82 +0x147
created by github.com/rook/rook/pkg/operator/ceph/pool.(*ReconcileCephBlockPool).reconcile
/home/runner/work/rook/rook/pkg/operator/ceph/pool/controller.go:298 +0xee5
```
Essentially, it's intermittent but when deleting the pool the
healthcheck kicked in and fetched the mirroring status, which returned
empty. THe subsequent code tried to access content of a nil pointer,
hence the error.
So now we stop monitoring first, then we proceed with the deletion.
Signed-off-by: Sébastien Han <seb@redhat.com>
Service serving certificates are intended to applications that require
encryption in openshift. These certificates are issued as TLS web server
certificates. Currently RGW supports TLS authentication with help of
certs passed as secrets, in this case we add following details as
`service.annotations` in the Objectstore Gateway Spec :
```
service:
annotations:
service.beta.openshift.io/serving-cert-secret-name: <name for
autogenerated secret>
```
More details about service serving cert can be found at :
https://docs.openshift.com/container-platform/4.6/security/certificates/service-serving-certificate.html
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
In case ssl enabled for RGW, bucket healthcheck won't work since access is denied from the server.
In that case configure S3 agent using "InsecureSkipVerify: true" option.
Fixes: 7288
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
This fixes libraries that were being imported multiple times in
pkg/operator/ceph/object and its subpackages.
Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
Let's make sure that Nautilus clients are using the pg_autoscaler when
setting up the cluster. We have seen cases where this was not enforced
after a re-installation.
Signed-off-by: Sébastien Han <seb@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Sometimes `radosgw-admin` succeeds after showing logs to stderr. We should
skip non-json strings if parsing output as json.
Here is an example.
```
2021-02-26 04:10:44.190418 I | op-bucket-prov: creating Ceph user "ceph-user-aSzNqgE7"
E0226 04:11:37.901310 8 controller.go:199] error syncing 'logging/loki-bucket': error provisioning bucket: Provision: can't create ceph user: error creating ceph user "ceph-user-aSzNqgE7": failed to unmarshal json. 2021-02-26T04:11:21.425+0000 7f6714be4980 1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:21.426+0000 7f6714be4980 0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.uid:ceph-user-aSzNqgE7
2021-02-26T04:11:32.168+0000 7f6714be4980 1 robust_notify: If at first you don't succeed: (110) Connection timed out
2021-02-26T04:11:32.168+0000 7f6714be4980 0 ERROR: failed to distribute cache for ceph-hdd-object-store.rgw.meta:users.keys:23Z8GUEXR0TJDO86PSJR
{
"user_id": "ceph-user-aSzNqgE7",
"display_name": "ceph-user-aSzNqgE7",
...
"mfa_ids": []
}: invalid character '-' after top-level value: failed to unmarshal json.
```
In this case, some logs like "robust_notify:..." was shown in stderr.
Unmarsharing was failed due to tried to parse these logs as json.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
If the ceph cli outputs an error with "error calling conf_read_file"
this means that the operator has not written its ceph configuration
file. Thus ceph cli commands will fail, so we can just ignore that since
the operator will soon write this file in its initialization sequence.
Signed-off-by: Sébastien Han <seb@redhat.com>
When creating object store, `radosgw-admin realm get ..` command is stuck forever when required number of OSDs are not available. Because of this the uninstall of object store is also stuck. User has to manually remove the finalizer to delete the object store.This PR uses `ExecuteCommandWithTimeout` for running `radosgw-admin` command. Timeout during installation will be reconciled. Cleanup will be treated as best effort. Any errors during uninstalling of single site object store will only be logged.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.
Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
After #6217, internal rgw spawning on external cluster was still broken,
operator claiming:
```
ceph-object-controller: failed to reconcile failed to create object
store deployments: failed to reconcile external endpoint:
failed to create or update object store "arch-cloud" endpoint:
failed to create endpoint "rook-ceph-rgw-arch-cloud". Endpoints
"rook-ceph-rgw-arch-cloud" is invalid: subsets[0]:
Required value: must specify `addresses` or `notReadyAddresses`
```
To spawn internal rgw for external cluster, changes detection of
what mean 'internal' for rgw and only use the externalRgwEndpoints
list for that independently of the status of the cluster
Now there is multiple posibilities:
* internal cluster and internal rgw pods: "normal case"
* external cluster and internal rgw pods: <= now working with the PR
* external cluster and external rgw pods: the external case
* internal cluster and external rgw pods: <= new case that could exist
Signed-off-by: Julien Girardin <jugirardin@free.fr>
Update to the latest lib bucket provisioner code.
Fixes issue 6650
Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.
Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.
Does not reintroduce bug 6767 from previous fix for 6650
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.
Signed-off-by: Sébastien Han <seb@redhat.com>
The RGW deployment's ceph-version label should now display the same
version as the image it is using.
It previously displayed an empty version: 0.0.0-0.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
Pools in stretch clusters will all use the same crush rule that
is generated by the operator when configuring the cluster for
stretch mode, with replica 4 and two replicas per failure domain.
Pools cannot create new crush rules, so we require that replica be
4 when the pools is created, EC is not allowed, and no new
crush rule will be created.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.
By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.
To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.
The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.
Fixes#4993.
Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
Add unit tests for when the "Zone" is configured
in an object store so that the object store joins
the CephObjectZone and the zone's corresponding
multisite configuration.
Signed-off-by: Ali Maredia <amaredia@redhat.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
currently, cephobjectstore monitoring is never
stopped for the following flow:
1. objectstore is deleted
2. go routine is stopped
3. objectstore is created
4. objectstore is deleted
so, removing clusterResourceDeleted flag and
working with the map only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
currently, the object store schema requires the
port to be non-zero. This does not allow http
access to the object store to be disabled.
so setting the port minimum to 0 to from 1 in the cr.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
Instead of setting the Phase of the CR once the reconcile is done, let's
actually set in from the healthcheck so it is more accurate.
Closes: https://github.com/rook/rook/issues/5249
Signed-off-by: Sébastien Han <seb@redhat.com>
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.
- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up
- Added more debug logging all around the object-store
code related to multisite
- files generated by rerun of `make codegen`
- change back the edit on the Copyright in object.go
Signed-off-by: Ali Maredia <amaredia@redhat.com>