Commit Graph
124 Commits
Author SHA1 Message Date
Blaine Gardner 17f0072d9d Merge pull request #12778 from BlaineEXE/multus-allow-cidr-spec
multus: allow using NADs without inspectable CIDRs
2023-09-07 13:42:41 -06:00
Blaine Gardner 3c43268d0a multus: detect network CIDRs via canary
Change how Rook detects network CIDRs for Multus networks. The IPAM
configuration is only defined as an arbitrary string JSON blob with a
"type" field and nothing more. Rook's detection of CIDRs for whereabouts
had already grown out of date since the initial implementation.
Additionally, Rook did not support DHCP IPAM, which is a reasonable
choice for users. And more, Rook did not support CNI plugin chaining,
which further complicates NADs. Based on the CNI spec, network chaning
can result in any changes to network CIDRs from the first-given plugin.

All these problems make it more and more difficult for Rook to support
Multus by inspecting the NAD itself to predict network CIDRs. Instead,
it is better for Rook to treat the CNI process as a black box. To
preserve legacy functionality of auto-detecting networks and to make
that as robust as possible, change to a canary-style architecture like
that used for Ceph mons, from which Rook will detect the network CIDRs
if possible.

Also allow users to specify overrides for CIDR ranges. This allows Rook
to still support esoteric and unexpected NAD or network configurations
where a CIDR range is not detectable or where the range detected would
be incomplete. Because it may be impossible for Rook to understand the
network CIDRs wholistically while residing only on a portion of the
network, this feature should have been present from Multus's inception.

Improving CIDR auto-detection and allowing users to specify overrides
for auto-detected CIDRs rounds out Rook's Multus support for CephCluster
(core/RADOS) installations. No further architectural changes should be
needed for CephClusters as regards application of public/cluster network
CIDRs for Multus networks.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-09-07 10:12:55 -06:00
subhamkrai 29d2b6a071 core: restart ceph daemons when network updated
We need to restart all the ceph daemons whenever
cephCluster network settings are modified like
requiremsgr2, encryption and compression. This
required for Ceph to consider the new settings
it require new ceph daemons all over.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-01 11:30:48 +05:30
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
Avan Thakkar 460900756c core: add service monitor for ceph-exporter service
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2023-02-15 15:44:56 +05:30
Avan Thakkar 2f8ee60374 core: introduce ceph-exporter
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2023-02-06 02:47:09 -05:00
Blaine Gardner 17c5cf7384 object: update startup, liveness, readiness probes
Based on the latest info from the Ceph RGW team, update all probes. Get
rid of the liveness probe entirely. Update startup and readiness probes
to support return codes for misconfiguration (500) and for server-side
throttling (498 or 503).

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-12-20 15:23:22 -07:00
Blaine Gardner a777b1d7d1 object: do not create service for external object stores
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.

This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.

The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-04 17:30:30 -06:00
Blaine Gardner a7c0c7ee93 object: remove health checker
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.

Full discussion: https://github.com/rook/rook/issues/11031

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-10-18 13:41:46 -06:00
Avan Thakkar 934aa91056 operator: make imagePullPolicy customizable for csi driver and ceph pods
Introduce a new env variable ROOK_CSI_IMAGE_PULL_POLICY in rook operator configmap which should be used to
customize the imagePullPolicy for the csi driver and imagePullPolicy property in cephVersionSpec for ceph pods.

Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2022-09-27 11:57:43 +05:30
zhucan b94f080548 object: gateway.port partially ignored when hostNetwork disabled
Signed-off-by: zhucan <zhucan.k8s@gmail.com>
2022-08-31 10:19:22 +08:00
Jiffin Tony Thottan 8003e764f9 rgw: add custom endpoint list option for zone
User can define his desired endpoint list in Zone CR so that it will
overwrite the default service name for rgw.

Resolves #6432

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2022-08-17 10:03:12 +05:30
Jiffin Tony Thottan 5c8ca01bd0 object: adding support for sse s3 for RGW
The RGW support server side encryption with help of s3 protocol, till
now the `sse:kms` was support in which keys will be provided by the user
and but it will be saved in external management service like vault. Now
the support for `sse:s3` is added so the entire encryption key
management is performed by RGW itsels.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-07-28 11:50:48 +05:30
Travis Nielsen 37a4d9e959 Merge pull request #10491 from zhucan/feat-10451
object: network mode can be set separately for cephcluster and rgw
2022-07-14 11:21:26 -06:00
Travis Nielsen dad97f3425 core: remove support for ceph octopus
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-07 15:03:26 -06:00
zhucan 5ee75d6652 object: support separated network for objectstore
Signed-off-by: zhucan <zhucan.k8s@gmail.com>
2022-06-30 10:56:03 +08:00
Alexander Trost 8686296e17 core: remove double imported packages
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-25 13:51:45 +02:00
subhamkrai bf7daccf60 core: make code changes to support latest cntrl runtime
making necessary code changes to support controller
runtime version.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-20 22:30:15 +05:30
subhamkrai 24802c559e core: fix golangci linter
fix golangci linter

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-04 20:59:31 +05:30
Jiffin Tony Thottan 5e72b26948 object: add service account for RGW pod
For supporting features like service account authentication for vault
KMS , a service account account need to attach with pod.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-04-04 11:24:50 +05:30
parth-gr 2cf89b11df rgw: use ceph config probes for rgw
use the common method ConfigureStartupProbe and
ConfigureLivenes probe for rgw probes configuration

Signed-off-by: parth-gr <paarora@redhat.com>
2022-02-22 20:36:40 +05:30
Jiffin Tony Thottan 0b4cdb992b object: fix backend path for transit engine for rgw kms
The backend path was added with additional `transit` to it.
Also added PR test case to check transit engine

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-02-15 18:49:31 +05:30
micalgenus a68151e2f2 rgw: gateway deployment strategy rolling on Pacific
When the strategy type of deployment is Recreate, from the moment the end is
started until the new pod is normally running, there is no pod in the
kubernetes service endpoints, so it goes to the service unavailable state.

Therefore, it is necessary to change it to `type: RollingUpdate` instead of
`type: Recreate` so that rolling updates can be made without downtime.

Before Pacific, creating a deployment as much as the instance was okay,
but after Pacific, the replicas was increased in one deployment, so rolling
updates were set up one by one as in the previous version.

Signed-off-by: micalgenus <micalgenus@gmail.com>
2022-01-21 00:27:37 +09:00
Sébastien Han a7acf490b6 osd: add ibm key protect kms
The cluster-wide encryption feature now has a new supported key
management system: IBM Key protect. More information about the backend
can be found in Rook documentation under the

Today's implementation stores OSD encryption keys has "Standard" keys.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-13 16:02:10 +01:00
Satoru Takeuchi af88b50fc7 rgw: fix startup probe
It's better to set the same handler to startupProbe as livenessProbe.
Otherwise, we might hit the following problem.

https://github.com/rook/rook/issues/6304

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-01-06 06:20:53 +00:00
Blaine Gardner c07d89d9ea core: rgw: allow specifying daemon startup probes
Allow specifying daemon startup probes where we also allow configuring
liveness probes. Startup probes allow Rook to tolerate when Ceph daemons
occasionally take a long time to start up while not also making
Kubernetes liveness probes slower to detect runtime failures of daemons.

Startup probes are beta in Kubernetes 1.18, so we should not enable
probes by default for earlier Kubernetes versions.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-21 15:12:41 -07:00
Sébastien Han 870114fc91 Merge pull request #9393 from y1r/add-context-k8sutil-endpoint
core: add context parameter to k8sutil endpoint
2021-12-13 10:04:19 +01:00
Yuichiro Ueno 03e976b52c core: add context parameter to k8sutil service
This commit adds context parameter to k8sutil service functions. By
this, we can handle cancellation during API call of service resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-12-11 11:40:28 +09:00
Yuichiro Ueno d1d252c5c2 core: add context parameter to k8sutil endpoint
This commit adds context parameter to k8sutil endpoint functions. By
this, we can handle cancellation during API call of endpoint resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-12-11 11:26:21 +09:00
parth-gr 0a86d26b2e core: create rook resources with k8s recommended labels
Adding Recommended Labels on the resources created by rook
    and using Recommended Labels in the helm chart,
    for better visuals and management of k8s object

Closes: https://github.com/rook/rook/issues/8400
Signed-off-by: parth-gr <paarora@redhat.com>
2021-12-07 18:16:44 +05:30
Jiffin Tony Thottan aba50d3ca9 object: add support in RGW to communicate vault with TLS
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-17 10:19:28 +05:30
Yuzuki Mimura 536b59ef0f rgw: change the way to livenessProbe and introduce readinessProbe
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.

Closes: #8407

Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-10-29 15:07:29 +00:00
parth-gr fc7905a7bd ceph: fixing ClientID of log-collector for RGW instance
The Client_ID generated by operator was
different from the log rotate file created
The Clinet_ID= rgwceph.client.rook.ceph.rgw.my.store.a
and log file name= ceph-client.rgw.my.store.a.log
So changed the CLient_ID to ceph-client.rgw.my.store.a for
correct working and this follow the patterns how other modules
Client_ID is generated

Closes: https://github.com/rook/rook/issues/8692
Signed-off-by: parth-gr <paarora@redhat.com>
2021-10-07 19:30:21 +05:30
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han d675969567 ceph: fix vault kv secret engine auto-detection
Passing struct by value essentially gives you a copy, so when modified
within a function, the scope is then reduced to that function. Using
pointers solves that you mutate the struct as many times as you want from
anywhere.
As a result, the auto-detection of the Vault KV backend was not working
correctly.
Also, added a ton of unit tests for Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-31 15:29:04 +02:00
Denis Egorenko fb04908315 ceph: add ability to specify ca bundle for rgw
Specify ca Bundle for RGW spec and mount inside pods.

Related-Issue: https://github.com/rook/rook/issues/8490
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
2021-08-10 23:18:26 +04:00
Sébastien Han 99e00dea1e ceph: auto detect vault k/v version
Rook will now auto detect the kv version of the vault server. This
allows users not having to pass the VAULT_BACKEND configuration in the
CephCluster CR.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-29 10:16:31 +02:00
Jiffin Tony Thottan 1665ad6ea7 ceph: add support for tls certs via k8s tls secrets for rgw
With this PR the RGW can accept TLS certs as K8s TLS secrets

Fixes: 2079
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-07-06 16:30:16 +05:30
Jiffin Tony Thottan d909c309f3 ceph: update the backend path for transit engine
The backend has for trasnit secret in vault kms for RGW is changed from pacific

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-05-27 22:48:21 +05:30
Travis Nielsen b0a63711f5 build: refactor to consolidate the rook.io/v1 package
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-18 19:55:37 -06:00
Jiffin Tony Thottan a5a5661e5f ceph: add security spec to object store spec
Add `SecuritySpec` to specify kms details for rgw in `ObjectStoreSpec` than using existing
`SecuritySpec` used for `ClusterSpec`. The vault secret engines conflict with OSD and RGW
kms encrytpion

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-28 14:39:39 +05:30
Jiffin Tony Thottan 3ef3e8b72e ceph: check for kv engine version for SSE RGW
The "https://docs.ceph.com/en/latest/radosgw/vault/" says RGW supports only version v2 of
kv engine, so pod won't start if the admin didn't provide "VAULT_BACKEND" in security.kms.ConnectionDetails

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-28 13:03:22 +05:30
Jiffin Tony Thottan aea21d9c84 ceph: service server cert support for rgw
Service serving certificates are intended to applications that require
encryption in openshift. These certificates are issued as TLS web server
certificates. Currently RGW supports TLS authentication with help of
certs passed as secrets, in this case we add following details as
`service.annotations` in the Objectstore Gateway Spec :
```
service:
  annotations:
    service.beta.openshift.io/serving-cert-secret-name: <name for
autogenerated secret>
```

More details about service serving cert can be found at :
https://docs.openshift.com/container-platform/4.6/security/certificates/service-serving-certificate.html

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-26 18:15:12 +05:30
Sébastien Han 20c302f72d ceph: use replicaset instead of deployments for rgw instances
Since Pacific, we can deploy multiple object gateway using the same
keying and they will appear in the service map separetly. So from now on
and on upgrades to Pacific Rook will remove all extra deployments to
only keep 1. In this single deployment the number of replica will be set
the desired `instances` count from the CRD spec.

You will now see gateways like this on Pacific:

```
rgw: 3 daemons active (1 hosts, 1 zones)
```

See an upgrade operator logs from Octopus to Pacific:

```
2021-04-14 10:05:16.883145 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-04-14 10:05:17.006492 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-04-14 10:05:17.006516 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-04-14 10:05:17.006523 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-04-14 10:05:17.006533 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-04-14 10:05:17.298176 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-04-14 10:05:17.307913 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" already exists. updating if needed
2021-04-14 10:05:17.317441 I | op-k8sutil: updating deployment "rook-ceph-rgw-my-store-a" after verifying it is safe to stop
2021-04-14 10:05:17.317463 I | op-mon: checking if we can stop the deployment rook-ceph-rgw-my-store-a
2021-04-14 10:05:25.552648 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mgr-a"
2021-04-14 10:05:25.552662 I | op-mon: checking if we can continue the deployment rook-ceph-mgr-a
2021-04-14 10:05:25.555077 I | op-mgr: setting services to point to mgr "a"
2021-04-14 10:05:25.608535 I | op-osd: start running osds in namespace "rook-ceph"
2021-04-14 10:05:25.608587 I | op-osd: wait timeout for healthy OSDs during upgrade or restart is "10m0s"
2021-04-14 10:05:25.618392 I | op-osd: start provisioning the OSDs on PVCs, if needed
2021-04-14 10:05:25.620512 I | op-osd: no storageClassDeviceSets or volumeSources are defined to configure OSDs on PVCs
2021-04-14 10:05:25.620542 I | op-osd: start provisioning the OSDs on nodes, if needed
2021-04-14 10:05:25.628416 I | op-osd: 1 of the 1 storage nodes are valid
2021-04-14 10:05:25.760702 I | op-k8sutil: Removing previous job rook-ceph-osd-prepare-minikube to start a new one
2021-04-14 10:05:25.773213 I | op-k8sutil: batch job rook-ceph-osd-prepare-minikube still exists
2021-04-14 10:05:26.595673 I | op-mgr: successful modules: prometheus
2021-04-14 10:05:27.129569 I | op-mgr: successful modules: mgr module(s) from the spec
2021-04-14 10:05:27.776720 I | op-k8sutil: batch job rook-ceph-osd-prepare-minikube deleted
2021-04-14 10:05:27.781881 I | op-osd: started OSD provisioning job for node "minikube"
2021-04-14 10:05:27.786693 I | op-osd: OSD orchestration status for node minikube is "starting"
2021-04-14 10:05:28.617937 I | op-mgr: successful modules: balancer
2021-04-14 10:05:28.706307 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:28.710245 I | op-osd: updating OSD 0 on node "minikube"
2021-04-14 10:05:31.596511 I | op-mgr: the dashboard secret was already generated
2021-04-14 10:05:31.926125 I | op-mgr: setting ceph dashboard "admin" login creds
2021-04-14 10:05:32.190343 E | op-mgr: failed modules: "dashboard". failed to initialize dashboard: failed to set login credentials for the ceph dashboard: failed to set login creds on mgr: failed to complete command for set dashboard creds: Invalid command: unused arguments: ["P.c(b6V$0pfK#)'70c5z"]
dashboard set-login-credentials <username> :  Set the login credentials. Password read from -i <file>
Traceback (most recent call last):
  File "/usr/bin/ceph", line 1310, in <module>
    retval = main()
  File "/usr/bin/ceph", line 1256, in main
    outf.write(outbuf)
TypeError: a bytes-like object is required, not 'str'
.
2021-04-14 10:05:35.665419 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-rgw-my-store-a"
2021-04-14 10:05:35.665553 I | op-mon: checking if we can continue the deployment rook-ceph-rgw-my-store-a
2021-04-14 10:05:35.675293 I | ceph-object-controller: config map "rook-ceph-rgw-my-store-mime-types" for object store "my-store" already exists, not overwriting
2021-04-14 10:05:35.691549 I | ceph-object-controller: found more rgw deployments 3 than desired 3 in object store "my-store", scaling down
2021-04-14 10:05:35.691580 I | op-k8sutil: removing deployment rook-ceph-rgw-my-store-c if it exists
2021-04-14 10:05:35.698382 I | op-k8sutil: Removed deployment rook-ceph-rgw-my-store-c
2021-04-14 10:05:35.704568 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:05:44.881129 I | op-osd: OSD orchestration status for node minikube is "orchestrating"
2021-04-14 10:05:44.882149 I | op-osd: OSD orchestration status for node minikube is "completed"
2021-04-14 10:05:45.756762 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:05:45.782039 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:45.785188 I | op-osd: updating OSD 1 on node "minikube"
2021-04-14 10:05:54.768851 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:54.773191 I | op-osd: updating OSD 2 on node "minikube"
2021-04-14 10:05:55.843879 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:06:05.255136 I | op-osd: finished running OSDs in namespace "rook-ceph"
2021-04-14 10:06:05.255155 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2021-04-14 10:06:05.560201 W | ceph-cluster-controller: upgrade orchestration completed but somehow we still have more than one Ceph version running. map[ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable):3 ceph version 16.2.0 (0c2054e95bcd9b30fdd908a79ac1d8bbc3394442) pacific (stable):6]:
2021-04-14 10:06:05.885001 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:06:08.321209 I | exec: timeout waiting for process radosgw-admin to return. Sending interrupt signal to the process
2021-04-14 10:06:10.645672 I | ceph-spec: object "rook-ceph-rgw-my-store-c" matched on delete, reconciling
2021-04-14 10:06:11.940994 I | op-k8sutil: confirmed rook-ceph-rgw-my-store-c does not exist
2021-04-14 10:06:11.972935 I | ceph-spec: object "rook-ceph-rgw-my-store-c-keyring" matched on delete, reconciling
2021-04-14 10:06:11.974394 I | ceph-object-controller: deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-c"
2021-04-14 10:06:13.721964 I | ceph-object-controller: successfully deleted rgw config for "client.rgw.my.store.c" in mon configuration database
2021-04-14 10:06:13.721989 I | cephclient: deleting ceph auth "client.rgw.my.store.c"
2021-04-14 10:06:14.073700 I | ceph-object-controller: completed deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-c"
2021-04-14 10:06:14.073724 I | op-k8sutil: removing deployment rook-ceph-rgw-my-store-b if it exists
2021-04-14 10:06:14.083950 I | op-k8sutil: Removed deployment rook-ceph-rgw-my-store-b
2021-04-14 10:06:14.090226 I | op-k8sutil: "rook-ceph-rgw-my-store-b" still found. waiting...
2021-04-14 10:06:24.174846 I | op-k8sutil: "rook-ceph-rgw-my-store-b" still found. waiting...
2021-04-14 10:06:30.505574 I | ceph-spec: object "rook-ceph-rgw-my-store-b" matched on delete, reconciling
2021-04-14 10:06:32.201263 I | op-k8sutil: confirmed rook-ceph-rgw-my-store-b does not exist
2021-04-14 10:06:32.219559 I | ceph-object-controller: deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-b"
2021-04-14 10:06:32.220219 I | ceph-spec: object "rook-ceph-rgw-my-store-b-keyring" matched on delete, reconciling
2021-04-14 10:06:33.821969 I | ceph-object-controller: successfully deleted rgw config for "client.rgw.my.store.b" in mon configuration database
2021-04-14 10:06:33.821992 I | cephclient: deleting ceph auth "client.rgw.my.store.b"
2021-04-14 10:06:34.182091 I | ceph-object-controller: completed deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-b"
2021-04-14 10:06:34.188659 I | ceph-object-controller: successfully scaled down rgw deployments to 1 in object store "my-store"
2021-04-14 10:06:34.188676 I | ceph-object-controller: enabling rgw dashboard
2021-04-14 10:06:49.487902 I | exec: timeout waiting for process radosgw-admin to return. Sending interrupt signal to the process
2021-04-14 10:06:49.494605 W | ceph-object-controller: failed to enable dashboard for rgw. failed to create user "dashboard-admin": failed to create s3 user: signal: interrupt
2021-04-14 10:06:49.494676 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
```

Running pods:

```
...
...
rook-ceph-rgw-my-store-a-6756bdbbd4-7265v            1/1     Running     1          41m
rook-ceph-rgw-my-store-a-6756bdbbd4-l5ccr            1/1     Running     1          41m
rook-ceph-rgw-my-store-a-6756bdbbd4-m679c            1/1     Running     1          41m
...
...
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-04-14 13:38:29 +02:00
Travis Nielsen 721acd1a8a Revert "ceph: added rook-ceph-default service account"
This reverts commit 737fb099fe.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-08 11:46:12 -06:00
parth-grandTareq Sharafy 737fb099fe ceph: added rook-ceph-default service account
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.

Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
2021-03-29 20:20:40 +05:30
subhamkraiandTravis Nielsen 319e4a41a4 ceph: placement in case of both PVC and non-PVC's
In the case of PVC,
We are giving lower priority to all placement.
We want deviceSet placement to applied and
override in case of overlapping settings and
we are merging nodeAffinity if applied in both
all placement and deviceSet.

In case of non-PVC,
we apply spec.placement

Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-18 10:32:57 +05:30
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Travis Nielsen e7331b0191 ceph: multiple mgrs have antiaffinity across hosts or zones
If there are multiple mgrs running, the operator will automatically
add pod antiaffinity across hosts by default. For stretch clusters,
antiaffinity across the stretch failure domain (e.g. zones)
will be added. Required antiaffinity will be created unless the
test setting allowMultiplePerNode is set to true.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00
Travis Nielsen 142ba63b6a ceph: enforce anti-affinity to multiple mgrs
If there are two mgr daemons, they should be running on different
hosts or zones, depending on the failure domain.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00