Commit Graph
94 Commits
Author SHA1 Message Date
Jiffin Tony Thottan aba50d3ca9 object: add support in RGW to communicate vault with TLS
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-17 10:19:28 +05:30
Yuzuki Mimura 536b59ef0f rgw: change the way to livenessProbe and introduce readinessProbe
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.

Closes: #8407

Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-10-29 15:07:29 +00:00
parth-gr fc7905a7bd ceph: fixing ClientID of log-collector for RGW instance
The Client_ID generated by operator was
different from the log rotate file created
The Clinet_ID= rgwceph.client.rook.ceph.rgw.my.store.a
and log file name= ceph-client.rgw.my.store.a.log
So changed the CLient_ID to ceph-client.rgw.my.store.a for
correct working and this follow the patterns how other modules
Client_ID is generated

Closes: https://github.com/rook/rook/issues/8692
Signed-off-by: parth-gr <paarora@redhat.com>
2021-10-07 19:30:21 +05:30
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han d675969567 ceph: fix vault kv secret engine auto-detection
Passing struct by value essentially gives you a copy, so when modified
within a function, the scope is then reduced to that function. Using
pointers solves that you mutate the struct as many times as you want from
anywhere.
As a result, the auto-detection of the Vault KV backend was not working
correctly.
Also, added a ton of unit tests for Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-31 15:29:04 +02:00
Denis Egorenko fb04908315 ceph: add ability to specify ca bundle for rgw
Specify ca Bundle for RGW spec and mount inside pods.

Related-Issue: https://github.com/rook/rook/issues/8490
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
2021-08-10 23:18:26 +04:00
Sébastien Han 99e00dea1e ceph: auto detect vault k/v version
Rook will now auto detect the kv version of the vault server. This
allows users not having to pass the VAULT_BACKEND configuration in the
CephCluster CR.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-29 10:16:31 +02:00
Jiffin Tony Thottan 1665ad6ea7 ceph: add support for tls certs via k8s tls secrets for rgw
With this PR the RGW can accept TLS certs as K8s TLS secrets

Fixes: 2079
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-07-06 16:30:16 +05:30
Jiffin Tony Thottan d909c309f3 ceph: update the backend path for transit engine
The backend has for trasnit secret in vault kms for RGW is changed from pacific

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-05-27 22:48:21 +05:30
Travis Nielsen b0a63711f5 build: refactor to consolidate the rook.io/v1 package
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-18 19:55:37 -06:00
Jiffin Tony Thottan a5a5661e5f ceph: add security spec to object store spec
Add `SecuritySpec` to specify kms details for rgw in `ObjectStoreSpec` than using existing
`SecuritySpec` used for `ClusterSpec`. The vault secret engines conflict with OSD and RGW
kms encrytpion

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-28 14:39:39 +05:30
Jiffin Tony Thottan 3ef3e8b72e ceph: check for kv engine version for SSE RGW
The "https://docs.ceph.com/en/latest/radosgw/vault/" says RGW supports only version v2 of
kv engine, so pod won't start if the admin didn't provide "VAULT_BACKEND" in security.kms.ConnectionDetails

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-28 13:03:22 +05:30
Jiffin Tony Thottan aea21d9c84 ceph: service server cert support for rgw
Service serving certificates are intended to applications that require
encryption in openshift. These certificates are issued as TLS web server
certificates. Currently RGW supports TLS authentication with help of
certs passed as secrets, in this case we add following details as
`service.annotations` in the Objectstore Gateway Spec :
```
service:
  annotations:
    service.beta.openshift.io/serving-cert-secret-name: <name for
autogenerated secret>
```

More details about service serving cert can be found at :
https://docs.openshift.com/container-platform/4.6/security/certificates/service-serving-certificate.html

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-26 18:15:12 +05:30
Sébastien Han 20c302f72d ceph: use replicaset instead of deployments for rgw instances
Since Pacific, we can deploy multiple object gateway using the same
keying and they will appear in the service map separetly. So from now on
and on upgrades to Pacific Rook will remove all extra deployments to
only keep 1. In this single deployment the number of replica will be set
the desired `instances` count from the CRD spec.

You will now see gateways like this on Pacific:

```
rgw: 3 daemons active (1 hosts, 1 zones)
```

See an upgrade operator logs from Octopus to Pacific:

```
2021-04-14 10:05:16.883145 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-04-14 10:05:17.006492 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-04-14 10:05:17.006516 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-04-14 10:05:17.006523 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-04-14 10:05:17.006533 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-04-14 10:05:17.298176 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-04-14 10:05:17.307913 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" already exists. updating if needed
2021-04-14 10:05:17.317441 I | op-k8sutil: updating deployment "rook-ceph-rgw-my-store-a" after verifying it is safe to stop
2021-04-14 10:05:17.317463 I | op-mon: checking if we can stop the deployment rook-ceph-rgw-my-store-a
2021-04-14 10:05:25.552648 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mgr-a"
2021-04-14 10:05:25.552662 I | op-mon: checking if we can continue the deployment rook-ceph-mgr-a
2021-04-14 10:05:25.555077 I | op-mgr: setting services to point to mgr "a"
2021-04-14 10:05:25.608535 I | op-osd: start running osds in namespace "rook-ceph"
2021-04-14 10:05:25.608587 I | op-osd: wait timeout for healthy OSDs during upgrade or restart is "10m0s"
2021-04-14 10:05:25.618392 I | op-osd: start provisioning the OSDs on PVCs, if needed
2021-04-14 10:05:25.620512 I | op-osd: no storageClassDeviceSets or volumeSources are defined to configure OSDs on PVCs
2021-04-14 10:05:25.620542 I | op-osd: start provisioning the OSDs on nodes, if needed
2021-04-14 10:05:25.628416 I | op-osd: 1 of the 1 storage nodes are valid
2021-04-14 10:05:25.760702 I | op-k8sutil: Removing previous job rook-ceph-osd-prepare-minikube to start a new one
2021-04-14 10:05:25.773213 I | op-k8sutil: batch job rook-ceph-osd-prepare-minikube still exists
2021-04-14 10:05:26.595673 I | op-mgr: successful modules: prometheus
2021-04-14 10:05:27.129569 I | op-mgr: successful modules: mgr module(s) from the spec
2021-04-14 10:05:27.776720 I | op-k8sutil: batch job rook-ceph-osd-prepare-minikube deleted
2021-04-14 10:05:27.781881 I | op-osd: started OSD provisioning job for node "minikube"
2021-04-14 10:05:27.786693 I | op-osd: OSD orchestration status for node minikube is "starting"
2021-04-14 10:05:28.617937 I | op-mgr: successful modules: balancer
2021-04-14 10:05:28.706307 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:28.710245 I | op-osd: updating OSD 0 on node "minikube"
2021-04-14 10:05:31.596511 I | op-mgr: the dashboard secret was already generated
2021-04-14 10:05:31.926125 I | op-mgr: setting ceph dashboard "admin" login creds
2021-04-14 10:05:32.190343 E | op-mgr: failed modules: "dashboard". failed to initialize dashboard: failed to set login credentials for the ceph dashboard: failed to set login creds on mgr: failed to complete command for set dashboard creds: Invalid command: unused arguments: ["P.c(b6V$0pfK#)'70c5z"]
dashboard set-login-credentials <username> :  Set the login credentials. Password read from -i <file>
Traceback (most recent call last):
  File "/usr/bin/ceph", line 1310, in <module>
    retval = main()
  File "/usr/bin/ceph", line 1256, in main
    outf.write(outbuf)
TypeError: a bytes-like object is required, not 'str'
.
2021-04-14 10:05:35.665419 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-rgw-my-store-a"
2021-04-14 10:05:35.665553 I | op-mon: checking if we can continue the deployment rook-ceph-rgw-my-store-a
2021-04-14 10:05:35.675293 I | ceph-object-controller: config map "rook-ceph-rgw-my-store-mime-types" for object store "my-store" already exists, not overwriting
2021-04-14 10:05:35.691549 I | ceph-object-controller: found more rgw deployments 3 than desired 3 in object store "my-store", scaling down
2021-04-14 10:05:35.691580 I | op-k8sutil: removing deployment rook-ceph-rgw-my-store-c if it exists
2021-04-14 10:05:35.698382 I | op-k8sutil: Removed deployment rook-ceph-rgw-my-store-c
2021-04-14 10:05:35.704568 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:05:44.881129 I | op-osd: OSD orchestration status for node minikube is "orchestrating"
2021-04-14 10:05:44.882149 I | op-osd: OSD orchestration status for node minikube is "completed"
2021-04-14 10:05:45.756762 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:05:45.782039 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:45.785188 I | op-osd: updating OSD 1 on node "minikube"
2021-04-14 10:05:54.768851 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:54.773191 I | op-osd: updating OSD 2 on node "minikube"
2021-04-14 10:05:55.843879 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:06:05.255136 I | op-osd: finished running OSDs in namespace "rook-ceph"
2021-04-14 10:06:05.255155 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2021-04-14 10:06:05.560201 W | ceph-cluster-controller: upgrade orchestration completed but somehow we still have more than one Ceph version running. map[ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable):3 ceph version 16.2.0 (0c2054e95bcd9b30fdd908a79ac1d8bbc3394442) pacific (stable):6]:
2021-04-14 10:06:05.885001 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:06:08.321209 I | exec: timeout waiting for process radosgw-admin to return. Sending interrupt signal to the process
2021-04-14 10:06:10.645672 I | ceph-spec: object "rook-ceph-rgw-my-store-c" matched on delete, reconciling
2021-04-14 10:06:11.940994 I | op-k8sutil: confirmed rook-ceph-rgw-my-store-c does not exist
2021-04-14 10:06:11.972935 I | ceph-spec: object "rook-ceph-rgw-my-store-c-keyring" matched on delete, reconciling
2021-04-14 10:06:11.974394 I | ceph-object-controller: deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-c"
2021-04-14 10:06:13.721964 I | ceph-object-controller: successfully deleted rgw config for "client.rgw.my.store.c" in mon configuration database
2021-04-14 10:06:13.721989 I | cephclient: deleting ceph auth "client.rgw.my.store.c"
2021-04-14 10:06:14.073700 I | ceph-object-controller: completed deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-c"
2021-04-14 10:06:14.073724 I | op-k8sutil: removing deployment rook-ceph-rgw-my-store-b if it exists
2021-04-14 10:06:14.083950 I | op-k8sutil: Removed deployment rook-ceph-rgw-my-store-b
2021-04-14 10:06:14.090226 I | op-k8sutil: "rook-ceph-rgw-my-store-b" still found. waiting...
2021-04-14 10:06:24.174846 I | op-k8sutil: "rook-ceph-rgw-my-store-b" still found. waiting...
2021-04-14 10:06:30.505574 I | ceph-spec: object "rook-ceph-rgw-my-store-b" matched on delete, reconciling
2021-04-14 10:06:32.201263 I | op-k8sutil: confirmed rook-ceph-rgw-my-store-b does not exist
2021-04-14 10:06:32.219559 I | ceph-object-controller: deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-b"
2021-04-14 10:06:32.220219 I | ceph-spec: object "rook-ceph-rgw-my-store-b-keyring" matched on delete, reconciling
2021-04-14 10:06:33.821969 I | ceph-object-controller: successfully deleted rgw config for "client.rgw.my.store.b" in mon configuration database
2021-04-14 10:06:33.821992 I | cephclient: deleting ceph auth "client.rgw.my.store.b"
2021-04-14 10:06:34.182091 I | ceph-object-controller: completed deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-b"
2021-04-14 10:06:34.188659 I | ceph-object-controller: successfully scaled down rgw deployments to 1 in object store "my-store"
2021-04-14 10:06:34.188676 I | ceph-object-controller: enabling rgw dashboard
2021-04-14 10:06:49.487902 I | exec: timeout waiting for process radosgw-admin to return. Sending interrupt signal to the process
2021-04-14 10:06:49.494605 W | ceph-object-controller: failed to enable dashboard for rgw. failed to create user "dashboard-admin": failed to create s3 user: signal: interrupt
2021-04-14 10:06:49.494676 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
```

Running pods:

```
...
...
rook-ceph-rgw-my-store-a-6756bdbbd4-7265v            1/1     Running     1          41m
rook-ceph-rgw-my-store-a-6756bdbbd4-l5ccr            1/1     Running     1          41m
rook-ceph-rgw-my-store-a-6756bdbbd4-m679c            1/1     Running     1          41m
...
...
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-04-14 13:38:29 +02:00
Travis Nielsen 721acd1a8a Revert "ceph: added rook-ceph-default service account"
This reverts commit 737fb099fe.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-08 11:46:12 -06:00
parth-grandTareq Sharafy 737fb099fe ceph: added rook-ceph-default service account
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.

Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
2021-03-29 20:20:40 +05:30
subhamkraiandTravis Nielsen 319e4a41a4 ceph: placement in case of both PVC and non-PVC's
In the case of PVC,
We are giving lower priority to all placement.
We want deviceSet placement to applied and
override in case of overlapping settings and
we are merging nodeAffinity if applied in both
all placement and deviceSet.

In case of non-PVC,
we apply spec.placement

Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-18 10:32:57 +05:30
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Travis Nielsen e7331b0191 ceph: multiple mgrs have antiaffinity across hosts or zones
If there are multiple mgrs running, the operator will automatically
add pod antiaffinity across hosts by default. For stretch clusters,
antiaffinity across the stretch failure domain (e.g. zones)
will be added. Required antiaffinity will be created unless the
test setting allowMultiplePerNode is set to true.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00
Travis Nielsen 142ba63b6a ceph: enforce anti-affinity to multiple mgrs
If there are two mgr daemons, they should be running on different
hosts or zones, depending on the failure domain.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00
Travis Nielsen ddedecabac Merge pull request #7334 from thotz/livenessprobergwsslenabled
ceph: handle ssl cases for RGW's livenessprobe
2021-03-02 11:42:06 -07:00
Jiffin Tony Thottan 785aa5a322 ceph: handle ssl cases for RGW's livenessprobe
The livenessprobe for swift/health always check with RGW internal port if host network disabled,
but won't work if only secure port is configured

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-03-02 14:39:57 +05:30
Sébastien Han e4c3a4cef7 ceph: revert kms VAULT_KV_VERS to VAULT_BACKEND
The vault env variable VAULT_BACKEND is read by the secret library so we
must not change it. This is reverting
ee6a03dfb8 with a few extra fixes for the
rgw code.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-02 10:08:39 +01:00
subhamkrai 4ef4be019b ceph: override default values in liveness probes
this commit set default values or allow overriding of liveness probe
fields if that individual field is 0 or nil.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2021-02-12 20:35:56 +05:30
Jiffin Tony Thottan da61c9a83e ceph: vault kms configuration for ceph object store
The first patch to configure vault for ceph object store. If the `security.kms` configured in
`clusterSpec` CRD, RGW will be configured with vault kms settings to handle SSE request from s3 clients.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-01-22 10:44:03 +05:30
Sébastien Han cb7d600841 ceph: set process working dir to /var/log/ceph
By setting the working directory of a Ceph daemon to the log direcor
(which is bindmounted to the host), we can ensure that the coredumps
will be available.

On CentOS 7 **only**, the kernel is configured with:

```
cat /proc/sys/kernel/core_pattern
core
```

which means that the coredumps will end up in the process working
directory and will be named "core".

On CentOS 8, everything is different and this won't work but still
CentOS will be able to consume this and changing the working directory
is harmless.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-15 16:06:42 +01:00
Julien Girardin d63d9b7d4d ceph: change external rgw detection, not relying on cluster
After #6217, internal rgw spawning on external cluster was still broken,
operator claiming:
```
ceph-object-controller: failed to reconcile failed to create object
store deployments: failed to reconcile external endpoint:
failed to create or update object store "arch-cloud" endpoint:
failed to create endpoint "rook-ceph-rgw-arch-cloud". Endpoints
"rook-ceph-rgw-arch-cloud" is invalid: subsets[0]:
Required value: must specify `addresses` or `notReadyAddresses`
```

To spawn internal rgw for external cluster, changes detection of
what mean 'internal' for rgw and only use the externalRgwEndpoints
list for that independently of the status of the cluster
Now there is multiple posibilities:

 * internal cluster and internal rgw pods: "normal case"
 * external cluster and internal rgw pods: <= now working with the PR
 * external cluster and external rgw pods: the external case
 * internal cluster and external rgw pods: <= new case that could exist

Signed-off-by: Julien Girardin <jugirardin@free.fr>
2021-01-08 11:28:41 +01:00
Travis Nielsen 8913a14ce4 ceph: rgw service selector should not change
If the selector changes on the rgw service, during upgrade the
clients will briefly not be able to connect to the rgw pods.
After all the pods are updated with any new labels, the connections
would be restored, but we want to avoid that temporary outage.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-02 16:45:02 -07:00
Sébastien Han c6a87203ca ceph: add log collector
We can now collect logs directly into a side-car container.
A new CRD spec has been added:

spec:
  logCollector:
    enabled: true
    periodicity: 24h

Every 24h we will rotate log files for each Ceph daemon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-12-01 16:30:17 +01:00
Travis Nielsen b91f4211c9 ceph: configure a stretched cluster
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:17:52 -07:00
Santosh Pillai 70c5f70701 ceph: support IPv6 single-stack for ceph
Add --ms-bind-ipv6 true args to mom, osd, mgr, object, mds, rbd daemons.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-10-02 01:03:03 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
Sébastien Han af72502fe4 ceph: do not use label selector on external svc
When the cluster is external there are no running pods so setting a
Selector on the Service Spec does not make sense, even worse, it makes
Kubernetes reconciling and overriding the configured endpoint.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Alexander Trost a87c00ddee ceph: allow user to add labels to pods
This allows users to specify additional labels to be added to the components
created by the Rook Ceph operator.
Additionally the rbd mirrors are now getting their annotations and
labels set.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2020-08-25 19:38:51 +02:00
subhamkrai e21854145a ceph: disable object store port when secureport specified
currently, the object store schema requires the
port to be non-zero. This does not allow http
access to the object store to be disabled.
so setting the port minimum to 0 to from 1 in the cr.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-08-12 21:57:01 +05:30
Travis Nielsen 06c21954e1 ceph: labels are immutable for match selectors
The labels cannot be updated on a deployment match selector.
In v1.4 a new label was added for ceph_daemon_type that is intended
to be on the pod labels, but cannot be applied to the match selectors.
Therefore, we suppress any new labels that are added to the daemons.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-04 16:47:00 -06:00
Blaine Gardner dcf8ae2526 ceph: rename pod labels function for more clarity
Rename PodLabels function to CephDaemonAppLabels for more clarity about
what the function's purpose is.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:30 -06:00
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
subhamkrai 23f0ac7f9c ceph: ability to configure livenessprobe for cephobjectstore crd
this commit allows us to disable livenessprobe for
cephobjectstore crd.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-20 22:53:32 +05:30
Sébastien Han 34ccb84aa3 ceph: add external prometheus endpoint
We can now connect an external prometheus exporter to Rook to collect
metrics and generate alerts from Prometheus.
Just enable this in the CephCluster CR:

```
spec:
  external:
    enable: true
   monitoring:
    enabled: true
    rulesNamespace: rook-ceph
    externalMgrEndpoints:
    - ip: 192.168.39.182
```

Closes: https://github.com/rook/rook/issues/5516
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-17 11:29:48 +02:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Sébastien Han e4eaa91ede ceph: add rgw endpoint healthcheck
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.

A good status will look like:

status:
  endpointStatus:
    lastChanged: "2020-06-25T13:47:45Z"
    lastChecked: "2020-06-25T13:48:46Z"
  phase: Connected

A failed status:

status:
  endpointStatus:
    details: |-
      error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
      caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
    health: ERROR

This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.

Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-02 16:34:10 +02:00
Ali Maredia fc579f4520 ceph: initial commit for ceph rgw multisite resources
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone

Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-06-09 16:28:53 -04:00
n.fraison c8df6c0faf ceph: add preferred pod anti-affinity on rgw if not host network
If host networking is not enabled add preferred pod anti-affinity on rgw pods to
improve spread of them

Signed-off-by: n.fraison <n.fraison@criteo.com>
2020-06-04 13:34:30 +02:00
Sébastien Han bc581da75e ceph: change rgw certificate mode
Because the secret is mounted and owned by root and the rgw is running
under the 'ceph' user then it cannot access it. So moving from 0400 to
0444 fixes the issue.

Closes: https://github.com/rook/rook/issues/5265
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-05-13 18:07:53 +02:00
Sébastien Han 812f1bd4bb ceph: revert "objectstore: Add ceph group as FsGroup"
This reverts commit 68fc50686a since it
now conflict with OCP environments with the other scc in place.

Closes: https://github.com/rook/rook/issues/5468
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-05-13 15:58:10 +02:00
Sébastien Han 2e5d6c749e Merge pull request #5362 from thotz/ssl_regression_rgw
objectstore: Add ceph group to FsGroup
2020-05-07 17:18:49 +02:00
Jiffin Tony Thottan 68fc50686a objectstore: Add ceph group as FsGroup
ceph gid (167) provided as supplymentary group, so that all the mounts can
be accessed by the same. This is required for ssl certificate consumed by
the rgw server from ceph octopus onwards.

Closes: https://github.com/rook/rook/issues/5265
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-05-07 18:37:36 +05:30
Madhu Rajanna 81688398f2 cleanup: use err.Wrap when the formatting is not required
Replaced err.Wrapf with err.Wrap when the formatting
is not required.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-29 17:42:59 +05:30
Satoru Takeuchi f78efc50aa cleanup: use the constants of environment variables properly
There are some code where constants of environment variables
are defined but not used properly. Let's avoid to use string
literals as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-04-24 17:31:00 +00:00