From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The Client_ID generated by operator was
different from the log rotate file created
The Clinet_ID= rgwceph.client.rook.ceph.rgw.my.store.a
and log file name= ceph-client.rgw.my.store.a.log
So changed the CLient_ID to ceph-client.rgw.my.store.a for
correct working and this follow the patterns how other modules
Client_ID is generated
Closes: https://github.com/rook/rook/issues/8692
Signed-off-by: parth-gr <paarora@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook will now auto detect the kv version of the vault server. This
allows users not having to pass the VAULT_BACKEND configuration in the
CephCluster CR.
Signed-off-by: Sébastien Han <seb@redhat.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
Recently, the builds of `ceph/ceph` image moved to quay.io, see
https://github.com/ceph/ceph-build/pull/1883 for more details.
Current images will remain but new builds will happen on quay.io only.
This means that tags such as `v14.2`, `v15.2`,`v16.2` will need to
switch to quay.io to get updates.
Signed-off-by: Sébastien Han <seb@redhat.com>
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Add `SecuritySpec` to specify kms details for rgw in `ObjectStoreSpec` than using existing
`SecuritySpec` used for `ClusterSpec`. The vault secret engines conflict with OSD and RGW
kms encrytpion
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.
Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The livenessprobe for swift/health always check with RGW internal port if host network disabled,
but won't work if only secure port is configured
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
After #6217, internal rgw spawning on external cluster was still broken,
operator claiming:
```
ceph-object-controller: failed to reconcile failed to create object
store deployments: failed to reconcile external endpoint:
failed to create or update object store "arch-cloud" endpoint:
failed to create endpoint "rook-ceph-rgw-arch-cloud". Endpoints
"rook-ceph-rgw-arch-cloud" is invalid: subsets[0]:
Required value: must specify `addresses` or `notReadyAddresses`
```
To spawn internal rgw for external cluster, changes detection of
what mean 'internal' for rgw and only use the externalRgwEndpoints
list for that independently of the status of the cluster
Now there is multiple posibilities:
* internal cluster and internal rgw pods: "normal case"
* external cluster and internal rgw pods: <= now working with the PR
* external cluster and external rgw pods: the external case
* internal cluster and external rgw pods: <= new case that could exist
Signed-off-by: Julien Girardin <jugirardin@free.fr>
Pools in stretch clusters will all use the same crush rule that
is generated by the operator when configuring the cluster for
stretch mode, with replica 4 and two replicas per failure domain.
Pools cannot create new crush rules, so we require that replica be
4 when the pools is created, EC is not allowed, and no new
crush rule will be created.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The pool spec has now a new property called "replicasPerFailureDomain"
which essentially represents the number of replicas to store in each
failure domain.
Assuming the failure domain is a datacenter (if the cluster is
stretched) then you will have 2 replicas per datacenter where each
replica ends up on a different host. This gives you a total of 4
replicas and for this, the "size" must be set to 4.
Closes: https://github.com/rook/rook/issues/5591
Signed-off-by: Sébastien Han <seb@redhat.com>
The labels cannot be updated on a deployment match selector.
In v1.4 a new label was added for ceph_daemon_type that is intended
to be on the pod labels, but cannot be applied to the match selectors.
Therefore, we suppress any new labels that are added to the daemons.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator should only connect to ceph with a single set of creds.
In a converged cluster this will be the admin creds and in an external
cluster it will be lower-privileged creds. Independent clusters were
implemented with a separate set of creds. To simplify the code these
are now merged to a single set.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Update the base operator image and cluster examples to use the latest
octopus release v15.2.4. Also cleanup some old examples and unit tests
that were still based on mimic.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.
A good status will look like:
status:
endpointStatus:
lastChanged: "2020-06-25T13:47:45Z"
lastChecked: "2020-06-25T13:48:46Z"
phase: Connected
A failed status:
status:
endpointStatus:
details: |-
error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
health: ERROR
This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.
Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).
Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
Code for pod anti affinity mechanism on ceph gateways is a duplica of the one used by ceph mon.
This is a refacto to use same code on both components
Signed-off-by: n.fraison <n.fraison@criteo.com>
As performed with mon, gateways from same ceph object store
must not run on the same host if hostNetwork is set to true
Signed-off-by: n.fraison <n.fraison@criteo.com>
ceph: Use the new NetworkSpec api
ceph: Change all tests to use the NetworkSpec
ceph: Remove TypeMeta field from NetworkSpec
ceph: Change IsHost network func signature
Signed-off-by: giovanism <giovanism@outlook.co.id>
When configuring an external cluster the orchestration of discover and all the ceph
daemons will be skipped. The operator will still watch for the creation of
crds for this namespace. When a filesystem, object, object user, or
ganesha crd are created, the operator will simply print an error
to the log.
Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit does multiple things:
* remove support for AllNodes where we would deploy one rgw per node on
all the nodes.
* a transition path is implemented in the code so that if someone has an
existing deployment, daemonsets will be removed and replaced by an
deployments.
* when using "instances", each rgw deployed has its own key which makes
Ceph reporting the exact number of rgw running, see:
```
[root@rook-ceph-operator-775cf575c5-bh4sr /]# ceph -s
cluster:
id: 611fcf39-0669-4864-9a12-debb35c0397a
health: HEALTH_OK
services:
mon: 3 daemons, quorum a,b,c (age 12h)
mgr: a(active, since 12h)
osd: 3 osds: 3 up (since 12h), 3 in (since 12h)
rgw: 3 daemons active (my.store.a, my.store.b, my.store.c)
data:
pools: 6 pools, 600 pgs
objects: 235 objects, 3.8 KiB
usage: 3.0 GiB used, 84 GiB / 87 GiB avail
pgs: 600 active+clean
```
Closes: https://github.com/rook/rook/issues/2474, https://github.com/rook/rook/issues/2957 and https://github.com/rook/rook/issues/3245
Signed-off-by: Sébastien Han <seb@redhat.com>
It's quite convinient to expose /var/log/ceph so that we can decide to
activate logs locally on the machine and see what's going on.
This is only a placeholder when a daemon is stuck crashlooping and we
want to allow administrator to gather log files.
We still do not log on file but this can be activated via a config
option passed to the centralized config option store.
For some daemons, which typically do not store any data (rgw, rbd-mirror
and mds) we had to propagate dataDirHostPath from the cluster spec to
each creation call so that the bindmount can happen.
Fixes: https://github.com/rook/rook/issues/2881
Signed-off-by: Sébastien Han <seb@redhat.com>
The basic idea here is to replace the release name in the ClusterInfo
object with the fine grained version information queried at runtime.
This patch also removes the Name field from the cluster spec, which was
also runtime determined and was redundant with ClusterInfo relase name.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
If someone sets limits to pod, we want to ensure the possible
experience, so we want to make sure that people do not configure
inapropriate values for certain daemons.
We decide to fail if the memory.limit is too low.
Signed-off-by: Sébastien Han <seb@redhat.com>
Configure the Ceph rgw daemon completely from the operator a la the
recent changes to the Ceph mon, mgr, and mds operators.
Create the rgw deployment or daemonset first, and then create the
keyring secret for the object store with its owner reference as the
corresponding deployment or daemonset. When the replication controller
is deleted, the secret is also deleted.
The RGW's mime.types file is now stored in a configmap with a different
file created for each object store. This is primarily just a means to
get the mime.types file into the rgw pod, but the added benefit is that
the administrator can modify the configmap, which could reduce
susceptibility to file type execution vulnerabilities (worst case).
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
The mgr orchestrator modules need the container image to run
the ceph image for blinking the lights when a disk is down
Signed-off-by: travisn <tnielsen@redhat.com>
This commit introduces the necessary changes to support the new
messenger feature coming with Ceph Nautilus (currently in development).
What changes? Now the monitor listens on two port:
* old 6789 for messenger v1, which will help us support older client
(e,g: krbd)
* new 3300 for messengers v2, which brings new improvement in the
messaging layer. This new transport layer brings numerous advantages
such as encryption improvement, speed improvement, pluggable nature to
support different network stack than TCP and many more.
We still have one Service IP, however it has 2 ports, see:
```
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
rook-ceph-mon-a ClusterIP 10.106.217.160 <none> 3300/TCP,6789/TCP 4h
rook-ceph-mon-b ClusterIP 10.99.36.175 <none> 3300/TCP,6789/TCP 4h
rook-ceph-mon-c ClusterIP 10.108.220.74 <none> 3300/TCP,6789/TCP 4h
````
The `ceph.conf` has changed and we don't force the port when using an IP
address (public addr etc). Ceph, depending on its version will naturally
start the monitors on their right port, 6789.
A new --ceph-version-name CLI argument has been added to the Rook binary
so that when the pod starts it passes the ceph version name and the
configuration of the ceph.conf, as well as daemon startup flags, happen
properly.
Given that the Rook Operator remembers the port of all the monitors it
deployed (through Pod definition), this change is not an issue and will
maintain backward compatibility.
Note that to test this you must build rook with dev container image,
which contains the dev Nautilus version. So you should do something
like:
`make -j4 BASEIMAGE='ceph/daemon-base:latest-master' IMAGES='ceph' build`
Resolves: #2525
Signed-off-by: Sébastien Han <seb@redhat.com>