Commit Graph
549 Commits
Author SHA1 Message Date
Elise Gafford 2ca0d8f93d ceph: add cleanupPolicy to cephcluster spec
In order to perform full automation of cephcluster deletion safely
and protect against catastrophic data loss, the end user must be
able to signify that they intend to irrecoverably delete the data
in their cluster. The cleanupPolicy field of the cluster spec is
intended to communicate this. This patch does not implement
automated deletion, but only creates the spec field and the
safety feature of halting orchestration other than deletion on a
cluster with a set cleanup policy value.

Partially-fixes: #3222
Signed-off-by: Elise Gafford <egafford@redhat.com>
2020-03-30 19:11:54 +05:30
Adler Fleurant 70e7f39db0 ceph: Removed 2nd env variable ROOK_HOSTPATH_REQUIRES_PRIVILEGED
It is not necessary to have 2 same environment variables and not desired as
changing only one may lead to the second variable being applied with unexpected
value. It is better to have only 1 ROOK_HOSTPATH_REQUIRES_PRIVILEGED environment
variable.

Closes: https://github.com/rook/rook/issues/5107
Signed-off-by: Adler Fleurant <Adler.Fleurant@PicoChange.com>
2020-03-30 07:50:30 -04:00
Umanga Chapagain 0e932c15eb Ceph: add CSI configurations to ConfigMap
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-27 15:17:30 +05:30
Travis Nielsen da3bf49b61 Merge pull request #5043 from cybozu-go/remove-unnecessary-napespace-fields
remove unnecessary namespace fields
2020-03-20 17:18:21 -06:00
Satoru Takeuchi a5ea2858c9 manifests: remove unnecessary namespace fields
There are many namespace fields in ClusterRole{,Binding}. However,
ClusterRole{,Binding} are not namespaced. So we can remove these.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-03-18 15:58:31 +00:00
Travis Nielsen 5702d29a3c Merge pull request #5033 from leseb/clarify-block-pool
ceph: update pool.yaml with more options
2020-03-18 08:28:24 -06:00
Sébastien Han 9f7745e8e7 ceph: update pool.yaml with more options
Expose:

* crushRoot
* deviceClass

To the example pool.yaml for awareness.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-18 15:22:27 +01:00
Sébastien Han 08dad79832 Merge pull request #5038 from travisn/design-docs-cleanup
Design docs cleanup
2020-03-18 10:51:36 +01:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Travis Nielsen f79bc3d6a0 noobaa: remove design docs due to inactivity
The NooBaa operator was proposed for addition to Rook, but
the team went a different direction and created an independent
operator. Removing the stale docs.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-17 14:30:56 -06:00
Travis Nielsen 2c70e10bd6 docs: add cluster-on-pvc and remove cluster-minimal examples
The cluster-minimal.yaml example is more confusing than helpful. Most users
seem to think that minimal includes a fully working cluster with OSDs.

The cluster-on-pvc.yaml example is frequently used and was missing from the
examples doc.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-16 15:13:56 -06:00
Travis Nielsen 9c5c78cf01 Merge pull request #5025 from jmolmo/blinking_lights
Ceph: Added rook manager module rights for read pod logs
2020-03-16 13:57:13 -06:00
Travis Nielsen 017a560db3 Merge pull request #5021 from samkulkarni20/yb-mem-limits
YugabyteDB: Add resource limits to YugabyteDB pods
2020-03-16 11:12:07 -06:00
Juan Miguel Olmo Martínez c87a01bb8c Ceph: Added rook manager module rights for read pod logs
This is needed to allow the ceph manager rook module to read the output
of the <lsmcli> command used to switch disk identification light on/off on
physical disk devices

[test ceph]

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2020-03-16 18:03:23 +01:00
Sameer Kulkarni 909770a346 YugabyteDB: Add resource limits to YugabyteDB pods
Current YugabyteDB operator code creates Master and TServer pods without any resource requests/limits (specifically CPU and memory).
This causes the operator to run into soft/hard memory limit issue. The fix adds recommended resource requests and limits as defaults to the pods it creates.

Closes: https://github.com/yugabyte/yugabyte-db/issues/3884
Signed-off-by: Sameer Kulkarni <samkulkarni20@gmail.com>
2020-03-14 12:23:14 +05:30
Travis Nielsen 7de388ce92 ceph: grant access to the mgr to delete pods
The Ceph mgr modules need access to restart the daemon pods
when the configuration changes for the daemons.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-13 17:08:27 -06:00
Sébastien Han 14cb46ae39 ceph: bump to 14.2.8
Ceph 14.2.8 just got released so let's use it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-11 19:32:48 +01:00
Sébastien Han 9a8d539d57 ceph: use crd subressource
The .metadata.generation field is updated if and only if the value at the .spec subpath changes.
Additionally, if the spec does not change, .metadata.generation is not updated.
This is a must-have for the controller-runtime work, if we don't have
this, the generation field will always be incremented, resulting in a
endless reconcile loop.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-05 12:37:44 -07:00
Stefan Haas b3cc4aeb44 ceph: ceph-csi version detection #3824
Checks the version of the configured ceph-csi image while starting the operator. The operator will fail if the image is not supported.
Added an additional parameter to operator to disable the check e.g. to test not yet supported csi images.

Signed-off-by: Stefan Haas <shaas@suse.com>
2020-02-28 14:14:40 +01:00
Sébastien Han dab9233c8b ceph: add CRD setting for pool size 1
As of Octopus, Ceph will prevent you from creating a pool with a
replica size of 1. Allowing such pool could lead to data loss, so enable
the new option: requireSafeReplicaSize: false if you are **ABSOLUTELY**
certain that is what you want.

Closes: https://github.com/rook/rook/issues/4889
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-25 16:28:58 +01:00
Sébastien Han 5ca4880ed0 Merge pull request #4888 from leseb/pvc-hybrid-dev-class
ceph: set crush device class via annotations on PVC
2020-02-25 14:07:22 +01:00
Travis Nielsen e35939b26c ceph: bump ceph version to v14.2.7
The ceph/ceph:v14.2.7 image is released so we can pick these fixes
up as the recommended version of ceph.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-20 11:06:40 -07:00
Sébastien Han a702f463cc ceph: set crush device class via annotations on PVC
When the OSD on PVC is backed by a metadata block PVC, the Ceph CRUSH
device class should be set to something else rather than the rotational
property of the drive.

Closes: https://github.com/rook/rook/issues/4881
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-20 18:53:31 +01:00
Travis Nielsen a9903ed8d7 ceph: fix status for EC pool creation
Creating an EC pool was succeeding, but then the update to the
status was failing because of an incorrect check for changing EC
parameters. Now we correctly check if EC parameters are changing
unexpectedly.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-18 15:34:01 -07:00
Sébastien HanandSean Micklethwaite b04e35dab5 ceph: add support for disk selection by full path
Rook-Ceph now supports specify device by their fullpath instead of links
created by the kernel like /dev/sdb.

Closes: https://github.com/rook/rook/issues/1228
Co-authored-by: Sean Micklethwaite <sean@wayve.ai>
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-13 17:31:38 +01:00
Umanga Chapagain 7b201b8dae Ceph: adds "rook-ceph-operator-config" configmap
"rook-ceph-operator-config" can be used to override Env Var
initialized during operator deployment.

This commit makes use of the ConfigMap to determine NodeAffinity
and Tolerations for ceph-csi drivers.

Closes #3239
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-02-12 14:07:03 +05:30
Sébastien Han c08c3ced05 ceph: add support for metadata PVC for OSD on PVC
We now support the addition of the PVC that acts as a metadata device
for a given OSD.
For this, you need to create a new `volumeClaimTemplates`, its name must
be "metadata" otherwise, Rook won't pick it up.

A template will look like this:

```
volumeClaimTemplates:
- metadata:
    name: data
  spec:
    resources:
      requests:
        storage: 10Gi
    # IMPORTANT: Change the storage class depending on your environment (e.g. local-storage, gp2)
    storageClassName: gp2
    volumeMode: Block
    accessModes:
      - ReadWriteOnce
- metadata:
    name: metadata
  spec:
    resources:
      requests:
        storage: 6Gi
    # IMPORTANT: Change the storage class depending on your environment (e.g. local-storage, gp2)
    storageClassName: gp2
    volumeMode: Block
    accessModes:
      - ReadWriteOnce
```

We now map block and block.db directly inside the container instead of
running ceph-volume activate. This is much cleaner.

Closes: https://github.com/rook/rook/issues/3852
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-11 17:30:02 +01:00
Travis Nielsen dbfcef8c19 ceph: set min size to 0 on pool replication
The replication and erasure code settings are mutually exclusive
so must both be treated as optional by the schema validation.
The minimum must be set to 0 to treat them as optional.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-06 13:48:52 -07:00
Sébastien Han cface95317 ceph: ability to set target_size_ratio
target_size_ratio is used the tell the Ceph cluster the excepted
occupancy of a pool in the cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-06 09:13:47 +01:00
Maksim Nabokikh 7dc9ec1191 cassandra: JMX prometheus exporter
Export prometheus metrics for cassandra in sidecar container

Signed-off-by: Maksim Nabokikh <maksim.nabokikh@flant.com>
2020-02-03 11:40:38 +04:00
Travis Nielsen 57fb7cda2e Merge pull request #4714 from Madhu-1/log-level
CSI: Add LogLevel for csi templates
2020-01-31 11:54:32 -07:00
Sébastien Han 720c06cfec Merge pull request #4791 from galexrt/pickup_3689
doc: fix CephFS CSI StorageClass name
2020-01-31 14:18:13 +01:00
Alexander Trost a4f9981d83 doc: fix CephFS CSI StorageClass name
This fixes the StorageClass name example and doc to be `rook-cephfs`
instead of just `csi-cephfs`.

Closes #4130.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2020-01-31 11:49:35 +01:00
Madhu Rajanna 48a6e493e8 CSI: Add LogLevel for csi templates
Currently there is no option to turn on and off the
log level verbosity in csi containers, This PR adds
the functionality to select logging levels for csi containers.

Fixes: #4690

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-01-31 15:04:11 +05:30
Nizamudeen 4b55a74ac4 ceph: Implementing Conditions on rook-ceph
Fixed conditions getting resetted after the operator restart.
Did the changes which required to implement conditions on the rook ceph cluster
Conditions will eliminate the current status.State and incorporates a type which
provides much more description to the current status of the cluster.

Signed-off-by: Nizamudeen <nia@redhat.com>
2020-01-30 10:21:11 +05:30
Travis Nielsen ec7f06b52b Merge pull request #4680 from jmolmo/mdszones
Ceph: MDS pod placement adheres to fault domain topology
2020-01-28 16:19:53 -07:00
Madhu Rajanna 5a685fba77 CSI: Add support for rbd erasure coded pool
In ceph-csi v2.0.0 added a support for specifing
the erasure coded pool in storageclass which will be
used to store data.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-01-28 23:50:23 +05:30
Madhu Rajanna cf16184899 CSI: Add support for CSI volume expansion
This PR makes the necessary changes required
for volume expansion support for both cephfs and
rbd.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-01-28 23:50:23 +05:30
Madhu Rajanna 8c1a90ca01 CSI: upgrade node-driver-registrar from v1.1.0 to v1.2.0
latest released upstream version is v1.2.0,this
PR updates it to latest releasd version.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-01-28 23:50:15 +05:30
Madhu Rajanna 06d2dd2994 CSI: update csi-attacher from v1.2.0 to v2.1.0
external-attacher v2.x is not compatible with v1.x
This PR makes the require changes in csi templates
and upgrade documentation for easily upgrade

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-01-28 23:46:24 +05:30
Madhu Rajanna 29f8749ded CSI: Remove containerized flag from rbd daemonset template
In ceph-csi release v2.0.0, deprecated `containerized` flag
is removed.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-01-28 12:20:13 +05:30
Madhu Rajanna a5605e8063 CSI: add missing image tag ENV to openshift operator
CSI image tag ENV required to select the image tag
is missing in operator-openshift.yaml, This commit
adds the missing image tag ENV to operator-openshift.yaml

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-01-28 12:20:13 +05:30
Madhu Rajanna 0abb3b9e1b CSI: update ceph-csi image tag to v2.0.0
ceph-csi v2.0.0 is released. This commit updates
the image tag to latest released ceph-csi version

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-01-28 12:20:13 +05:30
Travis Nielsen d40cb217ae Merge pull request #4694 from egafford/secure-port-for-ceph-object-manifests
ceph: add securePort values for objectstore integration tests
2020-01-27 14:49:47 -07:00
Elise Gafford d7906266e1 ceph: remove validation on SecurePort for k8s <= 1.13
Versions of k8s prior to 1.14 do not support the nullable field in
CRD properties. As k8s 1.13 remains part of our support matrix and
as this field must remain nullable, we must remove CRD validation
of SecurePort. Range validation of this field has been moved into
the logic of rgw.validateStore.

Fixes: #4693
Signed-off-by: Elise Gafford <egafford@redhat.com>
2020-01-27 13:49:35 -05:00
Sébastien Han 227d2d527a ceph: osd store refactor
Multiple things:

1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.

2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.

3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.

Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-23 19:13:09 +01:00
Juan Miguel Olmo Martínez f62b2fa305 Ceph: MDS pod placement adheres to fault domain topology
With this modification is guaranteed that at least 1 MDS pod is going to be
placed in each of the available zones in the k8s cluster.
This would be improved in the future using:
<Pod Topology Spread Constraints> (still in alpha state since kubernetes V.1.16)

This modification obtain the number of zones in the k8s cluster and appends an
<antiaffinity> term using the topologyKey
<topology.kubernetes.io/zone> in the same number of MDS pods.
In the case that will be more MDS pods than zones, the <antiaffinity> term won't
be added in these extra pods.

Important Note:
The antiaffiniy term only will work with nodes labeled using the NEW label:
 <topology.kubernetes.io/zone> present in k8s V.1.17 clusters.
For previous versions of k8s clusters the label used is:
 <failure-domain.beta.kubernetes.io/zone>
And therefore, in this kind of clusters this modification won't work.

Resolves #4641

[test ceph]

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2020-01-17 19:01:48 +01:00
Travis Nielsen 6ad62bebf4 Merge pull request #4688 from leseb/external-cluster-crash
ceph: ability to disable crash controller
2020-01-16 11:13:48 -07:00
Sébastien Han bb214b2f27 ceph: ability to disable crash controller
A new CR setting has been introduced to disable the crash controller.
To disable it, add `disable: true` on the crashCollector section.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-16 16:05:22 +01:00
Sébastien Han 81dd279062 Merge pull request #4653 from leseb/pick-14.2.6
ceph: bump to 14.2.6
2020-01-15 17:44:52 +01:00