Commit Graph
354 Commits
Author SHA1 Message Date
Travis Nielsen 53ed11f15b build: remove the edgefs operator from rook
The EdgeFS operator has been deprecated for some time in Rook.
If the replacement is added back to Rook it can be completed
according to the new guidelines in the documentation.
https://rook.io/docs/rook/master/storage-providers.html

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-25 17:51:24 -07:00
Travis Nielsen 2294851ca3 build: remove the cockroachdb operator from rook
The cockroachDB operator has not had community support in Rook.
Therefore, the time has come to deprecate and remove it.
If the sources are still needed, there is always git history.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-20 17:28:58 -07:00
Travis Nielsenandshenjiatong a0a27bc426 ceph: apply deviceClass properly to new osds
The deviceClass property was being ignored when creating the
non-pvc OSDs. Now the deviceClass will be specified as a property
for individual devices, all devices on a node, or all OSDs in the
cluster, depending on the level where the config is applied in the
cluster CR.

Co-authored-by: shenjiatong <yshxxsjt715@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-11 10:16:00 -07:00
Sébastien Han ad24990473 ceph: ability to abort orchestration
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):

* request for cluster deletion
* request for cluster upgrade

If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.

This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.

Operator log example:

```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
                Image:            "ceph/ceph:v15.2.5",
-               AllowUnsupported: true,
+               AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
-               Image:            "ceph/ceph:v15.2.5",
+               Image:            "ceph/ceph:v15.2.6",
                AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```

Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-25 15:32:04 +01:00
Arun Kumar Mohan ded16f779d ceph: changes for 'sigs.k8s.io/sig-storage-lib-external-provisioner/v6'
Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:03 +05:30
Travis Nielsen b9f692a56e ceph: disable the discovery daemon by default
The discovery daemon is not needed in most scenarios, therefore we disable it
by default. More and more clusters are moving to the cluster-on-pvc scenario
which certainly does not need the local discovery. Even where clusters are not
running on PVCs, the discovery is not needed since the device discovery is again
performed in the osd prepare job.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-10 13:16:52 -07:00
Jared Watts 8d445c919d Merge pull request #6202 from prksu/nfs-quota
nfs: implement disk quota
2020-11-02 08:34:54 -08:00
Jonas Schäfer 7ace6ac255 ceph: make root= CRUSH label value configurable
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.

By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.

To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.

The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.

Fixes #4993.

Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
2020-10-29 09:22:41 +01:00
Ahmad Nurus S 72bf7873c9 nfs: implement disk quota
Signed-off-by: Ahmad Nurus S <prksu.sh@gmail.com>
2020-10-19 10:20:12 +07:00
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
subhamkrai f9fafe62d4 ceph: handle golangci-lint linter unused
this commit will enable one more linter
in golangci-lint.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 22:53:24 +05:30
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
Blaine Gardner 76f1d9944e ceph: fix drive group deployment failure
Drive Groups will fail to deploy due to a message: `failed to parse
device list (""): failed to JSON unmarshal configured devices (""):
unexpected end of JSON input`

This is due to trying to JSON unmarshal an empty string which does not
contain any devices to provision when Drive Groups are specified. Fix so
that an empty string represents no non-Drive Group devices.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-09-16 09:32:06 -06:00
Servesha Dudhgaonkar 0c323439a6 ceph: purge a down osd with a job
If an OSD is down and needs to be removed from the cluster,
a job can be run that will remove the OSD deployment
and purge the OSD from ceph. If the OSD is still up,
the purge will be rejected.

To purge multiple OSDs at the same time, the OSD IDs can be
specified as a comma-separated list.

Signed-off-by: Servesha Dudhgaonkar <sdudhgao@redhat.com>
2020-09-01 09:36:20 +02:00
Madhu Rajanna 1989e0d8c5 ceph: remove csi support for kubernetes 1.13
as the kubernetes 1.13 is EOL and there is no major
functionalities available in 1.13 (resize,snapshot,clone
metrics etc) we are removing the support for kubernetes
for the same.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-08-05 21:19:59 +05:30
Sébastien Han 89b3225441 ceph: add encryption support for osd pvc
We can now encrypted OSD device that were provisioned via a storage
class using the PV interface.
The encryption works at the storageClassDeviceSets level, which means we
can have encrypted and non-encrypted sets.
Using the new key `encrypted` we can turn it on.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-05 10:06:11 +02:00
subhamkrai c938849cf8 ceph: closing file which are open for writing
using defer for closing file which are open
for writing is not safe. so closing file again
following  below steps:
1. open files
2. defer file.close()
3. write
4.file.close()

these will make sure files are closed.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-30 22:37:56 +05:30
subhamkrai 0279025e9e ceph: handling all the gosec errors
a few of the gosec errors were left. so
this commit will resolve all the errors.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-30 08:19:55 +05:30
subhamkrai de8e93274d ceph: handling gosec errors code g104
this commit handles all the gosec g104
(i.e Audit errors not checked) error.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-28 23:52:23 +05:30
subhamkrai bcd7faed4e ceph: suppress gosec errors for g204, g304, g101
this commit suppress the gosec errors for

g204: Audit use of command execution.
g304: File path provided as taint input.
g101: Look for hard coded credentials.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-28 11:10:48 +05:30
Travis Nielsen cf53467380 core: suppress gosec errors for closing files
To ensure a file handle is closed, we defer the close command
so it is guaranteed to run when the method returns. The closing
of the file handle is not going to fail in our usage since we aren't
using the SetDeadline on the files that would cancel a request
and return an error.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-24 12:31:13 -06:00
Sébastien Han 10ed662bdb ceph: enhancement disk sanitize options
Now, we can not only cleanup monitor data, logs and crashes but the
disks too. As part of the cleanupPolicy CR spec, we have a new setting
called sanitizeDisk which holds more details:

* method: indicates if the entire disk should be sanitized or simply ceph's metadata.
Possible choices are 'complete' or 'quick' (default)
* dataSource: indicate where to get random bytes from to write on the disk.
Possible choices are 'zero' (default) or 'random'
Using random sources will consume entropy from the system and will take much more time then the zero source
* iteration: overwrite N times instead of the default (1). Takes an integer value

See the documentation for more details.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-21 14:07:11 +02:00
Blaine Gardner 825f1e3eb2 ceph: support device names with special chars
Pass desired devices to the OSD provisioning container by
JSON-marshalling/-unmarshalling a disk ID with StorageConfig settings.
This will allow disks to be specified that contain special characters.
Notably, this will support /dev/disk/by-path/pci-HHHH:HH:HH.H devices
that were unsupported previously due to the format which separated
devices from params with colons (:).

Fixes #5056
Fixes #5535

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-17 09:44:26 -06:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen 4681f9e73d ceph: consolidate ceph config and client packages
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:43 -06:00
Travis Nielsen 631b13b906 ceph: refactor creds used by operator
The operator should only connect to ceph with a single set of creds.
In a converged cluster this will be the admin creds and in an external
cluster it will be lower-privileged creds. Independent clusters were
implemented with a separate set of creds. To simplify the code these
are now merged to a single set.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:42 -06:00
Blaine Gardner 7117fc12b7 ceph: osd: add drive groups spec to cluster CR
Add the ability to provision Ceph OSDs with Drive Groups.
This adds Drive Groups to the CephCluster CRD, and it sets code
in place for propagating this config to the OSD provisioning pod.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-14 09:43:03 -06:00
Travis Nielsen 7dcd64f1a2 Merge pull request #5761 from vbnrh/rook-ac-controller-runtime
ceph: implement controller-runtime for admission controllers
2020-07-13 16:31:23 -06:00
Vineet Badrinath 98874ccf4d ceph: implement controller-runtime for admission controllers
moves the admission-controller server to use controller-runtime library for better support, maintainibility of code.

Signed-off-by: Vineet Badrinath <vbadrina@redhat.com>
2020-07-14 01:10:03 +05:30
Ahmad Nurus S 90da5ec7b9 nfs: add validation admission webhook using controller-runtime
Signed-off-by: Ahmad Nurus S <prksu.sh@gmail.com>
2020-07-03 14:33:36 +07:00
Ahmad Nurus S f7f8270f58 nfs: rewrite nfs operator controller to use controller-runtime
Signed-off-by: Ahmad Nurus S <prksu.sh@gmail.com>
2020-07-03 14:33:25 +07:00
Vineet Badrinath ce1003aef8 ceph: adds scripts and components to support admission controllers
adds deploy.sh script to deploy validatingwebhookconfiguration and create secrets.
adds new command ceph admission-controller to start webhook servers.
adds validation for various rook custom resources

Signed-off-by: Vineet Badrinath <vbadrina@redhat.com>
2020-06-24 14:59:00 +05:30
Sébastien Han 68e62836c5 ceph: use newer octopus time format
radosgw-admin as of Octopus uses a different time format, it uses
"2006-01-02T15:04:05.999999999Z". It's close from RFC3339 but not quite
the same. This change is needed in order for the operator image (with a
Ceph Octopus based image) to perform radosgw-admin call correctly.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-17 23:00:17 +02:00
Sébastien Han aa8bf4c763 ceph: expose operator ceph base image
We now have a new env variable ROOK_CEPH_BASE_IMAGE_VERSION which
contains the ceph version on the operator image.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-16 15:13:40 +02:00
Sébastien Han c799e11032 ceph: enhancement cleanup with disks
Now, we can not only cleanup monitor data, logs and crashes but the
disks too. As part of the cleanupPolicy CR spec, we have a new setting
called sanitizeDisks which holds more details:

* confirmation: the confirmation message to sanitize disks,
use "yes-really-sanitize-disks" to confirm

This will **only** wipe the metadata, so it's a fast cleanup allowing
you to re-install later but won't remove all the data from the drive.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-05-27 18:14:02 +02:00
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Sébastien Han 20d1543507 ceph: add multus support
You can now use Rook along with Multus. Multus must be up and running
and the right ressources must exist such as NetworkAttachmentDefinition
CR.
The Cluster CR spec has new fields to work with multus:

network:
  provider: multus (or 'host' for hostNetworking)
  selectors:
    public: NetworkAttachmentDefinition name
    cluster: NetworkAttachmentDefinition name

If only a single NetworkAttachmentDefinition is provided Rook will use
both anyway for the Ceph traffic.

Please refer to the doc to learn more.

Closes: https://github.com/rook/rook/issues/4716
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-03 18:00:15 +02:00
Santosh Pillai e1a76bfb96 ceph: Cluster CleanupPolicy: Delete correct mon directories under the dataDirHostPath
In case of multiple clusters, we don't want to delete all the mon directories under the dataDirHostPath during cluster cleanup.
This PR deletes the mon directory only if the montior secret key matches.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-04-03 10:01:36 +05:30
Santosh Pillai 2bfd7c42bc ceph: cleanup cluster.Spec.DataDirHostPath on cluster deletion
In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-03-30 19:36:01 +05:30
Umanga Chapagain 0e932c15eb Ceph: add CSI configurations to ConfigMap
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-27 15:17:30 +05:30
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 7d2811bfb6 ceph: remove obsolete ipv4 command line flags
Long ago the ipv4 flags were renamed to public-ip and private-ip
so we can go ahead and remove the obsolete flags.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
moricho 2beb731a7e go: move to gomodules
This switches Rook to use `go mod` instead of `dep` for
dependencies management.

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:47:53 +09:00
Stefan Haas b3cc4aeb44 ceph: ceph-csi version detection #3824
Checks the version of the configured ceph-csi image while starting the operator. The operator will fail if the image is not supported.
Added an additional parameter to operator to disable the check e.g. to test not yet supported csi images.

Signed-off-by: Stefan Haas <shaas@suse.com>
2020-02-28 14:14:40 +01:00
Sébastien Han 227d2d527a ceph: osd store refactor
Multiple things:

1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.

2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.

3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.

Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-23 19:13:09 +01:00
n.fraison 7bac027796 ceph: ensure crush-location osd command args is always the same
Currently each time the ceph operator restart, osds also restart.
This is due to a change in the crush-location args with location not ordered
At each restart the location order can change and lead to that kind of change in the osd deployment
104c105
<                             "--crush-location=root=default host=hostname datacenter=PAR pod=2 rack=2",
---
>                             "--crush-location=root=default host=hostname pod=2 rack=2 datacenter=PAR",
Sorting the topology to ensure stability of this command arg

Signed-off-by: n.fraison <n.fraison@criteo.com>
2020-01-21 15:04:16 +01:00
rohan47 fda2b5d96c osd: Added check for --crush-location
Added a check for --crush-location, if --crush-location is not present,
the node topology is not applied and the operator will need to determine
the value based on the node labels instead of skipping this setting.

Signed-off-by: rohan47 <rohgupta@redhat.com>
2020-01-15 02:47:34 +05:30
Travis Nielsen 078e87a722 ceph: fix non-portable osd crush host name
The host name of an OSD in the CRUSH map should be the real
host name for non-portable OSDs. It was incorrectly being set
to the PVC name. Now the non-portable OSDs based on PVCs will
corretly have the host name set to the node name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-13 12:17:49 -07:00