Commit Graph
103 Commits
Author SHA1 Message Date
subhamkrai 39b5c057ce core: faster recovery from rbd rwo node loss
in the existing node watcher, we'll check for node update
event and see if there are `out-of-service` taints are applied
and `ROOK_WATCH_FOR_NODE_FAILURE` is enabled in rook-ceph-operator-configmap,
if then we'll create the networkFence cr and delete the cr if nodes come back.
And, added the unit test too.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-07 21:16:25 +05:30
Javier 884d5e9855 osd: allow to use filter device
this feature is to allow the use of filters using the deviceFilter flag
by skipping devices that do not match the filter

Closes: #10340
Signed-off-by: Javier <sjavierlopez@gmail.com>
2023-04-21 13:13:59 -06:00
Shinya Hayashi 05875a3f4f osd: support loop devices for test clusters
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.

This feature is intended to be used for testing purposes only.

Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
2022-11-09 06:45:41 +00:00
dkeven dc95af5c33 osd: allow raw partitions to be picked up by discover daemon
Signed-off-by: dkeven <dkvvven@gmail.com>
2022-09-27 14:18:30 +08:00
Satoru Takeuchi e66b35aa55 Merge pull request #10230 from leseb/fix-9470
core: append filesystem property on disk
2022-05-10 09:07:09 +09:00
Sébastien Han 8328ec6d31 core: add mountpoint detection to the prepare pod
Amount various filters let's add a mountpoint property and skip the
device if its filesystem is mounted.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-05-09 15:26:50 +02:00
Sébastien Han 8c1bf0f3f0 core: append filesystem property on disk
When detecting the device property the filesystem was not passed but
used later to validate if the device should be taken or not. Now we read
the filesystem info from "lsblk" and populate it in the device type.

Closes: https://github.com/rook/rook/issues/9470
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-05-09 15:26:48 +02:00
Satoru Takeuchi 4c46dd1abd osd: fix disk uuid management
- disk UUID should only be got for "disk" type device.
- It's not necessary to fail prepare job when failing to get disk UUID
- We should consider that `sgdisk` reports UUID even if there is no GPT.

Closes: https://github.com/rook/rook/issues/9948

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-09 18:25:09 +09:00
Sébastien Han 05775b068d osd: handle removal of encrypted osd deployment
This is handling a tricky scenario where the OSD deployment is manually
removed and the OSD never reconvers. This is unlikely to happen, but
still OSD should be able to run after that action. Essentially after a
manual deletion, we need to run the prepare job again to re-hydrate the
OSD information so that the OSD deployment can be deployed.
On encryption, it is a little bit tricky since ceph-volume list again
the main block won't return anything, so we need to target the encrypted
block to list.
There is another case this PR does not handle, which is the removal of
the OSD deployment and then the node is restarted. This means that the
encrypted container is not opened anymore. However, opening it requires
more work like writing the key on the filesystem (if not coming from the
Kubernete secret, eg,. KMS vault) and then run luksOpen. This is an
extreme corner case probably not worth worrying about for now.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 10:51:09 +01:00
Blaine Gardner 01e2feaef5 ceph: get rid of dynamic clientset
Stop using the dynamic clientset in favor of the controller-runtime
clientset.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-19 17:05:38 -06:00
Travis Nielsen 87eee8afbc Merge pull request #8924 from leseb/rm-unused-client
core: remove unused clientset
2021-10-06 07:40:52 -06:00
Sébastien Han 723f144452 core: remove unused clientset
Removing ancient clientset which was not used anywhere any more.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-06 11:34:06 +02:00
Blaine Gardner 7586cea049 core: create TRACE_INSECURE log level
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.

If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.

Resolves #8778

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-01 12:09:19 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 8fdb9fe2b9 ceph: remove old network types
Those types are legacy, not really used anywhere and not needed anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-04 10:56:01 +02:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Blaine Gardner 21e290e003 ceph: implement dependencies for CephCluster
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-10 10:06:10 -06:00
Sébastien Han e6e9c8c901 ceph: do not print pointer memory address
When printing pointers of struct (regardless if it is also a slice or
not) we must implement a stringer interface so they can be displayed
properly, this is especially useful when looking at the prepare job
logs.

Before this patch:

```
2021-02-10 06:39:06.783508 D | inventory: discovered disks are [0xc00029c240 0xc00024a000 0xc00086c480 0xc00024a480]
```

After this patch:

```
2021-02-11 13:05:15.730500 D | inventory: discovered disks are:
2021-02-11 13:05:15.730636 D | inventory: &{Name:dm-0 Parent: HasChildren:false DevLinks:/dev/disk/by-id/dm-uuid-LVM-aIFRq5P3hwE2YhukbOmsEZoLY61HxvX62A1xxb6ajTymVE0N8S4Ga0KWlHSgpvNw /dev/disk/by-id/dm-name-ceph--e29bf82e--86f1--4c8e--8412--c468def5a844-osd--block--2044a1ea--6e5e--4423--88cf--8947bf6957d1 Size:32208060416 UUID:91a852c2-8727-4277-9d92-879ee1996503 Serial: Type:lvm Rotational:true Readonly:false Partitions:[] Filesystem:ceph_bluestore Vendor: Model: WWN: WWNVendorExtension: Empty:false CephVolumeData: RealPath:/dev/mapper/ceph--e29bf82e--86f1--4c8e--8412--c468def5a844-osd--block--2044a1ea--6e5e--4423--88cf--8947bf6957d1 KernelName:dm-0 Encrypted:false}
2021-02-11 13:05:15.730657 D | inventory: &{Name:dm-1 Parent: HasChildren:false DevLinks:/dev/disk/by-id/dm-uuid-LVM-GyLVW5UxDRcws8PwsbQuY2CjrQwdn1VxgPQjsKZ2pGVaS1r2Ct0OGZ4P80yEMtqW /dev/disk/by-id/dm-name-ceph--334c640f--7a7f--4fdb--9db4--2f9459e19f9a-osd--block--66ce11c8--20f0--4600--ae50--e8849c9a0af6 Size:32208060416 UUID:cdbcb218-1b03-4d8c-b7d2-25cb88e33b6a Serial: Type:lvm Rotational:true Readonly:false Partitions:[] Filesystem:ceph_bluestore Vendor: Model: WWN: WWNVendorExtension: Empty:false CephVolumeData: RealPath:/dev/mapper/ceph--334c640f--7a7f--4fdb--9db4--2f9459e19f9a-osd--block--66ce11c8--20f0--4600--ae50--e8849c9a0af6 KernelName:dm-1 Encrypted:false}
2021-02-11 13:05:15.730686 D | inventory: &{Name:dm-2 Parent: HasChildren:false DevLinks:/dev/disk/by-id/dm-name-ceph--a011b5c2--cc3f--497a--b19a--6d284461cbc4-osd--block--4de5491c--ecf8--455d--9858--547001fd0321 /dev/disk/by-id/dm-uuid-LVM-QGzgO4E228fvT3NsighNABGdvAKxQEhFOh9rKiDLnSOTk7jLZWu2r1RIMXe5GNqW Size:32208060416 UUID:2238ce24-79e4-41a0-b8ce-e381ac41d5b0 Serial: Type:lvm Rotational:true Readonly:false Partitions:[] Filesystem:ceph_bluestore Vendor: Model: WWN: WWNVendorExtension: Empty:false CephVolumeData: RealPath:/dev/mapper/ceph--a011b5c2--cc3f--497a--b19a--6d284461cbc4-osd--block--4de5491c--ecf8--455d--9858--547001fd0321 KernelName:dm-2 Encrypted:false}
2021-02-11 13:05:15.730713 D | inventory: &{Name:vda1 Parent:vda HasChildren:false DevLinks:/dev/disk/by-path/pci-0000:00:05.0-part1 /dev/disk/by-path/virtio-pci-0000:00:05.0-part1 /dev/disk/by-uuid/0fdf5e49-6b3c-42f2-a1c5-bf77e490dfa4 /dev/disk/by-partuuid/f9dd32c9-a4a5-654f-aee4-651276fdc754 /dev/disk/by-label/boot2docker-data Size:19998934528 UUID: Serial: Type:part Rotational:true Readonly:false Partitions:[] Filesystem:ext4 Vendor: Model: WWN: WWNVendorExtension: Empty:false CephVolumeData: RealPath:/dev/vda1 KernelName:vda1 Encrypted:false}
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-11 14:06:42 +01:00
Sébastien Han ad24990473 ceph: ability to abort orchestration
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):

* request for cluster deletion
* request for cluster upgrade

If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.

This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.

Operator log example:

```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
                Image:            "ceph/ceph:v15.2.5",
-               AllowUnsupported: true,
+               AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
-               Image:            "ceph/ceph:v15.2.5",
+               Image:            "ceph/ceph:v15.2.6",
                AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```

Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-25 15:32:04 +01:00
Satoru Takeuchi 88f4b9c148 ceph: support crypt type device in OSD on PVC
Ceph supported encrypted OSD in the following two ways.

- Encrypt OSD by Ceph itself
- Encrypt OSD by user

However, Rook supports only the first way. Let's support this way too.

With supporting this way, Rook can use variety of encryption methods.
For example, TPM can be used to manage encrypt key. In fact, I use TPM
as follows.

https://github.com/cybozu-go/sabakan/blob/57ff3ab560acb99d23aa8901c1018bf865f9ff4d/docs/disk_encryption.md#disk-encryption

Signed-off-by: UMEZAWA Takeshi <takeshi-umezawa@cybozu.co.jp>
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-08-03 21:17:06 +00:00
Travis Nielsen 32730123ea ceph: add support for multipath devices
Add mpath to the list of supported device types,
although it will only work for OSDs on PVCs.
OSDs on a raw device without PVCs have not
yet been tested.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-06 17:12:25 -06:00
Dmitry Yusupov e88704120b edgefs: raw disk empty check regression fix
Signed-off-by: Dmitry Yusupov <dmitry.yusupov@nexenta.com>
2020-04-29 11:48:09 -07:00
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Satoru Takeuchi afb6dbe1ed ceph: fix a wrong comment and a warning message about populating device info
PopulateDeviceInfo may fail for reasons other than lsblk failure.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-04-24 13:00:15 +00:00
Sébastien Han 20d1543507 ceph: add multus support
You can now use Rook along with Multus. Multus must be up and running
and the right ressources must exist such as NetworkAttachmentDefinition
CR.
The Cluster CR spec has new fields to work with multus:

network:
  provider: multus (or 'host' for hostNetworking)
  selectors:
    public: NetworkAttachmentDefinition name
    cluster: NetworkAttachmentDefinition name

If only a single NetworkAttachmentDefinition is provided Rook will use
both anyway for the Ceph traffic.

Please refer to the doc to learn more.

Closes: https://github.com/rook/rook/issues/4716
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-03 18:00:15 +02:00
morimoto-cybozu ab83738aac ceph: fix device path passed to "ceph-volume inventory"
This commit fixes the argument for "ceph-volume inventory".
When a device "/dev/mapper/foo" is being checked for its availability,
the argument should not be "/dev/foo" nor "/dev/dm-1".

Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
2020-03-27 11:24:47 +00:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Sébastien Han 227d2d527a ceph: osd store refactor
Multiple things:

1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.

2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.

3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.

Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-23 19:13:09 +01:00
Sébastien Han dd659de46f ceph: add partition support
We now support partitions via 2 ways:

* if `useAllDevice: true`: partitions will be taken into account and
presented as OSD candidate
* if specified in the cluster CR: it'll picked up as well

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-16 09:34:46 +01:00
dulltzandSatoru Takeuchi dfe45ac6b0 ceph: support OSD on PVC backed by LV
"OSD on PVC" doesn't work for PV backed by LV. Fixing this problem
by the following changes.

- Rook accepts LVM disk type.
- If a LV-backed device is passed, Rook/Ceph invokes
  "ceph-volume lvm prepare" with "--data vg/lv"
  instead of "--data /path/to/device".
- If a LV-backed device is passed, Rook/Ceph suppresses
  activation/deactivation of VG that owns this LV.

Fixes: https://github.com/rook/rook/issues/4185
Signed-off-by: dulltz <isrgnoe@gmail.com>
Co-authored-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2019-12-05 17:26:27 +00:00
Juan Miguel Olmo Martínez 7c942604f6 ceph: Get <ceph-volume inventory> data in dev. configmaps
**Description of your changes:**
This modification adds the information extracted from 'ceph-volume inventory':
command to the device configmaps generated by the discovery daemon when
"rook discover" starts with the new boolean "--use-ceph-volume" parameter.

Resolves #
https://github.com/rook/rook/issues/2606

Now the <cephVolumeData> field contains all the information returned
from <ceph-volume inventory> command.

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2019-11-06 10:22:56 +01:00
morimoto-cybozu 815927d943 clusterd: return explicit errors
PopulateDeviceInfo() in pkg/clusterd/disk.go returns nil as *sys.LocalDisk
if an error has occurred.  This causes a nil pointer exception at
getAvailableDevices() in pkg/daemon/ceph/osd/daemon.go.  At least the returned
value should be checked at the caller.
I added an explicit error to the return value of PopulateDeviceInfo().  This
naturally revokes an error check at the caller.
I modified PopulateDeviceUdevInfo() in the same file too.

Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
2019-10-15 08:30:45 +00:00
rohan47andAshish Ranjan d2f52aebe5 Adds support for storageClassDeviceSet in rook-ceph operator
- Added code to support StorageClassDeviceSet spec provided in the cluster-on-pvc.yaml
- The code reads the StorageClassDeviceSet spec and creates pvc based on the ‘count’ field for each device set.
- OSD prepare job is started for each PVC which activates the ceph-volume on each PVC
- Finally OSD is started on each of the PVC device.

Co-authored-by: rohan47 <rohgupta@redhat.com>
Co-authored-by: Ashish Ranjan <aranjan@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2019-08-12 09:24:13 -06:00
Sébastien Han 928c4a01da Merge pull request #3418 from InfuseAI/feature/add-linear-type-to-supported-disks
Add linear to support diskType
2019-07-11 17:35:46 +02:00
Ash Wu e6d43d9d0b Add linear to support diskType
Linear raid can be created using `mdadm --create --level=linear`,

By creating a linear raid disk on top of a logical volume,
we can pass the linear device to `ceph-volume lvm batch --prepare`
to provision a new OSD on top of the logical volume since
`ceph-volume` does not take lv as the data device.

Signed-off-by: Ash Wu <hSATAC@gmail.com>
2019-07-10 11:14:26 +08:00
Ash Wu b0eb246919 Remove unused clusterd.GetAvailableDevices
clusterd.GetAvailableDevices is not used anywhere.

Signed-off-by: Ash Wu <hSATAC@gmail.com>
2019-07-09 17:16:37 +08:00
Guy Margalit 2ce2382676 Refactor operator context init and allow to run local
Signed-off-by: Guy Margalit <guymguym@gmail.com>
Co-Authored-By: Sébastien Han <seb@redhat.com>
Co-Authored-By: Travis Nielsen <tnielsen@redhat.com>

This change is meant to allow running operators locally on a developer machine.
The idea is to allow faster development cycles by reducing the time and complexity of building -> deploying -> debugging on cluster.

For operators that rely only on kubernetes API this works easily - see cockroachdb and minio examples in development-flow doc.

The change includes:

- rook.NewContext() - Refactored to remove repeating initialization code that was copy-pasted in most of the operators in order to create the clusterd.Context and the Clientsets. Also it detects the mode of working in-cluster vs external and sets up the external mode with standard user config (~/.kube/config) and a job executor.
- rook.GetOperatorImage() - Refactor this repeating code in many operators to detect the operator pod image. Also added a global flag --operator-image that developers can use to override this when running locally.
- rook.TerminateOnError() - Added a convenient function.
2019-06-19 20:47:55 +03:00
travisn f7c266f7ec skip sgdisk detection if not found in nautilus image
Signed-off-by: travisn <tnielsen@redhat.com>
2018-12-06 21:28:57 -07:00
Huamin Chen 82425aafcc prepare osd in a job per node. Once all osds are prepared, store osd info in orchestration configmap.
Operator watches the configmap, starts one osd replica set per osd.

Signed-off-by: Huamin Chen <hchen@redhat.com>
2018-07-05 16:02:27 -06:00
Alexander Trost 6fb6c212c6 Don't set default for IP flags
Fixes the issue that when the old IPv4 flags were set, the new flags
were still used because of the default being the new flags not being
empty.

Added `NetworkInfo()` func to cmd/ceph to get simplified network info

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2018-06-27 17:20:25 +02:00
Jared Watts 48b879e0ad Merge pull request #1769 from ucloudlab/ipv6-support
support ipv6
2018-06-19 20:33:12 -07:00
travisn d5d3540497 osd: detect disk uuid with sgdisk
Signed-off-by: travisn <tnielsen@redhat.com>
2018-06-19 13:05:51 -07:00
Zhang Miaolei 100e80250f support ipv6 host-port pair
deprecate NetworkInfo.PublicAddrIPv4, use NetworkInfo.PublicAddr instead
deprecate NetworkInfo.ClusterAddrIPv4, use NetworkInfo.ClusterAddr instead
deprecate command flag public-ipv4, use public-ip instead
deprecate command flag private-ipv4, use private-ip instead
change ROOK_PUBLIC_IPV4 to ROOK_PUBLIC_IP
change ROOK_PRIVATE_IPV4 to ROOK_PRIVATE_IP

Signed-off-by: Zhang Miaolei <zmlcc@outlook.com>
2018-06-15 14:26:42 +08:00
travisn 4078b1b7f9 osd: device discovery to load part uuid after restart
Signed-off-by: travisn <tnielsen@redhat.com>
2018-06-01 21:03:03 -07:00
Huamin Chen 6c11ff52d4 add device discovery daemon to operator:
run "rook discover" on storage nodes and discover devices on each node. The discovered disks are saved in a per node configmap, local-device-nodename.
Device information consits of name and persistent names, uuid, partition, filesystem, rotational, readonly, size, etc.

Signed-off-by: Huamin Chen <hchen@redhat.com>
2018-05-07 17:49:02 +00:00
David González Ruiz eb59858124 Recognize crypt type as avaiable type and test the feature
Signed-off-by: David González Ruiz <godboole@gmail.com>
2018-02-01 15:13:31 +01:00
David González Ruiz 093d881b9f [OSD] Add crypt type to the array of supported block devices
[skip ci]

Signed-off-by: David González Ruiz <godboole@gmail.com>
2018-01-31 17:41:22 +01:00
Travis Nielsen e2bd9276bf consume the generated rook clientset 2017-12-11 14:35:07 -08:00
Travis Nielsen 0ad60c065b osd: refactor proc manager for only the osd daemon 2017-10-26 17:57:25 -07:00