Commit Graph
71 Commits
Author SHA1 Message Date
parth-gr 01456e8683 osd: add a new devicetype label for osds
currently device class and device type
labels were clubbed with eachother
create a seperate label, as these two labels
can be seperate

Signed-off-by: parth-gr <partharora1010@gmail.com>
2024-12-05 16:56:04 +05:30
yingshanghuangqiao 3d672a726a core: fix some comments
Signed-off-by: yingshanghuangqiao <yingshanghuangqiao@foxmail.com>
2024-07-22 22:09:52 +08:00
Satoru Takeuchi 3e34ebeff8 osd: handle global or node-local device class configuration correctly
Rook uses global or node-local device class configuration if device-level
configurations doesn't exist. However, after introducing device-class
level resource configuration, rook set the default value of device class.
So global or node-level device class configuration has never been used
after that.

Closes: https://github.com/rook/rook/issues/11871
Closes: https://github.com/rook/rook/issues/11826

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2023-03-27 04:28:45 +00:00
Shinya Hayashi 05875a3f4f osd: support loop devices for test clusters
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.

This feature is intended to be used for testing purposes only.

Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
2022-11-09 06:45:41 +00:00
Satoru Takeuchi 7e571f6114 osd: support OSD on logical volume in host-based cluster
Rook supports raw mode OSD in host-based cluster. So we can also
support OSD on logical volume in this kind of cluster.

Logical volumes aren't picked by filters (i.e. `useAllDevices: true`
and `device{Path,}Filter` to avoid unwanted LV consumption on upgrade.

Closes: https://github.com/rook/rook/issues/2047

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-07-11 21:08:50 +00:00
Satoru Takeuchi e66b35aa55 Merge pull request #10230 from leseb/fix-9470
core: append filesystem property on disk
2022-05-10 09:07:09 +09:00
Sébastien Han 8328ec6d31 core: add mountpoint detection to the prepare pod
Amount various filters let's add a mountpoint property and skip the
device if its filesystem is mounted.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-05-09 15:26:50 +02:00
Sébastien Han 77b2e10c14 core: add more debug info to the prepare pod
When running the device property commands, let's also log their results
for better debugging/visibility.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-05-09 15:26:50 +02:00
Satoru Takeuchi 4c46dd1abd osd: fix disk uuid management
- disk UUID should only be got for "disk" type device.
- It's not necessary to fail prepare job when failing to get disk UUID
- We should consider that `sgdisk` reports UUID even if there is no GPT.

Closes: https://github.com/rook/rook/issues/9948

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-09 18:25:09 +09:00
Sébastien Han 05775b068d osd: handle removal of encrypted osd deployment
This is handling a tricky scenario where the OSD deployment is manually
removed and the OSD never reconvers. This is unlikely to happen, but
still OSD should be able to run after that action. Essentially after a
manual deletion, we need to run the prepare job again to re-hydrate the
OSD information so that the OSD deployment can be deployed.
On encryption, it is a little bit tricky since ceph-volume list again
the main block won't return anything, so we need to target the encrypted
block to list.
There is another case this PR does not handle, which is the removal of
the OSD deployment and then the node is restarted. This means that the
encrypted container is not opened anymore. However, opening it requires
more work like writing the key on the filesystem (if not coming from the
Kubernete secret, eg,. KMS vault) and then run luksOpen. This is an
extreme corner case probably not worth worrying about for now.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 10:51:09 +01:00
Sébastien Han 1e9d24ac77 ceph: print the c-v output when inventory command fails
Hopefully, this will give us more hints when a failure occurs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-13 18:08:06 +02:00
Sébastien Han 4471f3a92e osd: do not hide errors
The previous exit 32 check for loop device is 5 years old. Also, if the
device cannot be read it will be skipped anyway so let's report the
error and not hide it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-07 17:39:58 +02:00
Sébastien Han b7f55362b2 ceph: print the output on errors
Sometimes the error does not tell much, so as `exit status 1` and
printing the output along returning the error is useful.

For instance, I saw a job failing with no osd and the prepare job had
those lines:

```
exec: Running command: lsblk /dev/sdb1 --bytes --nodeps --pairs ....
inventory: skipping device "sdb1". exit status 1
```

We need to understand more about the lsblk issue.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-09 12:42:01 +02:00
Blaine Gardner c8a2db368a ceph: disable raw mode for disks
Even if we can use raw mode, do NOT use raw mode on disks. Ceph bluestore disks can
sometimes appear as though they have "phantom" Atari (AHDI) partitions created on them
when they don't in reality. This is due to a series of bugs in the Linux kernel when it
is built with Atari support enabled. This behavior does not appear for raw mode OSDs on
partitions, and we need the raw mode to create partition-based OSDs. We cannot merely
skip creating OSDs on "phantom" partitions due to a bug in `ceph-volume raw inventory`
which reports only the phantom partitions (and malformed OSD info) when they exist and
ignores the original (correct) OSDs created on the raw disk.

Resolves https://github.com/rook/rook/issues/7940

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-07-22 09:59:56 -06:00
Blaine Gardner 7f516b9e3d ceph: ignore atari partitions when scanning disks
Ceph bluestore raw disks can sometimes appear as though they have Atari
(AHDI) partitions on them. If a disk has Atari partitions, we just
ignore them as though they don't exist. This should be a safe assumption
since the hardware was last manufactured in 1992 and likely can't run
Kubernetes.

If we don't ignore the Atari partitions, Rook can create a new OSD on a
disk that is already running an OSD, corrupting the first and possibly
also the latest OSD. This can happen an arbitrary number of times per
disk.

More info: https://github.com/rook/rook/issues/7940

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-30 09:42:33 -06:00
Denis Egorenko 8a556fd847 ceph: add ability to set resource limit for OSDs based on device classes
Currently it is not possible to set resource limits based on device classes
for different OSDs. Now adding an ability to use predefined keys for main
cluster spec Resource section to reflect resource limits for different OSDs
with different device classes.

Closes: https://github.com/rook/rook/issues/8007
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
2021-06-08 17:46:41 +04:00
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
Sébastien Han 89b3225441 ceph: add encryption support for osd pvc
We can now encrypted OSD device that were provisioned via a storage
class using the PV interface.
The encryption works at the storageClassDeviceSets level, which means we
can have encrypted and non-encrypted sets.
Using the new key `encrypted` we can turn it on.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-05 10:06:11 +02:00
Satoru Takeuchi 88f4b9c148 ceph: support crypt type device in OSD on PVC
Ceph supported encrypted OSD in the following two ways.

- Encrypt OSD by Ceph itself
- Encrypt OSD by user

However, Rook supports only the first way. Let's support this way too.

With supporting this way, Rook can use variety of encryption methods.
For example, TPM can be used to manage encrypt key. In fact, I use TPM
as follows.

https://github.com/cybozu-go/sabakan/blob/57ff3ab560acb99d23aa8901c1018bf865f9ff4d/docs/disk_encryption.md#disk-encryption

Signed-off-by: UMEZAWA Takeshi <takeshi-umezawa@cybozu.co.jp>
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-08-03 21:17:06 +00:00
Travis Nielsen 32730123ea ceph: add support for multipath devices
Add mpath to the list of supported device types,
although it will only work for OSDs on PVCs.
OSDs on a raw device without PVCs have not
yet been tested.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-06 17:12:25 -06:00
morimoto-cybozu 0c3439c707 ceph: check LV availability by "ceph-volume lvm list"
This commit adds support for LVs to the device availability check
in the OSD prepare pod.
The availability of an LV is checked by "ceph-volume lvm list".
If it returns non-empty result, the LV is in use and not available.

Closes: https://github.com/rook/rook/issues/5075
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
2020-03-27 13:30:08 +00:00
morimoto-cybozu ab83738aac ceph: fix device path passed to "ceph-volume inventory"
This commit fixes the argument for "ceph-volume inventory".
When a device "/dev/mapper/foo" is being checked for its availability,
the argument should not be "/dev/foo" nor "/dev/dm-1".

Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
2020-03-27 11:24:47 +00:00
Satoru Takeuchi 8a13310ddb ceph: Consider the various paths of sgdisk
The path of sgdisk depends on the systems.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-03-25 12:27:14 +00:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Sébastien Han c08c3ced05 ceph: add support for metadata PVC for OSD on PVC
We now support the addition of the PVC that acts as a metadata device
for a given OSD.
For this, you need to create a new `volumeClaimTemplates`, its name must
be "metadata" otherwise, Rook won't pick it up.

A template will look like this:

```
volumeClaimTemplates:
- metadata:
    name: data
  spec:
    resources:
      requests:
        storage: 10Gi
    # IMPORTANT: Change the storage class depending on your environment (e.g. local-storage, gp2)
    storageClassName: gp2
    volumeMode: Block
    accessModes:
      - ReadWriteOnce
- metadata:
    name: metadata
  spec:
    resources:
      requests:
        storage: 6Gi
    # IMPORTANT: Change the storage class depending on your environment (e.g. local-storage, gp2)
    storageClassName: gp2
    volumeMode: Block
    accessModes:
      - ReadWriteOnce
```

We now map block and block.db directly inside the container instead of
running ceph-volume activate. This is much cleaner.

Closes: https://github.com/rook/rook/issues/3852
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-11 17:30:02 +01:00
Sébastien Han 227d2d527a ceph: osd store refactor
Multiple things:

1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.

2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.

3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.

Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-23 19:13:09 +01:00
Sébastien Han dd659de46f ceph: add partition support
We now support partitions via 2 ways:

* if `useAllDevice: true`: partitions will be taken into account and
presented as OSD candidate
* if specified in the cluster CR: it'll picked up as well

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-16 09:34:46 +01:00
Juan Miguel Olmo Martínez 7b3e2b006e ceph: Do not log sensitive information in debug mode
Fix security issues identified by TrailOfBits:
Logging of sensitive information in debug mode

Problematic debug lines which can contain sensitive data have been removed

Resolves: #4568

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2020-01-03 17:16:25 +01:00
Sébastien Han 04ad1636c8 ceph: reject ceph dm devices
Since https://github.com/rook/rook/pull/4219, lv devices are now
presented to ceph-volume when preparing the device.
So on this initial run, this won't fail because the dm hasn't been
created yet but if an orchestration is re-triggered, the prepare pod
ensures idempotency with the ceph-volume batch command.
Unfortunately, c-v seems to have a bug where it doesn't read the dm to
detect whether or not they are ceph members already.
Ceph bug: https://tracker.ceph.com/issues/43209

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-09 16:17:26 +01:00
dulltzandSatoru Takeuchi dfe45ac6b0 ceph: support OSD on PVC backed by LV
"OSD on PVC" doesn't work for PV backed by LV. Fixing this problem
by the following changes.

- Rook accepts LVM disk type.
- If a LV-backed device is passed, Rook/Ceph invokes
  "ceph-volume lvm prepare" with "--data vg/lv"
  instead of "--data /path/to/device".
- If a LV-backed device is passed, Rook/Ceph suppresses
  activation/deactivation of VG that owns this LV.

Fixes: https://github.com/rook/rook/issues/4185
Signed-off-by: dulltz <isrgnoe@gmail.com>
Co-authored-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2019-12-05 17:26:27 +00:00
Guangming Wang 950af10e62 cleanup: cleanup codebase
device.go: replace string func Index with Contains.

Signed-off-by: Guangming Wang <guangming.wang@daocloud.io>
2019-11-07 19:45:43 +08:00
Juan Miguel Olmo Martínez 7c942604f6 ceph: Get <ceph-volume inventory> data in dev. configmaps
**Description of your changes:**
This modification adds the information extracted from 'ceph-volume inventory':
command to the device configmaps generated by the discovery daemon when
"rook discover" starts with the new boolean "--use-ceph-volume" parameter.

Resolves #
https://github.com/rook/rook/issues/2606

Now the <cephVolumeData> field contains all the information returned
from <ceph-volume inventory> command.

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2019-11-06 10:22:56 +01:00
Santosh Pillai 8ea693a740 ceph: fix operator reconcile to restart osds daemons
OSD on PVC does not upgrade when the user upgrades the ceph version on cluster-on-pvc yaml.
The solution incudes:
   - upgrade osd prepare and daemon pods on upgrade
   - skip c-v prepare if filesystem is already present on the pvc device.
   - skip lvm release in case of upgrade.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2019-10-23 16:56:15 +05:30
Mikaël Cluseau 0267433484 fix: PARTNAME override ID_PART_ENTRY_NAME
When both PARTNAME and ID_PART_ENTRY_NAME are specified, rook takes
PARTNAME in account, overriding ID_PART_ENTRY_NAME. As PARTNAME is not
always updated by the kernel after sgdisk --change-name, it's better to
do the opposite. It also makes GetDevicePartitions coherent with
parsePartLabel (called by GetPartitionLabel).

Signed-off-by: Mikaël Cluseau <mikael.cluseau@gmail.com>
2019-10-22 09:23:48 +11:00
rohan47andAshish Ranjan d2f52aebe5 Adds support for storageClassDeviceSet in rook-ceph operator
- Added code to support StorageClassDeviceSet spec provided in the cluster-on-pvc.yaml
- The code reads the StorageClassDeviceSet spec and creates pvc based on the ‘count’ field for each device set.
- OSD prepare job is started for each PVC which activates the ceph-volume on each PVC
- Finally OSD is started on each of the PVC device.

Co-authored-by: rohan47 <rohgupta@redhat.com>
Co-authored-by: Ashish Ranjan <aranjan@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2019-08-12 09:24:13 -06:00
Ash Wu e6d43d9d0b Add linear to support diskType
Linear raid can be created using `mdadm --create --level=linear`,

By creating a linear raid disk on top of a logical volume,
we can pass the linear device to `ceph-volume lvm batch --prepare`
to provision a new OSD on top of the logical volume since
`ceph-volume` does not take lv as the data device.

Signed-off-by: Ash Wu <hSATAC@gmail.com>
2019-07-10 11:14:26 +08:00
travisn bdc3cf8146 osd: fix the device filter and improve device provisioning reliability
All devices detected by the discovery pod were being passed to the OSD provisioning pod
thus not always honoring the desired device list that should be provisioned.
Now the provisioning pod will be given the desired state from the crd,
then apply that state depending on the actual devices detected.
Also added a helper to ensure OSDsPerDevice is always valid.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-02-26 16:11:25 -07:00
travisn f7c266f7ec skip sgdisk detection if not found in nautilus image
Signed-off-by: travisn <tnielsen@redhat.com>
2018-12-06 21:28:57 -07:00
Alexander Trost b12d1e82cf Use truncated node name "everywhere"
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2018-11-23 10:02:35 +01:00
Dennis Chen 9cfafaa0a7 Enhancement of Ceph volume manager kmod detection
Current Ceph volume manager code doesn't take the builtin
kernel module into the account, so the `modinfo` will fail
in this case. Also, we need to add error handling code to
make sure the volume manager does work when proceeding to
continue, otherwise we will encounter the `rbd` related issue
in the following step even we have an artificial `agent` pod
in running state.

Signed-off-by: Dennis Chen <dennis.chen@arm.com>
2018-07-31 16:56:40 +08:00
Blaine Gardner fe5fb394c7 Fix spellcheck and trailing space/newline issues
Fix some basic spellcheck errors. Also remove trailing spaces and make
sure files have a newline (my editor does automatically).

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-07-26 11:36:09 -06:00
Huamin Chen dc1e827cce get partlabel from udev info PARTNAME
Signed-off-by: Huamin Chen <hchen@redhat.com>
2018-07-05 16:09:21 -06:00
travisn f1ad88e0ac tests: enable osds on devices
Signed-off-by: travisn <tnielsen@redhat.com>
2018-07-05 16:09:20 -06:00
Huamin Chen 82425aafcc prepare osd in a job per node. Once all osds are prepared, store osd info in orchestration configmap.
Operator watches the configmap, starts one osd replica set per osd.

Signed-off-by: Huamin Chen <hchen@redhat.com>
2018-07-05 16:02:27 -06:00
Evgeniy Klemin 759ec55f64 Fix retrieve device label
Signed-off-by: Evgeniy Klemin <evgeniy.klemin@gmail.com>
2018-06-27 13:31:46 +03:00
travisn d5d3540497 osd: detect disk uuid with sgdisk
Signed-off-by: travisn <tnielsen@redhat.com>
2018-06-19 13:05:51 -07:00
Aayush Sarva 80f3e8ebf4 Replace custom Awk and remove related tests
Signed-off-by: Aayush Sarva <checkaayush@gmail.com>

Return error if field not found

Signed-off-by: Aayush Sarva <checkaayush@gmail.com>
2018-06-19 05:03:28 +05:30
Huamin Chen 6c11ff52d4 add device discovery daemon to operator:
run "rook discover" on storage nodes and discover devices on each node. The discovered disks are saved in a per node configmap, local-device-nodename.
Device information consits of name and persistent names, uuid, partition, filesystem, rotational, readonly, size, etc.

Signed-off-by: Huamin Chen <hchen@redhat.com>
2018-05-07 17:49:02 +00:00