This commit adds support for LVs to the device availability check
in the OSD prepare pod.
The availability of an LV is checked by "ceph-volume lvm list".
If it returns non-empty result, the LV is in use and not available.
Closes: https://github.com/rook/rook/issues/5075
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
This commit fixes the argument for "ceph-volume inventory".
When a device "/dev/mapper/foo" is being checked for its availability,
the argument should not be "/dev/foo" nor "/dev/dm-1".
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Long ago the ipv4 flags were renamed to public-ip and private-ip
so we can go ahead and remove the obsolete flags.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook no longer relies on its own process management, now we can rely
completely on Kubernetes to manage the pod lifecycle. The code
to check for running processes and replacement them hasn't been
used for a long while.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit is to handle all those unhandled errors which raises the gosec warning.
Fixed G104: Unhandled Errors are handled now
Signed-off-by: Nizamudeen <nia@redhat.com>
This commit is also part of gosec error handling which handles the following issue:
Fixed G304: Potential File inclusion by cleaning up the file path
Signed-off-by: Nizamudeen <nia@redhat.com>
We now support the addition of the PVC that acts as a metadata device
for a given OSD.
For this, you need to create a new `volumeClaimTemplates`, its name must
be "metadata" otherwise, Rook won't pick it up.
A template will look like this:
```
volumeClaimTemplates:
- metadata:
name: data
spec:
resources:
requests:
storage: 10Gi
# IMPORTANT: Change the storage class depending on your environment (e.g. local-storage, gp2)
storageClassName: gp2
volumeMode: Block
accessModes:
- ReadWriteOnce
- metadata:
name: metadata
spec:
resources:
requests:
storage: 6Gi
# IMPORTANT: Change the storage class depending on your environment (e.g. local-storage, gp2)
storageClassName: gp2
volumeMode: Block
accessModes:
- ReadWriteOnce
```
We now map block and block.db directly inside the container instead of
running ceph-volume activate. This is much cleaner.
Closes: https://github.com/rook/rook/issues/3852
Signed-off-by: Sébastien Han <seb@redhat.com>
Multiple things:
1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.
2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.
3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.
Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
We now support partitions via 2 ways:
* if `useAllDevice: true`: partitions will be taken into account and
presented as OSD candidate
* if specified in the cluster CR: it'll picked up as well
Signed-off-by: Sébastien Han <seb@redhat.com>
The retry messages are too verbose with the stack trace
since the aggregated messages were added. This change
turns off the verbose message.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Fix security issues identified by TrailOfBits:
Logging of sensitive information in debug mode
Problematic debug lines which can contain sensitive data have been removed
Resolves: #4568
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
Since https://github.com/rook/rook/pull/4219, lv devices are now
presented to ceph-volume when preparing the device.
So on this initial run, this won't fail because the dm hasn't been
created yet but if an orchestration is re-triggered, the prepare pod
ensures idempotency with the ceph-volume batch command.
Unfortunately, c-v seems to have a bug where it doesn't read the dm to
detect whether or not they are ceph members already.
Ceph bug: https://tracker.ceph.com/issues/43209
Signed-off-by: Sébastien Han <seb@redhat.com>
"OSD on PVC" doesn't work for PV backed by LV. Fixing this problem
by the following changes.
- Rook accepts LVM disk type.
- If a LV-backed device is passed, Rook/Ceph invokes
"ceph-volume lvm prepare" with "--data vg/lv"
instead of "--data /path/to/device".
- If a LV-backed device is passed, Rook/Ceph suppresses
activation/deactivation of VG that owns this LV.
Fixes: https://github.com/rook/rook/issues/4185
Signed-off-by: dulltz <isrgnoe@gmail.com>
Co-authored-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
**Description of your changes:**
This modification adds the information extracted from 'ceph-volume inventory':
command to the device configmaps generated by the discovery daemon when
"rook discover" starts with the new boolean "--use-ceph-volume" parameter.
Resolves #
https://github.com/rook/rook/issues/2606
Now the <cephVolumeData> field contains all the information returned
from <ceph-volume inventory> command.
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
OSD on PVC does not upgrade when the user upgrades the ceph version on cluster-on-pvc yaml.
The solution incudes:
- upgrade osd prepare and daemon pods on upgrade
- skip c-v prepare if filesystem is already present on the pvc device.
- skip lvm release in case of upgrade.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
When both PARTNAME and ID_PART_ENTRY_NAME are specified, rook takes
PARTNAME in account, overriding ID_PART_ENTRY_NAME. As PARTNAME is not
always updated by the kernel after sgdisk --change-name, it's better to
do the opposite. It also makes GetDevicePartitions coherent with
parsePartLabel (called by GetPartitionLabel).
Signed-off-by: Mikaël Cluseau <mikael.cluseau@gmail.com>
- Added code to support StorageClassDeviceSet spec provided in the cluster-on-pvc.yaml
- The code reads the StorageClassDeviceSet spec and creates pvc based on the ‘count’ field for each device set.
- OSD prepare job is started for each PVC which activates the ceph-volume on each PVC
- Finally OSD is started on each of the PVC device.
Co-authored-by: rohan47 <rohgupta@redhat.com>
Co-authored-by: Ashish Ranjan <aranjan@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Linear raid can be created using `mdadm --create --level=linear`,
By creating a linear raid disk on top of a logical volume,
we can pass the linear device to `ceph-volume lvm batch --prepare`
to provision a new OSD on top of the logical volume since
`ceph-volume` does not take lv as the data device.
Signed-off-by: Ash Wu <hSATAC@gmail.com>
Signed-off-by: Guy Margalit <guymguym@gmail.com>
Co-Authored-By: Sébastien Han <seb@redhat.com>
Co-Authored-By: Travis Nielsen <tnielsen@redhat.com>
This change is meant to allow running operators locally on a developer machine.
The idea is to allow faster development cycles by reducing the time and complexity of building -> deploying -> debugging on cluster.
For operators that rely only on kubernetes API this works easily - see cockroachdb and minio examples in development-flow doc.
The change includes:
- rook.NewContext() - Refactored to remove repeating initialization code that was copy-pasted in most of the operators in order to create the clusterd.Context and the Clientsets. Also it detects the mode of working in-cluster vs external and sets up the external mode with standard user config (~/.kube/config) and a job executor.
- rook.GetOperatorImage() - Refactor this repeating code in many operators to detect the operator pod image. Also added a global flag --operator-image that developers can use to override this when running locally.
- rook.TerminateOnError() - Added a convenient function.
the existing exec interface with timeout is effectively the same as
ExecuteCommandWithOutput plus a timeout. this patch adds a variant of
ExecuteCommandWithOutputFile that uses a timeout.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
If someone sets limits to pod, we want to ensure the possible
experience, so we want to make sure that people do not configure
inapropriate values for certain daemons.
We decide to fail if the memory.limit is too low.
Signed-off-by: Sébastien Han <seb@redhat.com>
All devices detected by the discovery pod were being passed to the OSD provisioning pod
thus not always honoring the desired device list that should be provisioned.
Now the provisioning pod will be given the desired state from the crd,
then apply that state depending on the actual devices detected.
Also added a helper to ensure OSDsPerDevice is always valid.
Signed-off-by: travisn <tnielsen@redhat.com>
The io.MultiReader(r1, r2) reads r1 until EOF before moving on to r2.
When used for stdout/stderr stderr will not be written to the log until
stdout reaches EOF. This patch reads stderr in a go routine so that the
two streams can be interleaved properly.
Fixes: #2479
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
Previous code tried to interpret the string output as a `fmt.Sprintf`
string formatter, mangling `%` into `%!(MISSING)`, etc.
Thanks to go, I don't belive these are exploitable (unlike the similar
error in C).
Since this seemed to be a common error in the codebase, I did a quick
audit by visually inspecting the results of `git grep 'f([^"]'`. I
don't have a good suggestion for automated tests to prevent this in
future :(
Example error (look for `(MISSING)`):
```
E0927 05:31:07.618429 11227 driver-call.go:237] Failed to unmarshal output for command: unmount, output: "2018-09-27 05:31:07.191711 I | exec: Running command: df --type ceph /var/lib/kubelet/pods/95461479-c216-11e8-bcf0-02030782ac80/volumes/ceph.rook.io~rook/oe-scratch\n2018-09-27 05:46:43.808596 I | Filesystem 1K-blocks Used Available Use%!M(MISSING)ounted on\n2018-09-27 05:46:43.808659 I | 10.107.25.147:6790,10.109.173.79:6790,10.104.85.255:6790:/ 151678976 49410048 102268928 33%!/(MISSING)var/lib/kubelet/pods/95461479-c216-11e8-bcf0-02030782ac80/volumes/ceph.rook.io~rook/oe-scratch\n{\"status\":\"Success\"}\n", error: invalid character '-' after top-level value
```
Signed-off-by: Angus Lees <gus@inodes.org>
Current Ceph volume manager code doesn't take the builtin
kernel module into the account, so the `modinfo` will fail
in this case. Also, we need to add error handling code to
make sure the volume manager does work when proceeding to
continue, otherwise we will encounter the `rbd` related issue
in the following step even we have an artificial `agent` pod
in running state.
Signed-off-by: Dennis Chen <dennis.chen@arm.com>
Fix some basic spellcheck errors. Also remove trailing spaces and make
sure files have a newline (my editor does automatically).
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
deprecate NetworkInfo.PublicAddrIPv4, use NetworkInfo.PublicAddr instead
deprecate NetworkInfo.ClusterAddrIPv4, use NetworkInfo.ClusterAddr instead
deprecate command flag public-ipv4, use public-ip instead
deprecate command flag private-ipv4, use private-ip instead
change ROOK_PUBLIC_IPV4 to ROOK_PUBLIC_IP
change ROOK_PRIVATE_IPV4 to ROOK_PRIVATE_IP
Signed-off-by: Zhang Miaolei <zmlcc@outlook.com>
run "rook discover" on storage nodes and discover devices on each node. The discovered disks are saved in a per node configmap, local-device-nodename.
Device information consits of name and persistent names, uuid, partition, filesystem, rotational, readonly, size, etc.
Signed-off-by: Huamin Chen <hchen@redhat.com>