If there are at least three OSDs on a single node, we should
treat it as a potential production cluster and perform
the ok-to-stop checks during reconcile. Otherwise,
it may cause instability during upgrades on
single-node clusters.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
this PR improves the error logging when ok-to-stop requests fail
in the Rook operator. Instead of only showing a generic exit status
Signed-off-by: Oded Viner <oviner@redhat.com>
This change refactors inline roundup code to a function
roundup_size_MiB that takes a size(in bytes) as argument and returns
the smallest number of Mibibytes larger than or equal to the given size.
Signed-off-by: Michael Adam <obnox@samba.org>
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.
Additionally, it silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
and string format errors found by golangci-lint, while at it.
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
currently device class and device type
labels were clubbed with eachother
create a seperate label, as these two labels
can be seperate
Signed-off-by: parth-gr <partharora1010@gmail.com>
we were adding the device class as the
disk rotational type, but device class
can be user defined, so if user has
added the device class, use that value
Signed-off-by: parth-gr <partharora1010@gmail.com>
if the import from the peer cluster doesnt
work, because of some reasons like
cluster health, network, etc
the import cmd stuck forever and which
forever stucks the blockppol reconcile
Signed-off-by: parth-gr <partharora1010@gmail.com>
enable and disable the status of rados namespace
mirroring by looking at statusCheck spec of blockpool
Signed-off-by: parth-gr <partharora1010@gmail.com>
The upgrade of the mds daemons pauses to wait for the
stopping of standby daemons before the filesystem upgrade
is completed. If there are multiple filesystems, there
could be standbys for other filesystems that will not be
stopped at the time of the upgrade of another filesystem.
Thus, the standbys may not upgrade and be stuck on the
previous version. Now, the standby upgrade will only wait
for standbys for the same filesystem.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is similar to #14052 we did for radosnamespace
and this is an extension to support cleanup
at the blockpool level to cleanup the images
and the snapshots in a pool.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Modify the CR to allow mirroring of an rados namespace
to a differently named namespace on the remote cluster
1) enable rados namesapce mirroring only
if the blockpool mirrroing is enabled
2) disable blockpool mirroing only if
all the namesapce mirroing is disabled
if the rbd mirroring fails and ceph version is not supported
provide a error message with supported version details
and reason of failing
Signed-off-by: parth-gr <partharora1010@gmail.com>
Do not force delete resources in validation cleanup. Force deletion can
result in CNI IP address management (IPAM) addresses not being released
cleanly, exhausting them. This can then lead to Rook cluster deploy
failure if the CNI IPAM isn't able to garbage collect the IPs quickly.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Given that Ceph Quincy (v17) is past end of life,
remove Quincy from the supported Ceph versions,
examples, and documentation.
Supported versions now include only Reef and Squid.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
default application name is updated inside the `CreatePool` method. Send
pool spec as address in order to preserve this change.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
The crush rules may have a negative step num.
Rook had assumed negative values were not possible,
but just had not been encountered previously in
a custom crush rule.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When doing the clean-up of raw device OSDs, it was not taking metadata
and wal devices into account. This commit discover them via the output
of `ceph-volume raw list`, which is returning `device_db` and/or
`device_wal` if they are present.
After discovering the metadata and wal device it also clean them and
close the encrypted device if any.
Signed-off-by: Mathias Chapelain <mathias.chapelain@proton.ch>
Normally the device class of an OSD is determined at provisioning
and is not updated thereafter. In some scenarios the admin may
want to force update the device class to a new value. The device
classes can be updated by first setting allowDeviceClassUpdate
in the storage spec of the cephcluster, then updating the
device class specified on the deviceSets or other OSDs.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Reset the Multus validation tool debounce time to its intended 30 second
value. It was changed to 5 for testing, and the change was accidentally
committed.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
In order to help users check that they have implemented the newly-added
Multus host configuration prerequisites, add a check to the validation
tool to verify connectivity.
Because users who are already running clusters with Multus enabled, add
a flag that allows users to only check for host configuration
prerequisites. This mode will not start the large number of clients that
would normally be started because those clients could disrupt a running
Rook cluster negatively.
Host checking pods require host network access. Many Kubernetes
distributions have pod security features enabled. In order to allow
non-Vanilla distros to run this tool, allow specifying a service account
that pods will run as, which can be configured by the admin to allow
test pods.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Updating the device class swallowed any error if updated
for the pool. The error was not even logged, so we couldn't
troubleshoot why the new crush rule was not applied.
Log the error for troubleshooting and also fail the pool
reconcile since the desired configuration was not applied.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Pools in stretch clusters must all specify the same
CRUSH rule. No pools can use a different rule. When there
is a change in the device class, we do not even expect to update
the crush rules in a stretch cluster. Different device classes
are not supported in stretch clusters, and it's expected to be
a homogenous environment. Therefore, skip all crush rule updates
in stretch clusters.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
While adding a new encryption key to slot 1
if there exists a key in slot 1 which is not
equal to the one we want to update it with,
we kill the slot and then add the new key to it.
While killing the slot, the existing code uses
the new key, which is not valid in such cases.
This patch modifies the code to use the key in
slot 0 (the one that we know works) to kill the slot.
Signed-off-by: Niraj Yadav <niryadav@redhat.com>
this lead to creation of 2 crush rules because of difference in
deviceClass names:
eg:
creating a new crush rule for changed deviceClass
("default~hdd"-->"hdd") on crush rule "test-crush-bug-az-ab-4"
Signed-off-by: Deepika Upadhyay <deepikaupadhyay01@gmail.com>
this will be disabled by default but in scenarios where the user want to
update failureDomain, DeviceClass etc, this option can be enabled, to be
noted this can lead to lot of data rebalancing and remapping. Use with
caution
Signed-off-by: Deepika Upadhyay <deepika.upadhyay@clyso.com>
Add IPv6 support for multus validation tool. Also test that IPv6 support
works by specifying one of the NetAttachDefs with an IPv6 address range.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
When the clusters reach full, nearfull, or backfill full thresholds
ceph will raise health warnings and stop allowing IO or backfill
depending on the threshold. These settings require special ceph
commands instead of being generic ceph config. Allow these settings
to be set from the CephCluster CR in the spec.storage
section.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The following printcolumns were added
CephBlockPoolRadosNamespace: Phase,
BlockPoolName, Age
CephFilesystemSubVolumeGroup: Phase,
FilesystemName, PinningConfig, Age
Also improved the error messages in
SubVolumeGroupClient
Signed-off-by: NymanRobin <nyman.robin@gmail.com>