In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
In order to perform full automation of cephcluster deletion safely
and protect against catastrophic data loss, the end user must be
able to signify that they intend to irrecoverably delete the data
in their cluster. The cleanupPolicy field of the cluster spec is
intended to communicate this. This patch does not implement
automated deletion, but only creates the spec field and the
safety feature of halting orchestration other than deletion on a
cluster with a set cleanup policy value.
Partially-fixes: #3222
Signed-off-by: Elise Gafford <egafford@redhat.com>
The integration tests are failing to run when the disks
are not properly cleaned up and the bluestore label
is still found. Now we run sgdisk to more completely
zap the disk.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Run the tests on latest nautilus with the tag ceph/ceph:v14 so we
are constantly testing on the latest release.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit fixes the argument for "ceph-volume inventory".
When a device "/dev/mapper/foo" is being checked for its availability,
the argument should not be "/dev/foo" nor "/dev/dm-1".
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
There are many namespace fields in ClusterRole{,Binding}. However,
ClusterRole{,Binding} are not namespaced. So we can remove these.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
This is needed to allow the ceph manager rook module to read the output
of the <lsmcli> command used to switch disk identification light on/off on
physical disk devices
[test ceph]
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
Current YugabyteDB operator code creates Master and TServer pods without any resource requests/limits (specifically CPU and memory).
This causes the operator to run into soft/hard memory limit issue. The fix adds recommended resource requests and limits as defaults to the pods it creates.
Closes: https://github.com/yugabyte/yugabyte-db/issues/3884
Signed-off-by: Sameer Kulkarni <samkulkarni20@gmail.com>
The Ceph mgr modules need access to restart the daemon pods
when the configuration changes for the daemons.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests for yugabyte were pointing to an image
outside of the rook repo. Since that image changed unexpectedly
it broke the rook integration tests. The tests need to use
the image built by rook.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator logs are necessary for troubleshooting why the
pools may still exist at the end of the integration tests.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The .metadata.generation field is updated if and only if the value at the .spec subpath changes.
Additionally, if the spec does not change, .metadata.generation is not updated.
This is a must-have for the controller-runtime work, if we don't have
this, the generation field will always be incremented, resulting in a
endless reconcile loop.
Signed-off-by: Sébastien Han <seb@redhat.com>
The CI is failing currently in the OSD-on-pv scenario due to a change
in v14.2.8. Until that is resolved, we need to unblock the tests
by running against v14.2.7.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With a finalizer on the pools, the pools were not always being purged
during the integration tests. Now the multicluster suite will ensure
its pool is purged.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
As of Octopus, Ceph will prevent you from creating a pool with a
replica size of 1. Allowing such pool could lead to data loss, so enable
the new option: requireSafeReplicaSize: false if you are **ABSOLUTELY**
certain that is what you want.
Closes: https://github.com/rook/rook/issues/4889
Signed-off-by: Sébastien Han <seb@redhat.com>
The job to wipe devices fails the integration tests periodically.
Now the job will be retried once if it fails to retry wiping the
disks.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.
A number of other test improvements are also made for code
readability, test stability, and removing unused options.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The replication and erasure code settings are mutually exclusive
so must both be treated as optional by the schema validation.
The minimum must be set to 0 to treat them as optional.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration test was disabled due to removing support of pre-ceph-volume
OSDs as well as OSDs on directories. Now we re-enable the tests, with the
following approach:
- The base install in Rook v1.1 with the latest mimic release that
had c-v support
- The upgrade goes from v1.1 to Rook v1.2 then Rook master
- The final upgrade step is from mimic to nautilus
For efficiency the skipUpgradeChecks flag is added.
Logs are also collected between each upgrade step to improve
troubleshooting.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit enables testing OSDs over PVCs by modifying the
multi-cluster test to use PVC for provisioning OSDs when `manual`
storageClass is present in the cluster.
Signed-off-by: Ashish Ranjan <aranjan@redhat.com>
Fixed conditions getting resetted after the operator restart.
Did the changes which required to implement conditions on the rook ceph cluster
Conditions will eliminate the current status.State and incorporates a type which
provides much more description to the current status of the cluster.
Signed-off-by: Nizamudeen <nia@redhat.com>
The integration tests only have a single node so the upgrade
checks do not provide any real safety for the cluster during
the upgrade. We will be able to improve the reliability and speed
of the tests by disabling these checks.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Versions of k8s prior to 1.14 do not support the nullable field in
CRD properties. As k8s 1.13 remains part of our support matrix and
as this field must remain nullable, we must remove CRD validation
of SecurePort. Range validation of this field has been moved into
the logic of rgw.validateStore.
Fixes: #4693
Signed-off-by: Elise Gafford <egafford@redhat.com>
Somehow the job was stuck on 'fdisk -l', the logic to wipe the block
device changed and we don't need that command anymore.
We wipe the disk in the PV loop.
Signed-off-by: Sébastien Han <seb@redhat.com>
Multiple things:
1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.
2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.
3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.
Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
Running integration tests locally on k8s 1.13 was causing an error
with parsing the command line flags. We can't call flags.Parse()
from an init() method.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In the Ceph upgrade integration test, upgrade from Rook v1.0 with Ceph
Mimic v13 to install legacy disk-based OSDs. Then upgrade Ceph to
Nautilus, then Rook to v1.1, then v1.2, then v1.3 (master currently) to
make sure that Rook is able to run legacy OSDs created without
ceph-volume.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>