In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
In order to perform full automation of cephcluster deletion safely
and protect against catastrophic data loss, the end user must be
able to signify that they intend to irrecoverably delete the data
in their cluster. The cleanupPolicy field of the cluster spec is
intended to communicate this. This patch does not implement
automated deletion, but only creates the spec field and the
safety feature of halting orchestration other than deletion on a
cluster with a set cleanup policy value.
Partially-fixes: #3222
Signed-off-by: Elise Gafford <egafford@redhat.com>
The integration tests are failing to run when the disks
are not properly cleaned up and the bluestore label
is still found. Now we run sgdisk to more completely
zap the disk.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Run the tests on latest nautilus with the tag ceph/ceph:v14 so we
are constantly testing on the latest release.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit fixes the argument for "ceph-volume inventory".
When a device "/dev/mapper/foo" is being checked for its availability,
the argument should not be "/dev/foo" nor "/dev/dm-1".
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
There are many namespace fields in ClusterRole{,Binding}. However,
ClusterRole{,Binding} are not namespaced. So we can remove these.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The pool cleanup only needs to happen for an individual pool.
No need to query the block images in all pools. One of the rgw
pools is periodically causing a hang when it is queried,
but there is no need to query for it when we are cleaning
up the pool tests.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
This is needed to allow the ceph manager rook module to read the output
of the <lsmcli> command used to switch disk identification light on/off on
physical disk devices
[test ceph]
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
Current YugabyteDB operator code creates Master and TServer pods without any resource requests/limits (specifically CPU and memory).
This causes the operator to run into soft/hard memory limit issue. The fix adds recommended resource requests and limits as defaults to the pods it creates.
Closes: https://github.com/yugabyte/yugabyte-db/issues/3884
Signed-off-by: Sameer Kulkarni <samkulkarni20@gmail.com>
The Ceph mgr modules need access to restart the daemon pods
when the configuration changes for the daemons.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests for yugabyte were pointing to an image
outside of the rook repo. Since that image changed unexpectedly
it broke the rook integration tests. The tests need to use
the image built by rook.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator logs are necessary for troubleshooting why the
pools may still exist at the end of the integration tests.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The .metadata.generation field is updated if and only if the value at the .spec subpath changes.
Additionally, if the spec does not change, .metadata.generation is not updated.
This is a must-have for the controller-runtime work, if we don't have
this, the generation field will always be incremented, resulting in a
endless reconcile loop.
Signed-off-by: Sébastien Han <seb@redhat.com>
The CI is failing currently in the OSD-on-pv scenario due to a change
in v14.2.8. Until that is resolved, we need to unblock the tests
by running against v14.2.7.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With a finalizer on the pools, the pools were not always being purged
during the integration tests. Now the multicluster suite will ensure
its pool is purged.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
As of Octopus, Ceph will prevent you from creating a pool with a
replica size of 1. Allowing such pool could lead to data loss, so enable
the new option: requireSafeReplicaSize: false if you are **ABSOLUTELY**
certain that is what you want.
Closes: https://github.com/rook/rook/issues/4889
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests have failed intermittently due to needing just
a little longer to start the file test pod. This increases the
wait timeout.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The job to wipe devices fails the integration tests periodically.
Now the job will be retried once if it fails to retry wiping the
disks.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.
A number of other test improvements are also made for code
readability, test stability, and removing unused options.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The replication and erasure code settings are mutually exclusive
so must both be treated as optional by the schema validation.
The minimum must be set to 0 to treat them as optional.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests were always running two MDS daemons active,
with two standby. This is now parameterized so the test can request
how many MDS daemons to run. The smoke suite will run two active
and the rest will just run a single active MDS.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration test was disabled due to removing support of pre-ceph-volume
OSDs as well as OSDs on directories. Now we re-enable the tests, with the
following approach:
- The base install in Rook v1.1 with the latest mimic release that
had c-v support
- The upgrade goes from v1.1 to Rook v1.2 then Rook master
- The final upgrade step is from mimic to nautilus
For efficiency the skipUpgradeChecks flag is added.
Logs are also collected between each upgrade step to improve
troubleshooting.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit enables testing OSDs over PVCs by modifying the
multi-cluster test to use PVC for provisioning OSDs when `manual`
storageClass is present in the cluster.
Signed-off-by: Ashish Ranjan <aranjan@redhat.com>
Fixed conditions getting resetted after the operator restart.
Did the changes which required to implement conditions on the rook ceph cluster
Conditions will eliminate the current status.State and incorporates a type which
provides much more description to the current status of the cluster.
Signed-off-by: Nizamudeen <nia@redhat.com>