This type was a string already and was just making us doing string()
calls all the time to it's not worth it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Now Kubernetes will perform liveness checks on mon, mds and osd daemons.
The command will:
* call the socket (check for existence)
* execute a command and check the return code (success if 0)
This handles the case where the daemon is stuck locally and
unresponsive. It's unlikely but not impossible.
These checks bring more robustness to the implementation.
rbd-mirror and nfs have been leftover for the following reason. The
rbd-mirror socket name is different from other daemons (could be fixed
though): /run/ceph/ceph-client.rbd-mirror.a.1.94362516231272.asok also,
the command to call would need to be changed from "status" to "rbd
mirror status" so we can keep this for a later.
The nfs ganesha has no socket only a PID file which doesn't mean much.
No PID means the process does not run so Kubernetes will already handle
this and the pod will crash loop.
Signed-off-by: Sébastien Han <seb@redhat.com>
With 1.18 in the test matrix we need to update which versions
will run the different ceph test suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The go modules are changed by the install of the junit
tool for the unit tests. Before the release build we need
to ensure that the modules are tidied.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The master build is publishing with the dirty tag so we
need to ensure that the files do not remain modified.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).
Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
In order to perform full automation of cephcluster deletion safely
and protect against catastrophic data loss, the end user must be
able to signify that they intend to irrecoverably delete the data
in their cluster. The cleanupPolicy field of the cluster spec is
intended to communicate this. This patch does not implement
automated deletion, but only creates the spec field and the
safety feature of halting orchestration other than deletion on a
cluster with a set cleanup policy value.
Partially-fixes: #3222
Signed-off-by: Elise Gafford <egafford@redhat.com>
It is not necessary to have 2 same environment variables and not desired as
changing only one may lead to the second variable being applied with unexpected
value. It is better to have only 1 ROOK_HOSTPATH_REQUIRES_PRIVILEGED environment
variable.
Closes: https://github.com/rook/rook/issues/5107
Signed-off-by: Adler Fleurant <Adler.Fleurant@PicoChange.com>
The recent change to go modules is causing a file to change
during the build, which results in a dirty tag to be
generated.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This uses grep, sort and awk to auto generate the `make help` menu based
on comments after each target.
Removed `clean.all` and `prune.all` targets as they are not available
anymore.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
The PG count on metadata pools should default to rgw_rados_pool_pg_num_min
instead of the more general default pg count. This means rgw pools
will default to 8 PGs instead of 32 PGs, which means a lot more pools
can be created before hitting the default PG limit.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests are failing to run when the disks
are not properly cleaned up and the bluestore label
is still found. Now we run sgdisk to more completely
zap the disk.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Run the tests on latest nautilus with the tag ceph/ceph:v14 so we
are constantly testing on the latest release.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds support for LVs to the device availability check
in the OSD prepare pod.
The availability of an LV is checked by "ceph-volume lvm list".
If it returns non-empty result, the LV is in use and not available.
Closes: https://github.com/rook/rook/issues/5075
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
This fix is added to ensure that everything works as expected
when user tries to create ObjectStoreUser before creating
ObjectStore itself. ObjectStoreUser will wait for ObjectStore
to be up and running.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
This commit fixes the argument for "ceph-volume inventory".
When a device "/dev/mapper/foo" is being checked for its availability,
the argument should not be "/dev/foo" nor "/dev/dm-1".
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
Every developer should commit all the files that need to be changed
in a given set of changes. The go.sum file was not being generated
as expected since the update to go modules. Since the Jenkinsfile
was running make mod.check, the build was ending up with a modified
file, which caused a dirty tag on the build. Now it is expected
that each PR will merge with the corresponding go.sum changes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CrashCollector pod remains in pending state indefinitely after
running on a k8s node that was deleted. The code now deletes the
deployment after the node is deleted.
Signed-off-by: rohan47 <rohgupta@redhat.com>
If a deployment stays in pending we should give early by looking at
ProgressDeadlineExceeded, this will reduce the time to wait from 20 min
to 10 min because ProgressDeadlineExceeded default is 600 seconds.
Prior to this patch we would wait 20min since we take
currentDeployment.Spec.ProgressDeadlineSeconds which is typically 600
then retry every 2 seconds, which makes it 20min total.
Closes: https://github.com/rook/rook/issues/5090
Signed-off-by: Sébastien Han <seb@redhat.com>