This type was a string already and was just making us doing string()
calls all the time to it's not worth it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Now Kubernetes will perform liveness checks on mon, mds and osd daemons.
The command will:
* call the socket (check for existence)
* execute a command and check the return code (success if 0)
This handles the case where the daemon is stuck locally and
unresponsive. It's unlikely but not impossible.
These checks bring more robustness to the implementation.
rbd-mirror and nfs have been leftover for the following reason. The
rbd-mirror socket name is different from other daemons (could be fixed
though): /run/ceph/ceph-client.rbd-mirror.a.1.94362516231272.asok also,
the command to call would need to be changed from "status" to "rbd
mirror status" so we can keep this for a later.
The nfs ganesha has no socket only a PID file which doesn't mean much.
No PID means the process does not run so Kubernetes will already handle
this and the pod will crash loop.
Signed-off-by: Sébastien Han <seb@redhat.com>
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).
Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
In order to perform full automation of cephcluster deletion safely
and protect against catastrophic data loss, the end user must be
able to signify that they intend to irrecoverably delete the data
in their cluster. The cleanupPolicy field of the cluster spec is
intended to communicate this. This patch does not implement
automated deletion, but only creates the spec field and the
safety feature of halting orchestration other than deletion on a
cluster with a set cleanup policy value.
Partially-fixes: #3222
Signed-off-by: Elise Gafford <egafford@redhat.com>
The PG count on metadata pools should default to rgw_rados_pool_pg_num_min
instead of the more general default pg count. This means rgw pools
will default to 8 PGs instead of 32 PGs, which means a lot more pools
can be created before hitting the default PG limit.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds support for LVs to the device availability check
in the OSD prepare pod.
The availability of an LV is checked by "ceph-volume lvm list".
If it returns non-empty result, the LV is in use and not available.
Closes: https://github.com/rook/rook/issues/5075
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
This commit fixes the argument for "ceph-volume inventory".
When a device "/dev/mapper/foo" is being checked for its availability,
the argument should not be "/dev/foo" nor "/dev/dm-1".
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
The CrashCollector pod remains in pending state indefinitely after
running on a k8s node that was deleted. The code now deletes the
deployment after the node is deleted.
Signed-off-by: rohan47 <rohgupta@redhat.com>
If a deployment stays in pending we should give early by looking at
ProgressDeadlineExceeded, this will reduce the time to wait from 20 min
to 10 min because ProgressDeadlineExceeded default is 600 seconds.
Prior to this patch we would wait 20min since we take
currentDeployment.Spec.ProgressDeadlineSeconds which is typically 600
then retry every 2 seconds, which makes it 20min total.
Closes: https://github.com/rook/rook/issues/5090
Signed-off-by: Sébastien Han <seb@redhat.com>
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.
Signed-off-by: Sébastien Han <seb@redhat.com>
Let's not wait for the CephCluster to be done reconciling but instead
check for the Ceph cluster status, if it's closed to "ok" then we
proceed so HEALTH_OK and HEALTH_WARN are accepted.
Signed-off-by: Sébastien Han <seb@redhat.com>
Let's not apply the exponential backoff when waiting for the cluster to
be ready, let's only apply it when there is no cluster.
Closes: https://github.com/rook/rook/issues/5059
Signed-off-by: Sébastien Han <seb@redhat.com>
with older implementation, servicemonitor was not getting
updated due to missing resource version. This fix adds
resource version to the servicemonitor definition and
ensures that the object is properly created or
updated.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
The pools may already be created by the mgr module. If the pool specs
are not specified in the object CR, skip pool creation for the
object store. The pools are required to exist and the reconcile
will fail until they do exist.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook should always set the pgnum of the new block pool
to the default value("0"). Current implementation accidentally
works fine because newPool.Number is always 0 here.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The default pod anti-affinity for the mons that rook adds
automatically intended to be appended to any anti-affinity that
is specified in the cluster CR.
There is a bug in the ApplyToPodSpec() method that has long existed.
The issue is that when antiaffinity is appended, it will append not only
to the pod spec, but will modify the original placement spec. Thus, each
mon that is started will have one more antiaffinity clause than the previous mon.
This condition is rarely hit or noticed because it commonly is only hit
in the canary pods. Since these pods are immediately deleted,
there are no side effects of the duplicate affinity clauses. The place
where it becomes an issue is when the mons are backed by a PVC. In this case,
the mons do not have node affinity and will hit the code path that appends
to the antiaffinity and thus modifies the original antiaffinity.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is critical to get this information so that the label can be set
properly and also the configuration done appropriately.
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4941
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
The bump from 0.2 to 0.4 controller-runtime versions which happened
during the use of Go Modules (bumping Go version to 1.13) introduced an
issue where the deployment annotations keep getting changed during
watch update event.
The easy fix is to stop watching Deployment objects for now.
Signed-off-by: Sébastien Han <seb@redhat.com>
All processes created with the exec package will now only have debug logging
for the commands and their output. The osd prepare job relies heavily
on many of these commands so we can enable debug logging for the entire
job.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Long ago the ipv4 flags were renamed to public-ip and private-ip
so we can go ahead and remove the obsolete flags.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook no longer relies on its own process management, now we can rely
completely on Kubernetes to manage the pod lifecycle. The code
to check for running processes and replacement them hasn't been
used for a long while.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>