Some ceph health errors should not block the reconcile
of the cluster. Mgr modules do not have cause to block
the reconcile, as the cluster can usually work even
if a module is failing.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator settings loaded from the configmap have proven
inefficient for load time and frequently checking the configmap.
To avoid this ineffenciency, the configmap is only loaded once
each time it is created or updated. The values in the configmap
are applied as environment variables, which then are very efficient
to query throughout the various controllers, without needing
to be concerned about loading the configmap again.
Co-authored-by: Dmitry Mishin <dmitry.mishin@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Implement an allow list mechanism that disables potentially unsafe OBC
fields by default. OBC fields beyond `maxObjects` and `maxSize` don't
neatly fit into the OBC framework as it was originally envisioned and
implemented.
Some of the newly added configs could allow users to cause confusion for
themselves. Others might allow users to hijack others buckets. Some
might allow bricking the entire S3 store.
Out of an abundance of safety, allow-list the known-safe options by
default, and require administrators to enable potentially troublesome
options via the new operator-level config
`ROOK_OBC_ALLOW_ADDITIONAL_CONFIG_FIELDS`.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
This adds an operator config setting ROOK_REVISION_HISTORY_LIMIT
defaulting to kubernetes'value for RevisionHistoryLimit.
If configured, the provided value will be used as RevisionHistoryLimit
for all Deployments rook creates.
Fixes: #12722
Signed-off-by: Michael Adam <obnox@samba.org>
This new setting is of Boolean type and defaults to "false".
When set to "true", it changes the behavior of the
rook operator to
nable host network on all pods created by the cephcluster controller
new method to check the setting: opcontroller.EnForceHostNetwork()
Signed-off-by: Michael Adam <obnox@samba.org>
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.
This feature is intended to be used for testing purposes only.
Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This allows the catch-22 situation where the filesystem cannot be
reconciled because there is no MDS but there is no MDS because the
operator has not reconciled the filesystem and brought up the MDS pods.
Closes#5967, #5846
Signed-off-by: Lalit Maganti <lalitm@google.com>