add RecoverAndLogException() helper to log panics with stack trace.
added defer call in all rook controller Reconcile() methods for
better error visibility in operator logs
Signed-off-by: Oded Viner <oviner@redhat.com>
All existing controller runtime watches are converted to use "typed"
handlers and predicates instead of operating on `client.Object`. The
intent is to be bug for bug equivalent with the existing logic while
replacing run time type assertions and switch statements with compile
time type constraints and type casts. In several cases, functions using
assertions were split up such that each function only handles a single
Kind at a time. It is hoped that this will improve readability and
maintainability while facilitating future refactoring such as migrating
some watches to using IndexFields.
Of particular note is that the massive switch statement in
`WatchControllerPredicate()` from
`pkg/operator/ceph/controller/predicate.go` has been replaced with
generics, reflection, and splitting the obc logic into its own predicate
function. There are still many helper functions operating on
`client.Object`. These were not updated unless required by the compiler
in order to limit the size of this change. The type safety of these
funcs should be improved as followup work.
It is strongly suggested that going forward, handlers and predicates
only handle a single Kind (generic or not) and that switches / type
assertions are heavily discouraged or forbidden. This PR removed all
but a single switch statement in a predicate, which should be addressed
in future work.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The operator settings loaded from the configmap have proven
inefficient for load time and frequently checking the configmap.
To avoid this ineffenciency, the configmap is only loaded once
each time it is created or updated. The values in the configmap
are applied as environment variables, which then are very efficient
to query throughout the various controllers, without needing
to be concerned about loading the configmap again.
Co-authored-by: Dmitry Mishin <dmitry.mishin@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Implement an allow list mechanism that disables potentially unsafe OBC
fields by default. OBC fields beyond `maxObjects` and `maxSize` don't
neatly fit into the OBC framework as it was originally envisioned and
implemented.
Some of the newly added configs could allow users to cause confusion for
themselves. Others might allow users to hijack others buckets. Some
might allow bricking the entire S3 store.
Out of an abundance of safety, allow-list the known-safe options by
default, and require administrators to enable potentially troublesome
options via the new operator-level config
`ROOK_OBC_ALLOW_ADDITIONAL_CONFIG_FIELDS`.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
This adds an operator config setting ROOK_REVISION_HISTORY_LIMIT
defaulting to kubernetes'value for RevisionHistoryLimit.
If configured, the provided value will be used as RevisionHistoryLimit
for all Deployments rook creates.
Fixes: #12722
Signed-off-by: Michael Adam <obnox@samba.org>
This new setting is of Boolean type and defaults to "false".
When set to "true", it changes the behavior of the
rook operator to
nable host network on all pods created by the cephcluster controller
new method to check the setting: opcontroller.EnForceHostNetwork()
Signed-off-by: Michael Adam <obnox@samba.org>
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.
Signed-off-by: subhamkrai <srai@redhat.com>
It's better to move most of discover daemon setting
from env to configmap rook-ceph-operator-config.
Although, we are moving to configmap, we keep reading settings
from env but the priority will be configmap settings.
Signed-off-by: subhamkrai <srai@redhat.com>
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.
This feature is intended to be used for testing purposes only.
Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.
If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.
Resolves#8778
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>