Commit Graph
16 Commits
Author SHA1 Message Date
Oded Viner cf13deee6f core: log panics in controller reconcile functions
add RecoverAndLogException() helper to log panics with stack trace.
added defer call in all rook controller Reconcile() methods for
better error visibility in operator logs

Signed-off-by: Oded Viner <oviner@redhat.com>
2025-07-29 13:20:38 +03:00
Joshua Hoblitt 9f1ed201db core: typed watch handlers and predicates
All existing controller runtime watches are converted to use "typed"
handlers and predicates instead of operating on `client.Object`.  The
intent is to be bug for bug equivalent with the existing logic while
replacing run time type assertions and switch statements with compile
time type constraints and type casts. In several cases, functions using
assertions were split up such that each function only handles a single
Kind at a time. It is hoped that this will improve readability and
maintainability while facilitating future refactoring such as migrating
some watches to using IndexFields.

Of particular note is that the massive switch statement in
`WatchControllerPredicate()`  from
`pkg/operator/ceph/controller/predicate.go` has been replaced with
generics, reflection, and splitting the obc logic into its own predicate
function. There are still many helper functions operating on
`client.Object`. These were not updated unless required by the compiler
in order to limit the size of this change. The type safety of these
funcs should be improved as followup work.

 It is strongly suggested that going forward, handlers and predicates
 only handle a single Kind (generic or not) and that switches / type
 assertions are heavily discouraged or forbidden. This PR removed all
 but a single switch statement in a predicate, which should be addressed
 in future work.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2025-04-04 09:40:56 -07:00
Travis NielsenandDmitry Mishin d109dc9029 core: implement operator settings as env vars
The operator settings loaded from the configmap have proven
inefficient for load time and frequently checking the configmap.
To avoid this ineffenciency, the configmap is only loaded once
each time it is created or updated. The values in the configmap
are applied as environment variables, which then are very efficient
to query throughout the various controllers, without needing
to be concerned about loading the configmap again.

Co-authored-by: Dmitry Mishin <dmitry.mishin@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-02-24 12:12:45 -07:00
Blaine Gardner 0e33536539 object: disallow unsafe OBC fields by default
Implement an allow list mechanism that disables potentially unsafe OBC
fields by default. OBC fields beyond `maxObjects` and `maxSize` don't
neatly fit into the OBC framework as it was originally envisioned and
implemented.

Some of the newly added configs could allow users to cause confusion for
themselves. Others might allow users to hijack others buckets. Some
might allow bricking the entire S3 store.

Out of an abundance of safety, allow-list the known-safe options by
default, and require administrators to enable potentially troublesome
options via the new operator-level config
`ROOK_OBC_ALLOW_ADDITIONAL_CONFIG_FIELDS`.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2025-02-12 14:18:01 -07:00
Michael Adam ab8fd90aa6 core: add ROOK_REVISION_HISTORY_LIMIT operator setting
This adds an operator config setting ROOK_REVISION_HISTORY_LIMIT
defaulting to kubernetes'value for RevisionHistoryLimit.

If configured, the provided value will be used as RevisionHistoryLimit

for all Deployments rook creates.

Fixes: #12722

Signed-off-by: Michael Adam <obnox@samba.org>
2024-10-02 19:40:37 +02:00
Michael Adam e378588359 network: add a new operator config setting ROOK_ENFORCE_HOSTNETWORK
This new setting is of Boolean type and defaults to "false".

    When set to "true", it changes the behavior of the
     rook operator to
    nable host network on all pods created by the cephcluster controller

     new method to check the setting:  opcontroller.EnForceHostNetwork()

Signed-off-by: Michael Adam <obnox@samba.org>
2024-09-05 17:34:13 +02:00
subhamkrai d429ed8be4 build: update controller runtime to v0.18.4
this commit update cntrl runtime to v0.18.4 and other related deps/

Signed-off-by: subhamkrai <srai@redhat.com>
2024-07-05 09:05:12 +05:30
subhamkrai 28cc1ebc55 core: remove webhook & controller-runtime from apis
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-12-01 14:15:40 +05:30
subhamkrai fb39580067 operator: move most of discover pod setting to cm
It's better to move most of discover daemon setting
from env to configmap rook-ceph-operator-config.
Although, we are moving to configmap, we keep reading settings
from env but the priority will be configmap settings.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-16 22:17:47 +05:30
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
subhamkrai e9e2313126 webhook: disable webhook by default
sadly, we need to disable webhook by default until
we find solution to https://github.com/rook/rook/issues/10719

Signed-off-by: subhamkrai <srai@redhat.com>
2023-01-12 22:08:56 +05:30
subhamkrai 159fe3ae6f operator: disable webhook by default
disable webhook in downstream cluster.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-12-15 21:08:42 +05:30
Shinya Hayashi 05875a3f4f osd: support loop devices for test clusters
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.

This feature is intended to be used for testing purposes only.

Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
2022-11-09 06:45:41 +00:00
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
Blaine Gardner 7586cea049 core: create TRACE_INSECURE log level
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.

If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.

Resolves #8778

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-01 12:09:19 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00