Commit Graph
38 Commits
Author SHA1 Message Date
travisn 0240ee8563 core: skip reconcile if override configmap is empty
If the configmap rook-config-override is empty,
there is no need to trigger the reconcile to update
the ceph daemons. This configmap update is causing
unnecessary reconciles periodically in some clusters
even when it is empty.

Signed-off-by: travisn <tnielsen@redhat.com>
2024-02-01 14:29:20 -07:00
Jiffin Tony Thottan 3da2331fa3 object: watch updates for cosidriver crd
The changes made to cosi driver are reflected. Add the cosidriver crd to
predicate workflow.

Fixes: #12545
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2024-01-29 22:15:21 +05:30
travisn 995a64fb3a exporter: skip reconcile on exporter deletion
The exporter pod is ephemeral and frequently being deleted
and re-created when ceph daemons are created. Therefore,
we need to skip reconciling based on the deletion
of the exporter deployment.

Signed-off-by: travisn <tnielsen@redhat.com>
2024-01-19 14:49:44 -07:00
Travis Nielsen b693a5cca5 core: refactor crash collector for more node daemons
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md

Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-09 16:57:05 -07:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Madhu Rajanna b0fc7c9b92 namespace: add new CRD
This introduces a new CRD to add the ability
to create rados namespace for a given
ceph block pool. Typically the name of the pool
is the name of the blockpool created by rook.

Closes: #7035

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-04-05 10:10:04 +05:30
Sébastien Han 1208bc1410 core: reconcile operator configuration with env var
Previously, we were ignoring configuration changes coming from the
operator's pod env variables. We were assuming most users were using the
operator configmap "rook-ceph-operator-config" but most Helm users
don't. Now we will reconcile if a cephcluster is found during a CREATE
event (can be an operator restart or a cephcluster creation) AND no
"rook-ceph-operator-config" is found which means the operator's pod env
var are used.
Also, unit tests have been added (long due) for the predicate!

Closes: #9602, #9487, #9579
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-20 16:13:11 +01:00
Blaine Gardner bb0d0d39f6 Merge pull request #9384 from leseb/fix-7036
subvolumegroup: add new crd
2021-12-21 10:24:23 -07:00
Sébastien Han 6e9eb33782 subvolumegroup: add new crd
This introduces a new CRD to add the ability to create subvolumegroup
for a given ceph filesystem volume. Typically the name of the volume is
the name of the filesystem created by rook.

Closes: https://github.com/rook/rook/issues/7036
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 16:28:38 +01:00
Sébastien Han da76a772f8 core: disallow multiple clusters in the same namespace
Rook does not support running multiple clusters in the same namespace,
so the operator should not reconcile if a new cluster is added.
The scenario where a cluster is added while the operator is down is also
handled. CR updates are also handled.
When the operator detects more than one cluster it will refuse to
reconcile the CephCluster, and child CRDs will block too until the
operator is ready.
The user must remove one of the clusters before can continue to perform
any reconcile.

Closes: https://github.com/rook/rook/issues/9452
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-17 10:56:15 +01:00
Yuval Lifshitz f2b065a3ae rgw: refactor bucket notification code
this should make sure code reuse between the OBC label controller and
the CephBucketNotification controller.
done as part of the cleanup work from:
https://github.com/rook/rook/pull/8426

this also include adding unit tests for:
- topic controller
- notification controller
- obc label controller

Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-12-07 18:24:46 +02:00
Yuval Lifshitz 71ed45b69b rgw: implement bucket notifications for object storage
following the design from here:
https://github.com/rook/rook/blob/master/design/ceph/object/ceph-bucket-notification-crd.md

Closes: https://github.com/rook/rook/issues/5313
Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-11-04 11:20:40 +02:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Blaine Gardner 386eeb7eb9 ceph: fix detection of delete event reconciliation
Fix a bug where delete events on Rook-Ceph CRs would be detected
multiple times if deletion was blocked. Old method used pointer
comparison instead of using the (metav1.Time) Equal() method.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-08 16:48:53 -06:00
parth-gr bb23622b72 core: enhancement in debug logs
When DEBUG logging is enabled in the operator, there are a number of overwhelming messages in the log that are very overwhelming and don't seem useful, which makes debug mode difficult to use
Updated and Cutted down the messages that are not useful in debug mode

Closes: https://github.com/rook/rook/issues/7499
Signed-off-by: parth-gr <paarora@redhat.com>
2021-04-28 22:28:10 +05:30
Blaine Gardner 9d657f464c ceph: redact secret info from reconcile diffs
To avoid logging sensitive information, do not output diffs from secrets
when reconciling resources.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-15 08:14:34 -06:00
Sébastien Han c0123cf182 ceph: add cephfs mirroring support
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.

The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050

The same goes for integration tests, they will get added later once we
start testing on Pacific.

Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-28 19:21:18 +01:00
Sébastien Han ebbf332d8d ceph: update rgw and mds deployment for logCollector
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.

Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-22 18:19:07 +01:00
Sébastien Han 758226294b ceph: convert CephClient CRD to controller-runtime
This was the last remaining CRD to not use the controller-runtime
library.
Small additions were added with the transition:

* the Kubernetes Secret that contains the CephX key has now an owner
  reference to the CephClient object
* the secret name is present in the Status field of the CephClient:

```
status:
  info:
    secretName: rook-ceph-client-glance
  phase: Ready
```

The controller will reconcile on CR updates and also if the Kubernetes
Secret is deleted.

Closes: https://github.com/rook/rook/issues/4938
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-15 16:17:38 +01:00
Sébastien Han 70c6752e0e ceph: add dedicated predicate for CephCluster object
Since we want to pass a context to it, let's extract the logic into its
own predicate.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-24 18:21:07 +01:00
Sébastien Han 9b9f45b792 ceph: export isDoNotReconcile function
Since the predicate for the CephCluster object will soon move into its
own predidacte we need to export isDoNotReconcile so that it can be
consummed by the "cluster" package.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-24 18:21:02 +01:00
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
Sébastien Han bd90f9afe9 ceph: add more error handling and debug logging
The predicate was lacking from very useful errors as well as logging
messages.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Sébastien Han e7043edbd8 ceph: reconcile if cm is config override
Earlier, we were returning only when the object was not the config
override configmap, we want the opposite.
Also, this was blocking all subsequent conditions and add the correct
check to avoid cm to reconcile.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
Sébastien Han 56954d5cc8 ceph: fix CRs not reconciling on updates
Most of the CRs expect CephCluster and CephBlockPool were not
reconciling on updates. The predicate must return true on CR diff.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-15 11:15:33 +02:00
Ali Maredia fc579f4520 ceph: initial commit for ceph rgw multisite resources
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone

Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-06-09 16:28:53 -04:00
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Sébastien Han c1b8a7aa9e ceph: extract rbd-mirror to its own crd
Previously, the rbd-mirror daemon was integrated into the `CephCluster`
CRD. This wasn't really practical since we would have to wait for the
whole orchestration to be done to actually set it up. The same goes for
any CR update. Let's say you want to change the number of daemons, Rook
would go through mons, mgrs and osds until it get to rbd-mirror.
This triggers an undesired full orchestration.

With its own CRD this component just gains a lot more flexibility.

Closes: https://github.com/rook/rook/issues/5084
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-15 09:12:56 +02:00
Travis Nielsen ff1ab64d8d ceph: skip reconcile events where the spec not updated
The controllers should ignore events for status updates. The reconcile
only needs to happen when the spec is updated or the resource is marked
for deletion. Otherwise, the controllers may stay in an update loop as
the reconcile events are triggered again every time there is a status
update during the reconcile.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-09 11:20:45 -06:00
Sébastien Han 71799682ca ceph: print the reason for reconciling
General debugging helper without having to turn on DEBUG logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-06 16:07:01 +02:00
Sébastien Han c5075d55d1 ceph: trigger update on child controllers
When the CephCluster CR gets updated with a new image version, we need
to notify the child controllers of that upgrade.
For this, we set a 'ceph_version' label on the CR itself which will
effectively trigger a reconcile based on an update event.
This needs to be revisited and see how we can make use of
`EnqueueRequestsFromMapFunc` which might be a better approach using
controller-runtime built-in.

Closes: https://github.com/rook/rook/issues/5153
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-03 13:56:49 -06:00
Travis Nielsen 96dd92cc7f ceph: change verbose info log to debug in controller
Keeping the operator log clean...

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-03 09:15:25 -06:00
Travis Nielsen 0e55aa348b tests: reduce unhelpful debug logging
Removed some extremely verbose debug logging that makes
the logs difficult to read when debug is enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-30 13:32:47 -06:00
Sébastien Han 0e2f84f998 ceph: convert NFS controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4941
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-20 17:38:03 +01:00
Sébastien Han ca0a30f38d ceph: convert Filesystem controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 23:34:58 +01:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00