Commit Graph
11 Commits
Author SHA1 Message Date
Travis Nielsen 2ff429e479 csi: remove version check for k8s and cephcsi
The version checks for the csi driver are removed now
since they are all obsolete. The K8s version and cephcsi
versions are no longer checked. Anyway, the move to the
csi operator would take ownership of version checks
needed in the future, so for now we simplify rook
deployment of the csi driver.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-10-30 14:03:51 -06:00
Blaine Gardner 5539eedd1b multus: finish deprecating holder pods
Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-10-23 16:29:02 -06:00
Blaine Gardner 4f555dbbcb csi: allow force disabling holder pods
Add new CSI_DISABLE_HOLDER_PODS option for rook-ceph-operator.
This option will disable holder pods when set to "true".

In the long term, Rook plans to deprecate the holder pods entirely.
This new option will allow users to choose to migrate their clusters to
non-holder clusters when they are ready and able, giving them time to
gracefully migrate before the holders are permanently removed.

This option is set to "false" by default so that upgrading users don't
have their CSI pods modified unexpectedly.
Example manifests are modified to set this value to true so that new
clusters will not deploy holder pods.

Migrating users are provided with documentation to instruct them about
the new requirements they need to satisfy to successfully remove holder
pods, a procedure for migrating pods from holder to non-holder mounts,
and a way to delete holder pods once they are no longer in use.

When users set CSI_DISABLE_HOLDER_PODS="true", the CSI controller will
no longer deploy or update the holder pod Daemonsets, but it does not
delete any existing Daemonsets. This allows already-attached PVCs to
continue operating normally with their network connection continuing to
exist in the current holder pod. This is critical to avoid causing
ia cluster-wide storage outage.

More info: https://github.com/rook/rook/issues/13055

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-03-14 09:53:08 -06:00
parth-gr f007f2aca1 core: report node metrics using ceph telemetry
Add this reporting with the cephcluster reconcile,
Similar way we reported other telemetry's

Closes: https://github.com/rook/rook/issues/12344

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-16 19:13:22 +05:30
Blaine Gardner 17f0072d9d Merge pull request #12778 from BlaineEXE/multus-allow-cidr-spec
multus: allow using NADs without inspectable CIDRs
2023-09-07 13:42:41 -06:00
Blaine Gardner 3c43268d0a multus: detect network CIDRs via canary
Change how Rook detects network CIDRs for Multus networks. The IPAM
configuration is only defined as an arbitrary string JSON blob with a
"type" field and nothing more. Rook's detection of CIDRs for whereabouts
had already grown out of date since the initial implementation.
Additionally, Rook did not support DHCP IPAM, which is a reasonable
choice for users. And more, Rook did not support CNI plugin chaining,
which further complicates NADs. Based on the CNI spec, network chaning
can result in any changes to network CIDRs from the first-given plugin.

All these problems make it more and more difficult for Rook to support
Multus by inspecting the NAD itself to predict network CIDRs. Instead,
it is better for Rook to treat the CNI process as a black box. To
preserve legacy functionality of auto-detecting networks and to make
that as robust as possible, change to a canary-style architecture like
that used for Ceph mons, from which Rook will detect the network CIDRs
if possible.

Also allow users to specify overrides for CIDR ranges. This allows Rook
to still support esoteric and unexpected NAD or network configurations
where a CIDR range is not detectable or where the range detected would
be incomplete. Because it may be impossible for Rook to understand the
network CIDRs wholistically while residing only on a portion of the
network, this feature should have been present from Multus's inception.

Improving CIDR auto-detection and allowing users to specify overrides
for auto-detected CIDRs rounds out Rook's Multus support for CephCluster
(core/RADOS) installations. No further architectural changes should be
needed for CephClusters as regards application of public/cluster network
CIDRs for Multus networks.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-09-07 10:12:55 -06:00
subhamkrai a755ec156a csi: enable csi-addons-side when crds are deployed
when deploying csi, we'll check if csi-addons crds
are deployed we'll enable csi-addons sidecar by default.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-31 11:28:49 +05:30
Madhu Rajanna d7a4f0038c csi: set ceph cluster as ControllerRef for holder
Setting ceph cluster as the ControllerRef
for the holder daemonset set so that when
a ceph cluster is deleted the holder daemoset
will also gets deleted.

fixes: #12645

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2023-08-22 10:41:17 +02:00
Eng Zer Jun 8a25b1d903 test: use T.Setenv to set env vars in tests
This commit replaces `os.Setenv` with `t.Setenv` in tests. The
environment variable is automatically restored to its original value
when the test and all its subtests complete.

Reference: https://pkg.go.dev/testing#T.Setenv
Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
2022-07-15 23:07:19 +08:00
Sébastien Han 73b1347675 core: fix csi-cephfsplugin pod restart on non-hostnetworking env
Implementation of the design proposed in
https://github.com/rook/rook/pull/9903.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 17:53:01 +02:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00